A tiny reference manager backed by a Dropbox folder of PDFs and published as a static web page via GitHub Pages.
Live page: https://timodonnell.github.io/papersimreading/
A scheduled job scans a Dropbox folder for new PDFs, extracts each paper's
metadata (title, authors, journal, publication date, abstract, and a link to the
PDF) from public sources, and appends a record to data/references.json.
The web page (index.html) renders that JSON — searchable and
sortable, no build step.
Only real papers are published. A file is included only if it resolves to a
public identifier (a DOI or arXiv id). Anything else — books, grants, manuals,
reviewer copies, personal files — is skipped entirely and never linked. The list
of skipped files is remembered locally in a git-ignored .papersync-excluded.json
so those filenames never reach the repo and are not re-processed each run.
Link policy. Open-access and arXiv papers link to the free public copy (resolved via Unpaywall). Paywalled papers (a DOI, but no open copy) link to the user's own PDF in Dropbox. So a Dropbox share link is only ever generated for an identified paper — never for an unrecognized file. When an HTML edition exists (arXiv native HTML or a PubMed Central full-text page), it is linked too.
Dropbox folder (synced locally)
│ scan for *.pdf not yet recorded
▼
papersync.sync ──► extract DOI / arXiv id / title from the PDF
│ resolve metadata:
│ 1. Crossref by DOI
│ 2. arXiv API by arXiv id
│ 3. Crossref by filename-derived DOI (bioRxiv/Nature ids)
│ 4. Crossref by title
│ 5. LLM over first-page text (optional, needs API key)
│ no DOI/arXiv id? -> not a paper, excluded (never published)
│ link: OA/arXiv -> public copy; paywalled -> Dropbox
▼
data/references.json ──► git commit & push ──► GitHub Pages renders index.html
The Dropbox folder is synced to this machine by the Dropbox desktop client, so the job reads the local copy directly — no scraping, no API auth, no bulk downloads. The shared-folder URL and folder path are never committed; they live in a local config file.
- Install dependencies (poppler provides
pdftotext):pip install -r requirements.txt sudo apt-get install poppler-utils # if pdftotext is missing - Create the local config (outside the repo):
Set
mkdir -p ~/.config/papersimreading cat > ~/.config/papersimreading/config.json <<'JSON' { "papers_dir": "/home/you/Dropbox/your-folder", "crossref_mailto": "[email protected]", "anthropic_api_key": "", "generate_dropbox_links": true } JSON chmod 600 ~/.config/papersimreading/config.json
anthropic_api_keyto enable the LLM fallback for PDFs that no public source can identify. Leave it empty to stay API-only (no token cost).
python3 -m papersync.sync # process every new PDF
python3 -m papersync.sync --limit 20 # stop after 20 new papers
python3 -m papersync.sync --since 90 # only PDFs modified in the last 90 days
python3 -m papersync.sync --dry-run # list new PDFs, write nothingThe run is incremental and resumable: records are flushed to disk every few papers, and already-processed files are skipped next time. Files are tracked by content hash, so renaming or moving a PDF within the folder does not re-process it.
scripts/run.sh runs the sync and, if references.json changed, commits and
pushes. Install it as a daily cron job:
scripts/install-cron.sh # daily at 07:30
CRON_SCHEDULE="0 * * * *" scripts/install-cron.sh # hourlyBecause the PDFs live in this machine's local Dropbox sync, the job must run on
this machine (not a cloud agent). Logs go to papersimreading.log (git-ignored).
| Path | Purpose |
|---|---|
papersync/ |
the sync pipeline (config, PDF extraction, metadata lookups, links, store) |
papersync/refresh_links.py |
re-resolve pdf links for existing records (python -m papersync.refresh_links) |
papersync/refresh_html.py |
backfill HTML-version links (arXiv/PMC) for existing records |
data/references.json |
the reference database (source of truth for the page) |
.papersync-excluded.json |
local, git-ignored memory of non-paper files (kept out of the repo) |
index.html |
the GitHub Pages site, renders references.json client-side |
scripts/run.sh |
cron entry point: sync + commit + push |
scripts/install-cron.sh |
install/refresh the cron entry |