TankTechnology/zgca-paper-list

A bilingual, traceable index of research outputs from Zhongguancun Academy and ZGCI.

★ 0Forks 0GitHub ↗Compare

README

Awesome ZGCA Papers

Daily update GitHub Pages

A bilingual, traceable index of research outputs from Zhongguancun Academy (北京中关村学院) and the Zhongguancun Institute of Artificial Intelligence (中关村人工智能研究院).

Public-source coverage is maximized, but absolute completeness cannot be guaranteed. Every item retains inspectable institution evidence.

Discovery strategy

  • Search structured affiliation metadata across Crossref, OpenAlex, Europe PMC, arXiv and DataCite.
  • Backfill papers explicitly announced by the Academy's official research feed, then resolve their arXiv IDs and DOIs.
  • Scan official project repositories for exact affiliation lines that are present in a paper PDF or README but absent from arXiv metadata.
  • Merge formal publications, preprints and companion datasets without deleting existing records when a source is temporarily unavailable.

Latest outputs

Data sources

Source Status
arxiv ok (0 matched)
bza_official ok (21 matched)
core optional key
crossref unavailable (HTTPError)
datacite ok (17 matched)
europe_pmc ok (0 matched)
github_projects ok (7 matched)
lens optional key
openalex ok (0 matched)
semantic_scholar optional key

Use the data

Local development

npm ci
python3 scripts/pipeline.py build
npm run dev

Run networked discovery with python3 scripts/pipeline.py fetch. Optional API keys are documented in .env.example and should be stored as GitHub Actions secrets.

Polite arXiv HTML backfill

Historical affiliation discovery uses the partner-university list published by bjzgcai as a structured prefilter. OpenAlex first selects papers involving those institutions and an arXiv location; only that reduced queue is allowed to request https://arxiv.org/html/<id>v1.

python3 scripts/arxiv_html_backfill.py prefilter --from 2024-06-01
python3 scripts/arxiv_html_backfill.py start
python3 scripts/arxiv_html_backfill.py status

The local SQLite checkpoint is resumable and ignored by Git. arXiv requests are single-connection, at least 3.5 seconds apart by default, take a five-minute rest every 500 requests, and back off automatically after throttling or server errors. Only HTML matches in the author-affiliation region are published automatically; body-text matches remain audit-only associations.

Corrections

Open an issue or pull request. Stable overrides live in data/overrides.yml and are never replaced by the automated pipeline.

License

Code is MIT licensed. Aggregated metadata remains subject to its original source terms and always retains provenance links.

Contributors

github-actions[bot]longxiang-ai

Issues