Study repository: PeterPonyu/topic-model-endpoint-benchmark. Versioned archive: 10.5281/zenodo.22914505. Use CITATION.cff for release 1.0.0; release-manifest.json records distributed-file checksums. The archive and source repository describe this study only.
Author-owned software is MIT licensed. The author's manuscript, figures, generated results and model weights are CC BY 4.0. Source-study data, labels, annotations and other third-party materials retain their original terms; they are not relicensed. Read LICENSE and NOTICE.md before reusing mixed-content files.
This research package examines how annotation agreement, latent cluster geometry, and model-selection rules lead to different conclusions on the same retained single-cell benchmark. It contains the recorded score tables, frozen descriptive protocol, numerical analyses, independent checks, plotting code, and manuscript source. Reproduction uses these local scores and performs no model training.
The sixteen CSV files in analysis/inputs/ each contain one retained evaluation for Pure-VAE, Topic-FM-Base, Topic-FM-Transformer, and Topic-FM-Contrastive. The analysis uses 256 finite values: sixteen datasets × four configurations × four endpoints. Additional exported columns remain in the unchanged input files but are outside the primary analysis.
NMI (normalized mutual information) and ARI (adjusted Rand index) measure agreement with recorded annotations. ASW (average silhouette width) and DAV (Davies–Bouldin index) describe each model's own latent space using its K-means assignments. Higher NMI, ARI, and ASW and lower DAV indicate improvement on their respective scales. Geometry scores do not independently validate cell identity or biological mechanisms.
The dataset identifiers are blood_aged, bm, dentate, endo, hemato, hesc, hesc_times, ifnHSPC, lung, pansci_muscle, pansci_tcell, pituitary, retina, setty, spine, and teeth. A prespecified annotation-scope sensitivity excludes bm, hesc_times, ifnHSPC, and spine; their missing historical label provenance is not repaired by retaining their scores.
The frozen descriptive design in analysis/protocol.json specifies endpoint means, each of sixteen single-dataset deletions, comparisons with Pure-VAE, a leave-one-dataset-out variant selector, and the twelve-dataset annotation-scope restriction. DAV is sign-reversed only when orienting comparisons or rankings; its recorded values are unchanged.
The selector chooses among the three Topic-FM variants using the mean of (NMI + ARI + ASW) / 3 over the other fifteen datasets, with alphabetical tie resolution, then evaluates the omitted dataset. Its comparison with selecting the best variant on that same dataset quantifies selection optimism for this declared composite. The composite mixes annotation agreement and model-specific geometry and is not a validated biological utility.
Paired ranks, exploratory Friedman tests, a four-endpoint Holm adjustment, and the selector using (NMI + ARI) / 2 are explicitly post-hoc extensions. The two selector scores have different units and should not be compared as a common effect size. Pure-VAE leads the retained mean NMI and ARI, while Topic-FM-Transformer leads ASW and DAV; all sixteen deletions preserve those endpoint leaders.
Run the following from this directory. The first command independently checks the supplied numerical results before regeneration; the last command checks the regenerated results. Assertions must remain enabled: do not use Python's -O option.
python3 analysis/robustness_audit.py
python3 analysis/generate.py
python3 analysis/rank_sensitivity.py
python3 analysis/robustness_audit.pyThe generator writes analysis/outputs/summary.json, deletion.csv, and selection.csv; the rank calculation writes rank_sensitivity.json in the same directory. The independent audit reads the raw CSVs without importing either generator. It checks input SHA-256 values against the frozen protocol, all primary summaries, all 256 deletion records, all sixteen selector records, and all four rank tests. It creates validation/ when needed and writes robustness-audit.json and selection-without-asw.csv. A missing input, missing result, duplicate record, hash mismatch, or failed assertion stops the command with an error. No historical validation files are needed.
The numerical commands require Python 3, NumPy, pandas, and SciPy. They were checked with Python 3.13.5, NumPy 2.2.6, pandas 2.3.3, and SciPy 1.16.3. The analysis requires no GPU, training data download, network access, or optimizer updates.
Rscript --vanilla scripts/generate_figures.R
bash build.sh --checkThe R command reads the local inputs and numerical outputs, draws five data figures, and renders figures/architecture.svg as a sixth figure. It writes the figure PDFs and plot-data CSVs into analysis/outputs/. figure-build.json lists the figure inputs. Dependencies are R with ggplot2, gridExtra, jsonlite, Cairo support, and Arial; grid is included with R. The SVG renderer uses the system Python at /usr/bin/python3 with PyGObject, the Rsvg introspection binding, librsvg, and Cairo. The checked R stack was R 4.3.3, ggplot2 4.0.3, gridExtra 2.3, and jsonlite 2.0.0.
The manuscript requires XeLaTeX, BibTeX, Arial, and the TeX packages declared in paper.tex. bash build.sh --check compiles into build_check/paper.pdf without replacing the distributed PDF. To build a replacement PDF, run:
bash build.shThis writes build_independent/paper.pdf and copies it to paper.pdf. Both build modes record their file inputs and compilation logs in their output directory. No additional research files are required beyond this package; the software and fonts must be installed separately.
analysis/protocol.json retains the original freeze record, scientific rules, local input filenames, and input hashes. Its machine-specific source locations have been replaced with scientific descriptions. The recorded original protocol digest identifies the supplied bytes before that path transformation, not the edited public JSON. analysis/generator-provenance.json preserves historical source identities and unchanged input digests; it is not a checksum manifest for the current public code. AUDIT.md explains the independent calculations, checked assertions, and limits of the evidence.
The unit is one retained dataset evaluation per configuration. Related datasets, overlapping deletion panels, cells, donors, and model seeds cannot be counted as independent replications. Historical preprocessing, training exposure, matched encoder controls, and flow-off/on controls are incompletely recorded. The package has no per-cell latent arrays, assignments, donor identifiers, or seed replicates, so it cannot reconstruct the underlying metrics or establish causal architectural benefits, biological superiority, or population-level uncertainty. The analysis freeze precedes these score calculations; it does not establish prospective registration before the historical training runs.