Do frontier LLMs carry latent knowledge that can identify undiscovered cancer driver mutations?
Three preregistered experiments say no — under strict controls, on a task where the same pipeline separates canonical drivers from obvious passengers perfectly.
| # | Experiment | Contrast | n pairs | Effect | 95% CI | p |
|---|---|---|---|---|---|---|
| 1 | Pan-cancer TCGA | doubleton vs singleton | 89 | 0.539 | 0.430–0.646 | 0.263 |
| 2 | BLCA tissue-concordant | doubleton vs singleton | 313 | 0.514 | 0.458–0.571 | 0.326 |
| 3 | External validation | recurs in MSK-IMPACT vs not | 701 | 0.508 | 0.471–0.545 | 0.353 |
Effect = fraction of blinded pairwise judgments favouring the hypothesised variant; 0.5 is chance. Full write-up in results/RESULTS.md.
Conclusion. Given a gene symbol and an amino-acid substitution, a frontier LLM shows no detectable ability to distinguish somatic missense variants that recur in an independent 44,150-patient cohort from those that do not, once gene identity and mutational context are held fixed.
This is a claim about a hard task under strict controls, not about the model's biology knowledge. The same pipeline separates canonical drivers (BRAF V600E, KRAS G12D) from olfactory-receptor passengers at AUC = 1.000. What it cannot do is rank within a gene and mutational channel — which is exactly where undiscovered driver signal would have to live.
Experiment 3's naive form would have produced a confident false positive. External recurrence is dominated by mutability and ascertainment, not biology:
| Factor | External recurrence rate |
|---|---|
| CpG C>T vs not | 64.6% vs 15.3% |
| Substitution class | C>T 40.3% → T>A 3.4% |
| MSK panel coverage | 27.6% (54,331 samples) vs 7.3% (7,817) |
And the confounding pathway is real in our data: the LLM scores CpG variants higher than non-CpG (11.01 vs 10.26, p = 0.003), while CpG variants recur 4.2× more often. Arginine codons contain CpG, and R→W/Q/H are precisely the substitutions an LLM associates with drivers. Regressing recurrence on LLM score without matching recovers mutability and calls it biology.
Matching cases and controls exactly on (gene × SBS96 trinucleotide channel) removes that pathway by construction — and the effect disappears.
Enforced by 39 tests, not just documented:
- Recurrence identity is genomic —
(chrom, pos, ref, alt), never(gene, protein_change). - Counted by unique participant (barcode chars 1–12), after collapsing repeated records.
- No known cancer-driver genes in any cohort — gene-level exclusion.
- Same-gene matched pairs, always; Experiment 3 additionally matches on exact trinucleotide channel.
- Models are blinded. Inference consumes only
eval_variants_blinded_*.jsonl(variant_id,gene,protein_change). IDs are opaque blake2b digests; row order is shuffled;llm_io.assert_no_leakagegates every prompt. - Inference is at pair level, never per API call.
- Generator ≠ judge model family (self-preference bias).
Pairs are not matched on protein position or domain — that is where the signal under test may live.
make data # download + parse MC3, driver blacklist, annotate variants
make cohort # match pairs, freeze cohorts, QC report + figures
make test # 39 testsThen per experiment (see Makefile and results/manifest_*.json for exact
invocations):
PYTHONPATH=src python3 src/run_generator.py --tag <cohort>
PYTHONPATH=src python3 src/run_pairwise_judge.py --tag <cohort>
PYTHONPATH=src python3 src/analyze.py --tag <cohort>| Stage | Script |
|---|---|
| Acquire MC3 | src/download_mc3.py |
| External refs + TCGA case map | src/download_external.py |
| Parse, per-participant burden | src/parse_mc3.py |
| Driver-gene blacklist | src/build_driver_gene_blacklist.py |
| Candidate variants, SBS96 | src/build_candidate_variants.py |
| Matching / freezing | src/match_pairs.py, src/freeze_dataset.py |
| Calibration controls | src/build_controls.py |
| External cohort (MSK-IMPACT) | src/fetch_external_cohort.py, src/build_external_validation_set.py |
| Inference | src/run_generator.py, src/run_{absolute,pairwise,variant_only}_judge.py |
| Analysis | src/analyze.py, src/analysis_plots.py, src/plotting.py |
All frozen under data/frozen/ with SHA256 manifests.
| Tag | Description | Pairs |
|---|---|---|
v1_nohyper |
pan-cancer, hypermutator carriers excluded (primary, Exp 1) | 100 |
v1 |
pan-cancer, hypermutators retained (robustness, not run) | 100 |
blca |
bladder only, tissue-concordant doubletons (Exp 2) | 345 |
extval |
MSK-IMPACT recurrence, matched on gene × SBS96 (Exp 3) | 808 |
controls |
12 canonical drivers + 12 olfactory-receptor passengers | — |
| Resource | Version | Access |
|---|---|---|
| TCGA MC3 | mc3.v0.2.8.PUBLIC.maf.gz, GRCh37 |
GDC UUID 1c8cfe5f-e52d-41ba-94da-f15ea1337efc, open |
| MSK-IMPACT 50K | msk_impact_50k_2026, GRCh37 |
cBioPortal API, open |
| IntOGen | 2024-06-18 Compendium (633 genes) | open |
| cancerhotspots | v2 (240 genes) | open |
| OncoKB Cancer Gene List | — | requires ONCOKB_TOKEN (not applied) |
| COSMIC CGC | — | requires login (not applied) |
MC3 and MSK-IMPACT are both GRCh37, so no liftover is needed; 1,009 MSK rows carrying GRCh38 or no build are dropped rather than coordinate-mismatched.
Models: generator moonshotai/Kimi-K3, judge deepseek-ai/DeepSeek-V4-Pro-0813
(chosen by a bake-off on the controls), both via Together AI.
Total spend across all work: $248.99 over 20,615 API calls.
Every stage writes results/manifest_<stage>.json with git commit, input
hashes, seed and timestamp. Frozen datasets and prompts carry SHA256 checksums
(data/frozen/freeze_manifest_*.json, data/frozen/prompt_freeze.json). Seed
is 20260904. All raw model outputs are committed under
results/{generations,judgments}/; src/rebuild_cache.py reconstructs the API
cache from them without spending anything.
- Driver blacklist is 686 genes (IntOGen ∪ cancerhotspots). OncoKB and COSMIC CGC are credential-gated; the OncoKB sensitivity analysis was not run.
- Single generator and single judge model family. A second generator would separate "impossible from these inputs" from "this model can't do it".
- Cohort
v1(hypermutators retained) is frozen but was not run. - The Phase 21 memorisation audit was not run; for Experiment 3 the shortcut is closed by construction, since every cancerhotspots gene is already blacklisted.
- Experiment 3's judge showed strong position bias (slot A won 57.7% of decided judgments). It does not confound the contrast — A/B is randomised independently of condition — but any future pairwise design must randomise against it.