timodonnell/mutinfer

experiments in mutation effect prediction

★ 0Forks 0PythonGitHub ↗Compare

README

mutinfer

Do frontier LLMs carry latent knowledge that can identify undiscovered cancer driver mutations?

Three preregistered experiments say no — under strict controls, on a task where the same pipeline separates canonical drivers from obvious passengers perfectly.

Results

# Experiment Contrast n pairs Effect 95% CI p
1 Pan-cancer TCGA doubleton vs singleton 89 0.539 0.430–0.646 0.263
2 BLCA tissue-concordant doubleton vs singleton 313 0.514 0.458–0.571 0.326
3 External validation recurs in MSK-IMPACT vs not 701 0.508 0.471–0.545 0.353

Effect = fraction of blinded pairwise judgments favouring the hypothesised variant; 0.5 is chance. Full write-up in results/RESULTS.md.

Conclusion. Given a gene symbol and an amino-acid substitution, a frontier LLM shows no detectable ability to distinguish somatic missense variants that recur in an independent 44,150-patient cohort from those that do not, once gene identity and mutational context are held fixed.

This is a claim about a hard task under strict controls, not about the model's biology knowledge. The same pipeline separates canonical drivers (BRAF V600E, KRAS G12D) from olfactory-receptor passengers at AUC = 1.000. What it cannot do is rank within a gene and mutational channel — which is exactly where undiscovered driver signal would have to live.

The methodological result

Experiment 3's naive form would have produced a confident false positive. External recurrence is dominated by mutability and ascertainment, not biology:

Factor External recurrence rate
CpG C>T vs not 64.6% vs 15.3%
Substitution class C>T 40.3% → T>A 3.4%
MSK panel coverage 27.6% (54,331 samples) vs 7.3% (7,817)

And the confounding pathway is real in our data: the LLM scores CpG variants higher than non-CpG (11.01 vs 10.26, p = 0.003), while CpG variants recur 4.2× more often. Arginine codons contain CpG, and R→W/Q/H are precisely the substitutions an LLM associates with drivers. Regressing recurrence on LLM score without matching recovers mutability and calls it biology.

Matching cases and controls exactly on (gene × SBS96 trinucleotide channel) removes that pathway by construction — and the effect disappears.

Design invariants

Enforced by 39 tests, not just documented:

  1. Recurrence identity is genomic — (chrom, pos, ref, alt), never (gene, protein_change).
  2. Counted by unique participant (barcode chars 1–12), after collapsing repeated records.
  3. No known cancer-driver genes in any cohort — gene-level exclusion.
  4. Same-gene matched pairs, always; Experiment 3 additionally matches on exact trinucleotide channel.
  5. Models are blinded. Inference consumes only eval_variants_blinded_*.jsonl (variant_id, gene, protein_change). IDs are opaque blake2b digests; row order is shuffled; llm_io.assert_no_leakage gates every prompt.
  6. Inference is at pair level, never per API call.
  7. Generator ≠ judge model family (self-preference bias).

Pairs are not matched on protein position or domain — that is where the signal under test may live.

Pipeline

make data      # download + parse MC3, driver blacklist, annotate variants
make cohort    # match pairs, freeze cohorts, QC report + figures
make test      # 39 tests

Then per experiment (see Makefile and results/manifest_*.json for exact invocations):

PYTHONPATH=src python3 src/run_generator.py --tag <cohort>
PYTHONPATH=src python3 src/run_pairwise_judge.py --tag <cohort>
PYTHONPATH=src python3 src/analyze.py --tag <cohort>
Stage Script
Acquire MC3 src/download_mc3.py
External refs + TCGA case map src/download_external.py
Parse, per-participant burden src/parse_mc3.py
Driver-gene blacklist src/build_driver_gene_blacklist.py
Candidate variants, SBS96 src/build_candidate_variants.py
Matching / freezing src/match_pairs.py, src/freeze_dataset.py
Calibration controls src/build_controls.py
External cohort (MSK-IMPACT) src/fetch_external_cohort.py, src/build_external_validation_set.py
Inference src/run_generator.py, src/run_{absolute,pairwise,variant_only}_judge.py
Analysis src/analyze.py, src/analysis_plots.py, src/plotting.py

Cohorts

All frozen under data/frozen/ with SHA256 manifests.

Tag Description Pairs
v1_nohyper pan-cancer, hypermutator carriers excluded (primary, Exp 1) 100
v1 pan-cancer, hypermutators retained (robustness, not run) 100
blca bladder only, tissue-concordant doubletons (Exp 2) 345
extval MSK-IMPACT recurrence, matched on gene × SBS96 (Exp 3) 808
controls 12 canonical drivers + 12 olfactory-receptor passengers —

Data sources

Resource Version Access
TCGA MC3 mc3.v0.2.8.PUBLIC.maf.gz, GRCh37 GDC UUID 1c8cfe5f-e52d-41ba-94da-f15ea1337efc, open
MSK-IMPACT 50K msk_impact_50k_2026, GRCh37 cBioPortal API, open
IntOGen 2024-06-18 Compendium (633 genes) open
cancerhotspots v2 (240 genes) open
OncoKB Cancer Gene List — requires ONCOKB_TOKEN (not applied)
COSMIC CGC — requires login (not applied)

MC3 and MSK-IMPACT are both GRCh37, so no liftover is needed; 1,009 MSK rows carrying GRCh38 or no build are dropped rather than coordinate-mismatched.

Models: generator moonshotai/Kimi-K3, judge deepseek-ai/DeepSeek-V4-Pro-0813 (chosen by a bake-off on the controls), both via Together AI. Total spend across all work: $248.99 over 20,615 API calls.

Reproducibility

Every stage writes results/manifest_<stage>.json with git commit, input hashes, seed and timestamp. Frozen datasets and prompts carry SHA256 checksums (data/frozen/freeze_manifest_*.json, data/frozen/prompt_freeze.json). Seed is 20260904. All raw model outputs are committed under results/{generations,judgments}/; src/rebuild_cache.py reconstructs the API cache from them without spending anything.

Known limitations

  • Driver blacklist is 686 genes (IntOGen ∪ cancerhotspots). OncoKB and COSMIC CGC are credential-gated; the OncoKB sensitivity analysis was not run.
  • Single generator and single judge model family. A second generator would separate "impossible from these inputs" from "this model can't do it".
  • Cohort v1 (hypermutators retained) is frozen but was not run.
  • The Phase 21 memorisation audit was not run; for Experiment 3 the shortcut is closed by construction, since every cancerhotspots gene is already blacklisted.
  • Experiment 3's judge showed strong position bias (slot A won 57.7% of decided judgments). It does not confound the contrast — A/B is randomised independently of condition — but any future pairwise design must randomise against it.

Contributors

timodonnell

Issues