This repository aims to formalize the results from the manuscript contained in paper/Finding_the_threshold-rev4.tex using Lean 4.
paper/: source.texof the target paper.Formalization/: Lean source files for the project (currently a placeholder).Formalization.lean: entry point for the Lean library.
The project uses Lean v4.24.0-rc1 (see lean-toolchain). The plan below records the steps needed to turn the paper into a complete Lean development and tracks how intermediate claims will be verified mechanically.
- After the container launches there is no internet access. Commands such as
lake updateandlake exe cache getwill fail inside the session. Record any dependency changes directly inlakefile.leanand notify the maintainer so they can refresh caches outside the container. - Treat
lake-manifest.jsonas an immutable lockfile. Never revert or "clean up" this file, even if it appears modified—leave the tracked version untouched unless a maintainer explicitly supplies a replacement manifest produced outside the container.
The roadmap below interleaves construction steps with explicit strategies for verifying every intermediate definition, lemma, and theorem in Lean. Each stage identifies the concrete checks, supporting automation, and sanity tests that guarantee the mechanized statements coincide with those of the paper.
- Add
mathlibas a dependency inlakefile.lean. - Confirm with the maintainer that
lake updatesucceeds outside the container whenever new dependencies are added (internet access is unavailable within the session, so do not attempt the command locally). - Establish a project-wide namespace (currently
Codex) and record linting/CI commands (lake build,lake exe cache getif needed). - Create a scratch file (
Formalization/Scratch.lean) for experiments before incorporating statements into the main hierarchy. - Verification approach.
- Maintain a
#check/#evalscratchpad for each new definition before moving it into the hierarchy. - Refer to the Verification Checklist for linting guidance that applies across all stages.
- Maintain a
Status (Stage 0): The namespace Codex now lives in Formalization/Basic.lean, accompanied by a trivial SimpleGraph sanity check so future imports confirm access to mathlib. A dedicated scratchpad (Formalization/Scratch.lean) is available for experiments. Because the interactive container has no internet connectivity, dependency refreshes must be coordinated with the maintainer outside the session; record any required updates in lakefile.lean and leave the lake update checkbox unchecked until confirmation arrives.
Goal: formalize the deterministic combinatorial ingredients used across the paper.
Key tasks and Lean checks:
-
Finite simple graphs.
- Use
SimpleGraphfrommathlibwith vertex typeFin nand define helper constructors for labelled graphs onFin nwith explicit edge sets. - Verify basic facts about edge counts (
e(H)) and subgraph inclusion with Lean lemmas, tagging the most important ones with@[simp]or@[simp, aesop_safe]for later automation. - Create test lemmas (e.g.,
example : e (completeGraph (Fin 3)) = 3 := by decide) to confirm the definitions match combinatorial intuition.
- Use
-
Copy counting (
M_{H', H}).- Define
countCopies (H' H : SimpleGraph (Fin n)) : ℕas the number of embeddings ofH'intoH(usingSimpleGraph.embedding), quotienting by automorphisms if needed. - Finish the double-counting probability identity
π_H(J₀ ⊆ 𝐇) = M_{J,H} / M_Jin Lean. (Completed via the new lemmauniformProbability_contains_fixedwhich combines the constant-fibre argument withembeddingPairsEquiv.) - Validate the combinatorial identities with small examples (
Fin 3,Fin 4) inside Lean usingdec_trivialorsimp [countCopies]to ensure the formulas have the correct normalization factors.
- Define
-
Monotonicity and edge-induced subgraphs.
- Formalize the edge-induced subgraph construction and show that the number of edges is preserved as required.
- Provide automation lemmas showing the closure of subgraphs under intersection/union when needed for counting arguments.
- Use Lean's rewriting tools (
by_cases,simp,finset.induction) to verify every structural property, recording each as a lemma reusable in later stages.
Status (Stage 1): Stage 1 utilities in Formalization/Stage1/FiniteSimpleGraphs.lean now build graphs from explicit edge sets and prove the foundational edge-count lemmas (including monotonicity of edgeCount and the n.choose 2 formula for complete graphs). Edge-induced subgraphs, together with union/intersection closure lemmas and finite edge-count computations, are available to support the upcoming copy-counting and subgraph arguments. The copy-counting API confirms that isomorphic pattern or host graphs yield identical enumerations of labelled embeddings, and the new permutation construction leads to the uniform probability lemma uniformProbability_contains_fixed, completing the Stage 1 double-counting identity.
Goal: formalize the probabilistic objects (G(n,p)) and compute expectations used in the thresholds.
Lean tasks:
-
Probability space for
G(n, p).- Model
G(n, p)as the product measure on edge indicators. UseSimpleGraphand random edge subsets, employingMeasureTheoryandProbabilityAPIs inmathlib. - Define
gnp (n : ℕ) (p : ℝ)returning a random variable valued inSimpleGraph (Fin n). - Confirm measurability and integrability obligations explicitly with Lean proofs (
measurable_gnp,integrable_countCopies) and tag the statements with documentation notes referencing the paper. (These are provided byStage2.measurable_gnpandStage2.integrable_countCopies.)
- Model
-
Random variables counting subgraphs.
- Define
countCopiesRVfor a fixed patternH'and prove measurability/integrability of the associated real-valued random variable. - Verify the expectation formula for edges by proving
expected_countCopies_completeGraph_two. - Extend the expectation check to triangles via
expected_countCopies_completeGraph_three. - Generalize the expectation statement to an arbitrary finite graph
H'and package it as a reusable lemma. - Use the independence API to state tail bounds for fixed thresholds (
ℙ[Z_{H'} ≥ t]).
- Define
-
Tail bounds via Markov.
- Port or restate Markov's inequality from
mathlibin the form needed forZ_{H'}. - Instantiate the inequality for the random variables above to derive the inequalities
p_E ≤ p_Etilde ≤ p_crit. - Capture the instantiated inequalities as Lean lemmas (
pE_le_pEtilde,pEtilde_le_pCrit) and add@[simp]orlemmawrappers to make them directly reusable in Stage 3.
- Port or restate Markov's inequality from
Status (Stage 2): Formalization/Stage2/RandomGraph.lean now develops the Bernoulli product measure over edge indicators, defines gnp, and proves measurability/integrability of countCopiesRV. The file establishes cylinder-measure formulas for finitely many edge assignments and evaluates expectations for K₂ and K₃ embeddings, leaving the general expectation theorem and the subsequent Markov-based tail bounds for future work.
Goal: encode (p_E(H)), (\tilde{p}E(H)), and (p\mathsf{c}(H)) as Lean definitions and derive the easy inequalities.
Lean tasks:
-
Define thresholds.
- Implement
pE H,pEtilde H, andpCrit Hasℝdefined viaInf/Supover sets of parameters satisfying the corresponding probability or expectation constraints. - Show that the
Infis achieved for nonempty sets by proving nontriviality (e.g., using0 ≤ p ≤ 1). - Immediately prove characterization lemmas (
lemma pE_def) that rewrite each definition into the equivalent inequality from the paper, ensuring the Lean definition matches the informal one.
- Implement
-
Prove inclusion bounds.
- Encode the arguments from §2.1 (Markov-based inequalities) to show
pE ≤ pEtildeandpEtilde ≤ pCrit. - Verify the algebraic rewritings such as
(1 / (2 * M_{H'}))^(1 / e(H'))in Lean, ensuring hypotheses handle the nonempty subgraph case (e(H') ≥ 1). - Build automation lemmas using
gcongr,nlinarith, andpositivityto verify inequalities without manual rewriting, and capture intermediate results as@[simp]theorems where appropriate.
- Encode the arguments from §2.1 (Markov-based inequalities) to show
-
Continuity and monotonicity lemmas.
- Prove helper results to handle the
max/minformulations: e.g.,pE H = sup_{H' ⊆ H} .... - Sanity-check these results on enumerated finite graphs (e.g., compute
pEfor paths of length 2) using Lean'sby decideorinterval_casesto ensure the definitions produce reasonable values.
- Prove helper results to handle the
Goal: mechanize the probabilistic combinatorial lemma underpinning the main theorem.
Lean tasks:
-
Statement preparation.
- Define the notion of an
R-spread distribution over subgraphs using Lean's probability kernels. - Relate it to a finite list of subgraphs
G₁, …, G_Mby taking the uniform measure on that list. - Verify equivalence with the paper's definition by proving bidirectional lemmas (
spread_iff_uniform_support) and tagging them with@[simp]/@[iff]as appropriate.
- Define the notion of an
-
Leverage existing results.
- Formalize or port
Theorem 1.6fromfracKK_annalsif available; otherwise, implement the argument sketched in the paper by combining concentration bounds for edge counts with the black-boxspreadinequality. - Each imported result should be restated and proved in Lean, verifying prerequisites (e.g., Chernoff bounds) from
mathlib. - Use Lean's rewrite tactics to double-check the hypotheses line up: provide convenience lemmas that translate between Lean's statements and the constants/notation in the paper.
- Formalize or port
-
Final lemma proof.
- Assemble the above to prove Lemma 2.1 exactly as stated, keeping constants explicit (
∃ Csuch that …). - Introduce helper lemmas that isolate each probabilistic estimate and unit-test them by instantiating with toy parameters in Lean, ensuring the inequality structure is correct before combining them.
- Assemble the above to prove Lemma 2.1 exactly as stated, keeping constants explicit (
Goal: combine all previous stages to formalize the main inequality.
Lean tasks:
-
Define the uniform distribution
π_H.- Express
π_Hover copies ofHusingFintypeenumerations and show it isR-spread withR = 1 / (2 * pEtilde H). - Provide the Lean proof of the double-counting identity
π_H(J₀ ⊆ 𝐇) = M_{J,H} / M_J. - Verify the normalization (
∑ π_H = 1) usingsimplemmas and add a smallexamplewith a concrete graph confirming the probability measure is valid.
- Express
-
Apply the spread lemma.
- Instantiate Lemma 2.1 with
k = e(H)andR = 1 / (2 * pEtilde H)to deduce the desired bound. - Translate the result back into the statement about thresholds (
pCrit H ≤ L * pEtilde H * log (e(H))). - Record each translation as a named lemma (
main_theorem) and provide a structured proof script that clearly references the supporting lemmas, making it easy to audit the Lean derivation.
- Instantiate Lemma 2.1 with
-
Document constant dependencies.
- Track universal constants inside Lean proofs to output an explicit
L. Store them in a dedicated namespace for reuse. - Confirm constant propagation by writing Lean lemmas that show each constant remains positive/bounded; these serve as automated regression tests when constants change.
- Track universal constants inside Lean proofs to output an explicit
- Hamiltonian cycle calculation. Provide a Lean computation (or estimate) establishing
pEtilde(H) ≍ 1/nfor Hamiltonian cycles, mirroring the remark in the paper. Cross-check the computation with a small#evalornorm_numverification. - Bibliographic remarks. Optionally, create Lean doc-strings referencing the relevant literature for maintainability.
- Automation health checks. Use
#synth/#simp?andlibrary_searchwithin doc-strings to record the tactics that succeed on intermediate results, ensuring future refactors can reproduce the same proofs.
- Each stage introduces definitions and lemmas that should be accompanied by Lean proofs; placeholders (e.g.,
by admit) should be avoided in the final development. - Attach validation lemmas/examples to every new definition to show it behaves correctly on toy instances.
- After significant additions, run
lake build(andlake testif a test harness is added) to ensure the code compiles. - Maintain documentation within Lean files (
/-! ### ... -/blocks) describing the relationship between the formal proofs and the paper's arguments. - Use
#lintto catch missingsimp/instancelemmas and guarantee all intermediate statements are fully verified.
- Rebuild the Stage 1 double-counting probability identity in manageable steps:
- Re-establish the cardinality equality for embeddings containing a fixed copy using sigma-type bookkeeping.
- Upgrade the
uniformProbability_double_countlemma to the fully uniform statement once the cardinality step is stable.
- After Stage 1 is complete, begin Stage 2 by modeling
G(n,p)and introducing the associated expectation lemmas.
Progress and deviations from this plan should be recorded either in this README or in additional markdown notes within the repository.