mohsen1/causal-ai-handbook

★ 0Forks 0GitHub ↗Compare

README

Causal AI

Download the book

From Intervention to Intelligent Agents

A practical learning book on causal inference, causal discovery, representation learning, foundation models, active experimentation, and world models

Edition: August 24, 2026
Audience: Technical builders, researchers, data scientists, and advanced students
Suggested prerequisites: Basic probability, linear algebra, machine learning, and Python


How to use this book

Causal AI is not one algorithm. It is a family of ideas for building systems that can reason about what would happen if the world were changed, rather than merely predicting what tends to accompany what.

This book develops the subject in five layers:

  1. Causal questions: What intervention or counterfactual are we asking about?
  2. Causal models: What variables, mechanisms, and assumptions represent the system?
  3. Identification and estimation: Can the answer be recovered from the available data, and how precisely?
  4. Learning causal structure and representations: How can parts of the model be learned rather than fully specified?
  5. Causal AI systems: How can agents, foundation models, experimental loops, and world models use these ideas safely?

The chapters are cumulative, but there are three sensible routes:

  • Builder route: Chapters 1–6, 8, 10–14.
  • Research route: Chapters 1–4, 7–12, 16.
  • Applied causal inference route: Chapters 1–6, 13–15.

Every chapter ends with a short set of exercises. The exercises are designed to test whether you can formulate assumptions and reason about failure modes, not merely repeat definitions.

Core rule: A causal estimate is only as credible as the causal question, study design, and assumptions that make it identifiable. A more powerful model cannot repair a fundamentally unidentified question.


Contents


Map of the field

flowchart LR
    Q["Causal question"] --> M["Causal model"]
    M --> I["Identification"]
    I --> E["Estimation"]
    E --> D["Decision or intervention"]
    D --> O["New observations"]
    O --> M

    M --> CD["Causal discovery"]
    O --> CRL["Causal representation learning"]
    CRL --> M
    CD --> M
    M --> WM["Causal world model"]
    WM --> A["Agent or planner"]
    A --> D

    classDef question fill:#FFF4D6,stroke:#D97706,color:#3B2F0B,stroke-width:2px;
    classDef model fill:#E9F2FF,stroke:#2563EB,color:#102A43,stroke-width:2px;
    classDef inference fill:#E8F8EF,stroke:#15803D,color:#123524,stroke-width:2px;
    classDef action fill:#FCE8F3,stroke:#BE185D,color:#4A102A,stroke-width:2px;
    classDef learning fill:#F2EAFE,stroke:#7C3AED,color:#2E1065,stroke-width:2px;

    class Q question;
    class M,WM model;
    class I,E inference;
    class D,A action;
    class O,CD,CRL learning;
Loading

Notation

Symbol Meaning
$X, Y, Z$ Random variables
$A$ or $T$ Treatment, action, or intervention variable
$Y(a)$ Potential outcome under treatment value $a$
$do(A=a)$ Intervention that sets $A$ to $a$
$P(Y\mid A=a)$ Observational conditional distribution
$P(Y\mid do(A=a))$ Post-intervention distribution
$G$ Causal graph
$pa(X)$ Parents of $X$ in a graph
$U$ Exogenous or unobserved variables
ATE Average treatment effect
CATE Conditional average treatment effect
SCM Structural causal model
DAG Directed acyclic graph

Part I — Foundations

1. What Causal AI Is

1.1 Prediction is not intervention

Suppose a company observes that customers who receive help from a support agent renew at a lower rate. A predictive model may correctly learn:

$$ P(\text{renewal}=1\mid \text{support}=1) < P(\text{renewal}=1\mid \text{support}=0). $$

It would be a mistake to conclude that support causes churn. Customers with severe problems are more likely to contact support. Problem severity affects both support usage and renewal. The observed relationship mixes at least two processes:

  • the selection process that determines who receives support;
  • the causal effect of support on later renewal.

The decision question is not “What renewal rate is observed among supported customers?” It is:

$$ P(\text{renewal}=1\mid do(\text{support}=1)) $$

compared with:

$$ P(\text{renewal}=1\mid do(\text{support}=0)). $$

The first is a conditional probability. The second describes a changed data-generating process.

This distinction appears anywhere an AI system chooses actions:

  • Which users should receive an offer?
  • Which deployment should be rolled back?
  • Which component should be restarted?
  • Which drug should a patient receive?
  • Which experiment should a scientific agent run?
  • Which tool should an LLM agent call?

A system that only learns observational regularities can work well while the environment remains similar and the system remains passive. Once it starts choosing actions, it changes the distribution it was trained on. Causality becomes part of the control problem.

1.2 Three kinds of query

It is useful to separate three increasingly demanding questions.

Association

What tends to occur together?

Example:

$$ P(Y\mid X=x) $$

This is the normal territory of supervised learning.

Intervention

What would happen to $Y$ if we forced $X$ to a chosen value?

Example:

$$ P(Y\mid do(X=x)). $$

This is needed for policy, experimentation, and planning.

Counterfactual

Given what happened to this specific unit, what would have happened under a different action?

Example:

$$ P(Y_{x'}\mid X=x, Y=y). $$

Counterfactuals combine observed evidence about a particular case with an imagined intervention. They require more model structure than population-level intervention queries.

flowchart LR
    A["Association\nWhat is observed?"] --> I["Intervention\nWhat if we act?"] --> C["Counterfactual\nWhat would have happened instead?"]

    classDef assoc fill:#E8F1FF,stroke:#2563EB,color:#102A43,stroke-width:2px;
    classDef intervene fill:#E8F8EF,stroke:#15803D,color:#123524,stroke-width:2px;
    classDef counter fill:#FCE8F3,stroke:#BE185D,color:#4A102A,stroke-width:2px;
    class A assoc;
    class I intervene;
    class C counter;
Loading

These are not merely different prompt types for the same predictive model. They require different semantics and, often, different data.

1.3 The main subfields

“Causal AI” is best treated as an umbrella term.

Causal inference

The graph or identifying assumptions are supplied. The goal is to estimate a causal effect, such as the effect of a treatment on an outcome.

Causal discovery

The goal is to learn aspects of causal structure from observational data, interventional data, multiple environments, temporal information, or structural assumptions.

Causal representation learning

The meaningful causal variables are not directly observed. The goal is to recover high-level causal state from images, video, text, or entangled sensor streams.

Causal decision-making

A causal model is used to choose interventions, policies, or experiments under uncertainty and cost.

Actual causation and explanation

The question concerns which events or mechanisms caused a particular observed outcome, rather than an average population effect.

Causal reasoning in agents

An LLM or other agent formulates hypotheses, invokes formal causal tools, designs interventions, and updates a model based on evidence.

Causal world models

A learned model represents state, mechanisms, actions, and changes well enough to support intervention prediction, counterfactual planning, and adaptation outside the training distribution.

1.4 What causality does not give you automatically

Causal language is often attached to systems that do not possess causal semantics. The following are insufficient on their own:

  • An attention map.
  • A feature attribution score.
  • A graph produced by an LLM from variable names.
  • A time-ordered correlation.
  • A model that predicts the next frame.
  • A model that generates plausible explanations.
  • A high score on a benchmark whose causal graph is already written in the prompt.

A causal claim needs a link between a target query and assumptions about how data were generated or how interventions operate.

1.5 A minimal causal AI contract

A credible causal AI system should make the following explicit:

  1. Target: What intervention, population, outcome, and time horizon are being studied?
  2. Model: What variables and mechanisms are represented?
  3. Assumptions: Which variables are observed? Which confounders may be hidden? Are cycles, interference, or selection present?
  4. Identification: Why can the target be computed from the available data?
  5. Estimation: Which statistical procedure is used, and how uncertain is it?
  6. Validation: Which falsification, sensitivity, and out-of-distribution tests were performed?
  7. Decision rule: How does the estimate become an action, accounting for cost, risk, and uncertainty?
  8. Monitoring: Which mechanisms or data regimes may have changed since estimation?

A foundation model may automate pieces of this contract. It does not eliminate the contract.

1.6 Running example: deployment reliability

Throughout the book, we will use a software deployment system as one example. Imagine these variables:

  • change_size: magnitude of a code or configuration change;
  • test_depth: amount of pre-deployment testing;
  • traffic: production traffic at deployment time;
  • dependency_health: health of external dependencies;
  • deploy: whether a deployment is performed;
  • incident: whether a production incident occurs;
  • rollback: whether the change is rolled back;
  • recovery_time: time to restore service.

A predictive incident model may find that rollbacks correlate with long recovery times. That does not imply rollbacks cause slow recovery. Severe incidents cause both rollback and long recovery.

A causal system asks more precise questions:

  • What is the effect of increasing test_depth on incident probability?
  • For this change, would delaying deployment until traffic falls reduce incident risk?
  • Which intervention would most efficiently distinguish a dependency failure from an application regression?
  • Did the rollback reduce recovery time for this incident relative to the counterfactual of not rolling back?

These questions require different estimands, assumptions, and evidence.

1.7 Exercises

  1. A model finds that people who take a particular medicine have worse health outcomes. Give two causal graphs consistent with this association but with opposite treatment effects.
  2. Classify each query as associational, interventional, or counterfactual:
    • “What is the probability of failure after large deployments?”
    • “What would failure probability be if all large deployments received extended testing?”
    • “Would yesterday’s failed deployment have succeeded if it had received extended testing?”
  3. Pick one AI product you know. Write its minimal causal AI contract in eight lines.
  4. Explain why a highly accurate next-state predictor can still be unsafe for planning.

2. Graphs, Independence, and Causal Structure

2.1 Directed acyclic graphs

A directed acyclic graph contains nodes and directed edges, with no directed path that returns to its starting node. An edge $X\rightarrow Y$ says that $X$ is a direct cause of $Y$ relative to the variables represented in the model.

“Relative to the variables represented” matters. A causal graph is an abstraction, not a complete map of reality. If several intermediate mechanisms are omitted, an edge can summarize their net direct relationship at the chosen level.

A DAG encodes two kinds of information:

  • causal ordering and direct mechanisms;
  • conditional independence restrictions on the observational distribution, under the causal Markov assumption.

If $X_1,\dots,X_d$ form a DAG $G$, the joint distribution factorizes as:

$$ P(x_1,\dots,x_d)=\prod_{i=1}^{d}P(x_i\mid pa(x_i)). $$

This factorization is probabilistic. Interpreting arrows causally requires additional semantics.

2.2 Three structures you must recognize

Chain

flowchart LR
    X --> M --> Y
    classDef node fill:#E8F1FF,stroke:#2563EB,color:#102A43,stroke-width:2px;
    class X,M,Y node;
Loading

$X$ influences $Y$ through mediator $M$. Without conditioning, information can flow along the path. Conditioning on $M$ blocks the path.

Fork, or common cause

flowchart LR
    Z --> X
    Z --> Y
    classDef conf fill:#FFF4D6,stroke:#D97706,color:#3B2F0B,stroke-width:2px;
    classDef obs fill:#E8F1FF,stroke:#2563EB,color:#102A43,stroke-width:2px;
    class Z conf;
    class X,Y obs;
Loading

$Z$ causes both $X$ and $Y$. This creates association between $X$ and $Y$ even if neither causes the other. Conditioning on $Z$ blocks the path.

Collider

flowchart LR
    X --> C
    Y --> C
    classDef collider fill:#FCE8F3,stroke:#BE185D,color:#4A102A,stroke-width:2px;
    classDef obs fill:#E8F1FF,stroke:#2563EB,color:#102A43,stroke-width:2px;
    class C collider;
    class X,Y obs;
Loading

$X$ and $Y$ both cause $C$. The path is blocked by default. Conditioning on $C$, or on one of its descendants, can create association between $X$ and $Y$.

Collider bias is one of the most common sources of error because ordinary predictive modeling encourages adding variables. In causal analysis, “control for more variables” is not a valid general rule.

2.3 A deployment collider example

Suppose both code defects and infrastructure instability cause incidents:

flowchart LR
    D["Code defect"] --> I["Incident observed"]
    S["Infrastructure instability"] --> I

    classDef cause fill:#E8F1FF,stroke:#2563EB,color:#102A43,stroke-width:2px;
    classDef collider fill:#FCE8F3,stroke:#BE185D,color:#4A102A,stroke-width:2px;
    class D,S cause;
    class I collider;
Loading

Across all deployments, defects and infrastructure instability may be independent. Among deployments selected because an incident occurred, they can become negatively associated: if evidence for a severe defect is weak, infrastructure instability becomes a more plausible explanation, and vice versa. An incident-response dataset that only contains failures is therefore selected on a collider.

2.4 Paths and d-separation

A path is a sequence of adjacent edges, ignoring their direction while tracing the route. A path is blocked by a conditioning set $S$ when either:

  • it contains a non-collider that is in $S$; or
  • it contains a collider such that neither the collider nor any descendant of it is in $S$.

Two sets of variables are d-separated by $S$ when every path between them is blocked. Under the Markov property, d-separation implies conditional independence.

You do not need to perform d-separation mechanically at first. Learn to identify chains, forks, and colliders, then trace open paths.

2.5 Markov and faithfulness assumptions

Causal Markov assumption

Each variable is independent of its non-descendants given its parents. This links graph structure to observational independences.

Faithfulness

The observed conditional independences arise from graph separation rather than exact cancellation of causal effects.

For example, $X$ could affect $Y$ through two paths with equal and opposite coefficients, producing zero total association despite a causal relationship. Faithfulness rules out such finely tuned cancellation.

Causal discovery algorithms often depend on both assumptions. Violations, small samples, measurement error, or weak effects can make graph recovery unstable.

2.6 Causal sufficiency and hidden confounding

A set of observed variables is causally sufficient when there is no unobserved common cause of any pair of included variables. This is a strong assumption.

When a hidden variable $U$ causes both $X$ and $Y$, the relationship is often drawn with a bidirected edge:

flowchart LR
    X["X"] <--> Y["Y"]
    classDef latent fill:#F2EAFE,stroke:#7C3AED,color:#2E1065,stroke-width:2px;
    class X,Y latent;
Loading

The bidirected edge is shorthand for an omitted common cause, not a direct two-way causal relationship.

Algorithms such as FCI aim to represent equivalence classes that allow latent confounding. Their output is more cautious than a single fully directed DAG.

2.7 Selection bias

Selection occurs when inclusion in a dataset depends on variables relevant to the causal question. Examples include:

  • analyzing only users who remained active;
  • analyzing only incidents that triggered an alert;
  • analyzing only patients who visited a hospital;
  • analyzing only experiments that completed;
  • using human-rated model outputs when selection into rating depends on model confidence.

Selection can open paths that were closed in the population. A causal model should include the selection process whenever possible.

2.8 Time and cycles

A DAG cannot contain a directed cycle at a single modeled time slice. Many real systems have feedback:

  • demand affects price and price affects demand;
  • load affects latency and latency changes load-balancing behavior;
  • an agent acts, observes, and acts again;
  • biological systems settle into equilibrium through feedback.

Three common responses are:

  1. Unroll time: represent $X_t\rightarrow Y_{t+1}\rightarrow X_{t+2}$.
  2. Model equilibrium carefully: use causal formalisms designed for cyclic or equilibrium systems.
  3. Choose a coarser intervention-level abstraction: represent mechanisms whose update sequence is not itself the target.

Do not force a cyclic system into a static DAG without explaining what the nodes and intervention semantics mean.

2.9 Graphs are assumptions, not decorations

A graph should be treated like executable documentation of assumptions. For every edge, ask:

  • What mechanism does it summarize?
  • At what time scale?
  • Could the direction be reversed?
  • Is there an omitted common cause?
  • Is this relationship stable across environments?
  • What intervention would test it?

For every absent edge, ask what independence or exclusion restriction is being asserted.

2.10 Exercises

  1. Draw a graph in which adjusting for a variable removes confounding.
  2. Draw a graph in which adjusting for a variable creates bias.
  3. In an incident dataset containing only incidents, identify at least two potential collider structures.
  4. Explain the difference between an edge $X\rightarrow Y$ and the statement that $X$ is statistically associated with $Y$ after controlling for other variables.
  5. Give an example of feedback that should be unrolled over time and an example where an equilibrium model may be more natural.

3. Potential Outcomes and Structural Causal Models

Causal inference has two dominant languages. They overlap substantially but emphasize different objects.

  • The potential-outcomes framework begins with outcomes under alternative treatments.
  • The structural causal model framework begins with variables generated by mechanisms and supports graphical intervention and counterfactual reasoning.

A strong practitioner should be able to translate between them.

3.1 Potential outcomes

For each unit $i$, define:

  • $Y_i(1)$: outcome if treated;
  • $Y_i(0)$: outcome if untreated.

The individual treatment effect is:

$$ \tau_i = Y_i(1)-Y_i(0). $$

Only one potential outcome is observed:

$$ Y_i = A_iY_i(1)+(1-A_i)Y_i(0). $$

This is the fundamental missing-data problem of causal inference. For the same unit at the same time, we cannot observe both treatment worlds.

Common population targets include:

Average treatment effect

$$ ATE = E[Y(1)-Y(0)]. $$

Average treatment effect on the treated

$$ ATT = E[Y(1)-Y(0)\mid A=1]. $$

Conditional average treatment effect

$$ CATE(x)=E[Y(1)-Y(0)\mid X=x]. $$

These are different estimands. A treatment can have a positive ATE, negative ATT, and strong variation in CATE if treated units differ systematically from the broader population.

3.2 Core assumptions in the potential-outcomes language

Consistency

If a unit receives treatment $A=a$, then its observed outcome equals $Y(a)$. This requires the treatment to be sufficiently well defined.

“Received support” may not be a single treatment if support quality, timing, channel, and content vary substantially.

No interference

One unit’s treatment does not affect another unit’s outcome. This often fails in networks, markets, epidemics, shared infrastructure, and multi-agent systems.

Consistency plus no interference is often summarized under SUTVA, although treatment-version issues deserve separate attention.

Conditional exchangeability or ignorability

$$ Y(a) \perp A \mid X. $$

Given covariates $X$, treatment assignment behaves as though random with respect to the potential outcomes.

This is not testable from observational data alone. It is a substantive claim that $X$ contains enough information to block confounding.

Positivity or overlap

$$ 0 < P(A=a\mid X=x) < 1 $$

for relevant $x$.

Every covariate profile in the target population must have a nonzero chance of receiving each treatment being compared. Without overlap, the data contain no direct support for the missing treatment condition in that region.

3.3 Structural causal models

An SCM consists of:

  • endogenous variables $V$;
  • exogenous variables $U$;
  • structural assignments $F$;
  • a distribution over $U$.

For example:

$$ \begin{aligned} S &:= f_S(U_S) && \text{severity} \\ A &:= f_A(S,U_A) && \text{support assignment} \\ Y &:= f_Y(A,S,U_Y) && \text{renewal} \end{aligned} $$

The assignments describe mechanisms. The graph contains edges from variables appearing on the right side of an equation to the variable on the left.

flowchart LR
    S["Problem severity"] --> A["Support"]
    S --> Y["Renewal"]
    A --> Y

    classDef conf fill:#FFF4D6,stroke:#D97706,color:#3B2F0B,stroke-width:2px;
    classDef treat fill:#E8F1FF,stroke:#2563EB,color:#102A43,stroke-width:2px;
    classDef out fill:#E8F8EF,stroke:#15803D,color:#123524,stroke-width:2px;
    class S conf;
    class A treat;
    class Y out;
Loading

3.4 Interventions as surgery

The intervention $do(A=a)$ replaces the structural assignment for $A$ with the constant assignment:

$$ A := a. $$

The incoming arrows into $A$ are removed. Other mechanisms remain unchanged unless the intervention is defined to change them.

flowchart LR
    S["Problem severity"] -. incoming edge removed .-> A["Support fixed to a"]
    S --> Y["Renewal"]
    A --> Y

    classDef removed fill:#F8FAFC,stroke:#94A3B8,color:#475569,stroke-dasharray:5 5;
    classDef fixed fill:#FCE8F3,stroke:#BE185D,color:#4A102A,stroke-width:2px;
    classDef normal fill:#E8F8EF,stroke:#15803D,color:#123524,stroke-width:2px;
    class S removed;
    class A fixed;
    class Y normal;
Loading

This makes clear why conditioning is not intervention. Conditioning on $A=a$ selects units whose natural assignment mechanism produced $a$. Intervening changes the mechanism that produces $A$.

3.5 Counterfactual inference: abduction, action, prediction

An SCM supports unit-specific counterfactual reasoning in three conceptual steps.

1. Abduction

Use observed evidence to update beliefs about exogenous variables or latent state.

2. Action

Modify the structural model with the hypothetical intervention.

3. Prediction

Propagate the updated latent state through the modified model.

For a failed deployment, evidence such as logs and timing informs beliefs about latent defect severity or dependency instability. The counterfactual “Would delaying the deployment have prevented the incident?” keeps the inferred latent circumstances fixed while changing the deployment-time mechanism.

This is stronger than estimating an average effect across deployments.

3.6 Mapping potential outcomes to SCMs

In an SCM, the potential outcome $Y(a)$ is the value of $Y$ in the modified model where the treatment assignment is replaced by $A:=a$.

Potential outcomes are excellent for defining estimands and study designs. SCMs are especially useful for:

  • representing multiple variables and pathways;
  • determining adjustment sets;
  • expressing interventions on mechanisms;
  • deriving counterfactuals;
  • causal discovery;
  • world models and agent interaction.

The frameworks are complementary rather than competing.

3.7 Treatment versions and intervention semantics

An intervention must specify more than a variable name when multiple mechanisms can produce the same value.

Consider service_available = true. This value could be produced by:

  • restarting the existing process;
  • replacing the process with a healthy replica;
  • routing traffic to another region;
  • returning a cached response while the service remains down.

These interventions may have different downstream consequences even though the measured variable has the same value. Classical perfect interventions often abstract away this difference. For equilibrium systems and mechanism-specific actions, richer formalisms may be needed. The 2026 proposal of bipartite graphical causal models makes equation replacement explicit: the intervention identifies which mechanism is replaced, which variable is constrained, and to what value R39.

3.8 Population, unit, and mechanism uncertainty

Do not collapse all uncertainty into a single confidence interval.

  • Sampling uncertainty: We observed a finite sample.
  • Model uncertainty: Several functional forms or graphs fit the data.
  • Identification uncertainty: The data and assumptions may only bound the answer.
  • Unit-level uncertainty: We may know an average effect but not a specific counterfactual.
  • Mechanism uncertainty: The intervention may not preserve mechanisms as assumed.
  • Distribution uncertainty: The target environment may differ from the study environment.

A reliable causal AI system should expose which kind of uncertainty dominates.

3.9 Exercises

  1. Define ATE, ATT, and CATE for a deployment-testing intervention.
  2. Give an example where positivity fails in a company’s historical data.
  3. Explain why “send a notification” may violate treatment consistency.
  4. Write a three-equation SCM for traffic, latency, and autoscaling.
  5. Describe the abduction, action, and prediction steps for the counterfactual: “Would the incident have occurred if the deployment had been delayed by one hour?”

Part II — Identification and Estimation

4. Identification: Can the Question Be Answered?

Identification comes before estimation.

An estimand is identified when, under stated assumptions, it can be expressed as a function of the observed data distribution. If two causal models satisfy all assumptions, generate the same observed distribution, but imply different answers to the target query, the query is not point identified.

No amount of data, model capacity, or optimization can resolve that ambiguity without new assumptions or new kinds of data.

4.1 The four-step discipline

For every causal analysis, separate:

  1. Question: What is the intervention and target population?
  2. Model: What causal assumptions are represented?
  3. Identification: Which observed-data expression equals the causal quantity?
  4. Estimation: How will that expression be estimated from a finite sample?

Many failed analyses jump from a vague question directly to an estimator.

4.2 Backdoor adjustment

Suppose $A$ is treatment, $Y$ is outcome, and $Z$ is a common cause:

flowchart LR
    Z["Confounder Z"] --> A["Treatment A"]
    Z --> Y["Outcome Y"]
    A --> Y

    classDef conf fill:#FFF4D6,stroke:#D97706,color:#3B2F0B,stroke-width:2px;
    classDef treat fill:#E8F1FF,stroke:#2563EB,color:#102A43,stroke-width:2px;
    classDef out fill:#E8F8EF,stroke:#15803D,color:#123524,stroke-width:2px;
    class Z conf;
    class A treat;
    class Y out;
Loading

The path $A\leftarrow Z\rightarrow Y$ enters $A$ through an incoming arrow. It is a backdoor path. Conditioning on $Z$ blocks it.

The adjustment formula is:

$$ P(Y\mid do(A=a)) = \sum_z P(Y\mid A=a,Z=z)P(Z=z). $$

For a continuous $Z$, replace the sum with an integral.

A set $Z$ is a valid backdoor adjustment set when it blocks every backdoor path from $A$ to $Y$ and contains no descendant of $A$.

Why descendants are dangerous

A descendant of treatment can be:

  • a mediator, in which case controlling for it removes part of the total effect;
  • a collider or descendant of a collider, in which case controlling for it opens bias;
  • a post-treatment consequence sharing causes with the outcome, producing post-treatment confounding.

Pre-treatment status alone does not guarantee a variable is a good control, but post-treatment variables require particular care.

4.3 Good controls, unnecessary controls, and bad controls

Consider:

flowchart LR
    C["Change complexity"] --> T["Test depth"]
    C --> I["Incident"]
    T --> B["Bugs caught"]
    B --> I
    U["Engineer experience"] --> T

    classDef conf fill:#FFF4D6,stroke:#D97706,color:#3B2F0B,stroke-width:2px;
    classDef treat fill:#E8F1FF,stroke:#2563EB,color:#102A43,stroke-width:2px;
    classDef med fill:#F2EAFE,stroke:#7C3AED,color:#2E1065,stroke-width:2px;
    classDef out fill:#E8F8EF,stroke:#15803D,color:#123524,stroke-width:2px;
    class C conf;
    class T treat;
    class B med;
    class I out;
    class U treat;
Loading

To estimate the total effect of test_depth on incident:

  • change_complexity is a confounder and should be adjusted for.
  • bugs_caught is a mediator and should not be adjusted for if the target is the total effect.
  • engineer_experience predicts treatment but, in this graph, does not affect the outcome except through treatment. Adjusting for it is not required for identification and can sometimes increase variance.

A variable’s predictive usefulness is not the criterion. Its position in the causal structure is.

4.4 The frontdoor strategy

Frontdoor adjustment can identify an effect despite unobserved treatment–outcome confounding when a suitable mediator is observed.

flowchart LR
    U["Hidden U"] --> A["Treatment A"]
    U --> Y["Outcome Y"]
    A --> M["Mediator M"]
    M --> Y

    classDef hidden fill:#F2EAFE,stroke:#7C3AED,color:#2E1065,stroke-width:2px,stroke-dasharray:5 5;
    classDef treat fill:#E8F1FF,stroke:#2563EB,color:#102A43,stroke-width:2px;
    classDef med fill:#FFF4D6,stroke:#D97706,color:#3B2F0B,stroke-width:2px;
    classDef out fill:#E8F8EF,stroke:#15803D,color:#123524,stroke-width:2px;
    class U hidden;
    class A treat;
    class M med;
    class Y out;
Loading

The standard conditions are strong:

  1. $M$ intercepts every directed path from $A$ to $Y$.
  2. There is no unblocked backdoor path from $A$ to $M$.
  3. All backdoor paths from $M$ to $Y$ are blocked by $A$.

When these hold, the effect can be written using observational distributions. Frontdoor examples in the real world are rarer than textbook presentations suggest because mediator–outcome confounding and direct treatment effects are common.

4.5 Instrumental variables

An instrument $Z$ creates variation in treatment $A$ that can isolate causal effects even when $A$ and $Y$ are confounded.

flowchart LR
    Z["Instrument Z"] --> A["Treatment A"] --> Y["Outcome Y"]
    U["Hidden confounder U"] --> A
    U --> Y

    classDef inst fill:#FFF4D6,stroke:#D97706,color:#3B2F0B,stroke-width:2px;
    classDef treat fill:#E8F1FF,stroke:#2563EB,color:#102A43,stroke-width:2px;
    classDef hidden fill:#F2EAFE,stroke:#7C3AED,color:#2E1065,stroke-width:2px,stroke-dasharray:5 5;
    classDef out fill:#E8F8EF,stroke:#15803D,color:#123524,stroke-width:2px;
    class Z inst;
    class A treat;
    class U hidden;
    class Y out;
Loading

Typical assumptions include:

  • Relevance: $Z$ changes $A$.
  • Exclusion: $Z$ affects $Y$ only through $A$.
  • Independence: $Z$ is independent of unobserved causes of $Y$.
  • Monotonicity: for a common local-average-treatment-effect interpretation, the instrument does not make some units systematically move in the opposite treatment direction.

In the simplest linear case, the Wald estimand is:

$$ \frac{E[Y\mid Z=1]-E[Y\mid Z=0]} {E[A\mid Z=1]-E[A\mid Z=0]}. $$

An instrument does not automatically identify the population ATE. It often identifies an effect for “compliers,” the units whose treatment changes because of the instrument.

Weak instruments create severe finite-sample problems. Exclusion restrictions are substantive and frequently debatable.

4.6 Mediation

A mediator lies on a causal path from treatment to outcome.

flowchart LR
    A["Treatment"] --> M["Mediator"] --> Y["Outcome"]
    A --> Y

    classDef treat fill:#E8F1FF,stroke:#2563EB,color:#102A43,stroke-width:2px;
    classDef med fill:#FFF4D6,stroke:#D97706,color:#3B2F0B,stroke-width:2px;
    classDef out fill:#E8F8EF,stroke:#15803D,color:#123524,stroke-width:2px;
    class A treat;
    class M med;
    class Y out;
Loading

Different questions include:

  • Total effect: all pathways from $A$ to $Y$.
  • Controlled direct effect: effect of $A$ while fixing $M$ to a chosen value.
  • Natural direct and indirect effects: cross-world quantities that compare mediator values generated under different treatment conditions.

Natural effects require stronger assumptions and careful interpretation. In an engineering system, a controlled intervention on the mediator may be more actionable than a natural-effect decomposition.

4.7 Longitudinal treatment and time-varying confounding

Suppose treatment at time $t$ affects a later covariate $L_{t+1}$, which both influences future treatment and predicts the outcome:

flowchart LR
    L0 --> A0 --> L1 --> A1 --> Y
    L0 --> Y
    A0 --> Y
    L1 --> Y

    classDef state fill:#FFF4D6,stroke:#D97706,color:#3B2F0B,stroke-width:2px;
    classDef action fill:#E8F1FF,stroke:#2563EB,color:#102A43,stroke-width:2px;
    classDef out fill:#E8F8EF,stroke:#15803D,color:#123524,stroke-width:2px;
    class L0,L1 state;
    class A0,A1 action;
    class Y out;
Loading

Naively adjusting for $L_1$ can block part of the effect of $A_0$, while failing to adjust can confound the effect of $A_1$. This motivates g-methods such as:

  • the parametric g-formula;
  • marginal structural models with inverse-probability weights;
  • structural nested models;
  • longitudinal doubly robust estimators.

Sequential decision-making and agents routinely create time-varying confounding because earlier actions change later state and action selection.

4.8 Difference-in-differences and regression discontinuity

Not every identification strategy is expressed primarily through a small DAG.

Difference-in-differences

Compares changes over time between treated and comparison groups. Its central identifying condition is a form of parallel trends: absent treatment, the groups’ outcomes would have evolved similarly. Modern variants address staggered adoption and heterogeneous effects; simple two-way fixed-effects regressions can be misleading in those settings.

Regression discontinuity

Uses a treatment assignment rule that changes discontinuously at a threshold. Under continuity and no precise manipulation around the cutoff, units just above and below can identify a local causal effect.

These are designs, not mere regression specifications. Their credibility comes from assignment structure and diagnostic evidence.

4.9 Partial identification

Sometimes the correct answer is a set or interval rather than a point.

If hidden confounding cannot be ruled out, observed data may imply bounds:

$$ \tau \in [\tau_L,\tau_U]. $$

Bounds can be tightened by:

  • monotonicity assumptions;
  • instrumental variables;
  • mediator or covariate information;
  • shape constraints;
  • sensitivity parameters;
  • limited interventional data.

Partial identification is not failure. It prevents assumptions from being hidden inside a point estimate. A 2026 foundation-model proposal extends prior-data fitted networks to learn distributions over partially identified intervention and counterfactual queries, explicitly representing multiple values compatible with the observations and assumptions R38.

4.10 An identification decision tree

flowchart TD
    Q["Define treatment, outcome, population, horizon"] --> R{"Randomized treatment?"}
    R -->|Yes| E["Estimate intention-to-treat or protocol-specific effect"]
    R -->|No| G{"Credible causal graph or design?"}
    G -->|No| B["Collect domain knowledge, redesign study, or report association only"]
    G -->|Yes| C{"Observed set blocks confounding?"}
    C -->|Yes| A["Backdoor / g-formula / doubly robust estimation"]
    C -->|No| IV{"Valid instrument, frontdoor mediator, threshold, panel design, or intervention data?"}
    IV -->|Yes| S["Use design-specific identification"]
    IV -->|No| P["Partial identification or sensitivity analysis"]

    classDef start fill:#FFF4D6,stroke:#D97706,color:#3B2F0B,stroke-width:2px;
    classDef yes fill:#E8F8EF,stroke:#15803D,color:#123524,stroke-width:2px;
    classDef no fill:#FCE8F3,stroke:#BE185D,color:#4A102A,stroke-width:2px;
    classDef decision fill:#E8F1FF,stroke:#2563EB,color:#102A43,stroke-width:2px;
    class Q start;
    class R,G,C,IV decision;
    class E,A,S yes;
    class B,P no;
Loading

4.11 Exercises

  1. For the support-and-renewal example, state a backdoor adjustment set and explain its substantive meaning.
  2. Invent a plausible instrument for an engineering process. Then argue why its exclusion restriction might fail.
  3. Give a treatment whose effect is only partially identified from your current data. What additional intervention would tighten the bounds?
  4. Explain why identifying the total effect and identifying the direct effect require different control sets.
  5. Write a longitudinal graph for an AI agent whose past tool calls affect both future observations and future tool selection.

5. Estimation: From Identified Formula to Data

Once a target is identified, estimation turns an observed-data functional into a number, interval, or distribution.

This chapter focuses on the common setting with a binary treatment $A$, outcome $Y$, and pre-treatment covariates $X$, under conditional exchangeability and overlap.

Define:

$$ \mu_a(x)=E[Y\mid A=a,X=x] $$

and:

$$ e(x)=P(A=1\mid X=x). $$

The first is an outcome model. The second is a propensity model.

5.1 Randomized experiments

In a simple randomized experiment, the difference in sample means is an unbiased estimator of the intention-to-treat effect:

$$ \widehat{ATE}=\bar{Y}_{A=1}-\bar{Y}_{A=0}. $$

Randomization handles confounding in expectation. It does not automatically solve:

  • noncompliance;
  • attrition;
  • interference;
  • treatment variation;
  • low power;
  • multiple testing;
  • external validity;
  • logging or implementation bugs.

Covariate adjustment can improve precision when performed appropriately. Pre-registration and clear outcome definitions protect against analysis flexibility.

5.2 Outcome regression and standardization

Fit a model for $E[Y\mid A,X]$. Predict both potential treatment conditions for every unit:

$$ \widehat{Y_i(1)}=\hat\mu_1(X_i), \qquad \widehat{Y_i(0)}=\hat\mu_0(X_i). $$

Then average:

$$ \widehat{ATE}_{OR}

\frac{1}{n}\sum_{i=1}^n \left[\hat\mu_1(X_i)-\hat\mu_0(X_i)\right]. $$

This is the g-formula or standardization estimator in this simple setting.

It is consistent if the outcome model is correctly specified and the causal assumptions hold. Flexible ML can reduce functional-form error, but naive overfitting can damage statistical inference.

5.3 Propensity scores and inverse-probability weighting

The propensity score compresses observed treatment-assignment information into a scalar:

$$ e(X)=P(A=1\mid X). $$

An inverse-probability weighted estimator is:

$$ \widehat{ATE}_{IPW}

\frac{1}{n}\sum_{i=1}^n \left[ \frac{A_iY_i}{\hat e(X_i)}

\frac{(1-A_i)Y_i}{1-\hat e(X_i)} \right]. $$

Weighting creates a pseudo-population in which treatment is balanced with respect to measured covariates.

Common problems include:

  • extreme weights from poor overlap;
  • misspecified propensity models;
  • high variance;
  • hidden confounding;
  • treating propensity prediction accuracy as the main goal.

Propensity models should be evaluated for balance and overlap, not merely classification metrics.

5.4 Matching

Matching pairs or groups treated and untreated units with similar observed covariates or propensity scores.

Matching is intuitive, but it does not create randomization by magic. Key decisions include:

  • distance metric;
  • caliper;
  • matching with or without replacement;
  • exact matching on critical variables;
  • handling unmatched units;
  • defining the resulting target population.

The estimate may target the treated, the matched sample, or another overlap population rather than the original population ATE.

5.5 Doubly robust estimation

The augmented inverse-probability weighted estimator combines outcome and propensity models:

$$ \widehat{ATE}_{AIPW}

\frac{1}{n}\sum_{i=1}^{n} \left[ \hat\mu_1(X_i)-\hat\mu_0(X_i) + \frac{A_i(Y_i-\hat\mu_1(X_i))}{\hat e(X_i)}

\frac{(1-A_i)(Y_i-\hat\mu_0(X_i))}{1-\hat e(X_i)} \right]. $$

Under regularity conditions, it is consistent if either the outcome model or propensity model is correctly specified. “Doubly robust” does not mean robust to hidden confounding or positivity failure.

5.6 Orthogonalization and double machine learning

Modern ML estimators can fit nuisance functions such as $E[Y\mid X]$ and $E[A\mid X]$ well but introduce regularization and overfitting bias. Double/debiased machine learning uses:

  • orthogonal estimating equations that are locally insensitive to nuisance error;
  • sample splitting or cross-fitting so predictions are evaluated on observations not used to train the nuisance model.

For a partially linear model:

$$ Y = \theta A + g(X)+\varepsilon, $$

one can residualize:

$$ \tilde Y=Y-\hat E[Y\mid X], \qquad \tilde A=A-\hat E[A\mid X], $$

then estimate $\theta$ from the relationship between $\tilde Y$ and $\tilde A$ on held-out folds.

The conceptual benefit is separation of predictive nuisance learning from the causal target. The predictive model is a component, not the estimand.

5.7 Complete Python lab: confounding, adjustment, IPW, and AIPW

The script below generates synthetic observational data with a known ATE. It compares a naive difference in means with regression adjustment, IPW, and cross-fitted AIPW.

"""A complete synthetic causal-effect estimation example.

Dependencies:
    pip install numpy pandas scikit-learn statsmodels
"""

from __future__ import annotations

from dataclasses import dataclass
from typing import Tuple

import numpy as np
import pandas as pd
import statsmodels.api as sm
from sklearn.base import clone
from sklearn.ensemble import HistGradientBoostingClassifier, HistGradientBoostingRegressor
from sklearn.model_selection import KFold


@dataclass(frozen=True)
class SimulationConfig:
    n: int = 12_000
    true_ate: float = 2.0
    seed: int = 7


def sigmoid(x: np.ndarray) -> np.ndarray:
    return 1.0 / (1.0 + np.exp(-x))


def simulate_data(config: SimulationConfig) -> pd.DataFrame:
    """Simulate confounded observational data with a constant treatment effect."""
    rng = np.random.default_rng(config.seed)

    age = rng.normal(0.0, 1.0, config.n)
    severity = rng.normal(0.0, 1.0, config.n)
    prior_usage = 0.6 * age - 0.4 * severity + rng.normal(0.0, 0.8, config.n)

    propensity = sigmoid(
        -0.3
        + 1.1 * severity
        + 0.5 * age
        - 0.4 * prior_usage
        + 0.3 * severity * age
    )
    treatment = rng.binomial(1, propensity)

    baseline = (
        4.0
        + 1.8 * severity
        - 0.9 * age
        + 0.7 * prior_usage
        + 0.5 * severity**2
    )
    outcome = baseline + config.true_ate * treatment + rng.normal(0.0, 1.0, config.n)

    return pd.DataFrame(
        {
            "age": age,
            "severity": severity,
            "prior_usage": prior_usage,
            "treatment": treatment,
            "outcome": outcome,
            "true_propensity": propensity,
        }
    )


def naive_difference(df: pd.DataFrame) -> float:
    treated = df.loc[df["treatment"] == 1, "outcome"].mean()
    untreated = df.loc[df["treatment"] == 0, "outcome"].mean()
    return float(treated - untreated)


def regression_adjustment(df: pd.DataFrame) -> float:
    """Linear g-computation with explicit nonlinear and interaction features."""
    design = df[["treatment", "age", "severity", "prior_usage"]].copy()
    design["severity_sq"] = design["severity"] ** 2
    design["severity_age"] = design["severity"] * design["age"]
    design = sm.add_constant(design, has_constant="add")

    model = sm.OLS(df["outcome"], design).fit()

    treated_design = design.copy()
    treated_design["treatment"] = 1
    untreated_design = design.copy()
    untreated_design["treatment"] = 0

    y1 = model.predict(treated_design)
    y0 = model.predict(untreated_design)
    return float(np.mean(y1 - y0))


def cross_fitted_nuisance_predictions(
    x: np.ndarray,
    treatment: np.ndarray,
    outcome: np.ndarray,
    n_splits: int = 5,
    seed: int = 11,
) -> Tuple[np.ndarray, np.ndarray, np.ndarray]:
    """Return out-of-fold propensity, treated-outcome, and control-outcome predictions."""
    propensity_model = HistGradientBoostingClassifier(
        max_depth=4,
        learning_rate=0.05,
        max_iter=250,
        random_state=seed,
    )
    outcome_model = HistGradientBoostingRegressor(
        max_depth=4,
        learning_rate=0.05,
        max_iter=250,
        random_state=seed,
    )

    e_hat = np.empty(len(outcome), dtype=float)
    mu1_hat = np.empty(len(outcome), dtype=float)
    mu0_hat = np.empty(len(outcome), dtype=float)

    splitter = KFold(n_splits=n_splits, shuffle=True, random_state=seed)
    for train_idx, test_idx in splitter.split(x):
        x_train, x_test = x[train_idx], x[test_idx]
        a_train, a_test = treatment[train_idx], treatment[test_idx]
        y_train = outcome[train_idx]

        e_model = clone(propensity_model)
        e_model.fit(x_train, a_train)
        e_hat[test_idx] = e_model.predict_proba(x_test)[:, 1]

        treated_model = clone(outcome_model)
        control_model = clone(outcome_model)

        treated_mask = a_train == 1
        control_mask = a_train == 0
        if treated_mask.sum() == 0 or control_mask.sum() == 0:
            raise RuntimeError("A fold contains no treated or no control observations.")

        treated_model.fit(x_train[treated_mask], y_train[treated_mask])
        control_model.fit(x_train[control_mask], y_train[control_mask])

        mu1_hat[test_idx] = treated_model.predict(x_test)
        mu0_hat[test_idx] = control_model.predict(x_test)

    return e_hat, mu1_hat, mu0_hat


def ipw_and_aipw(df: pd.DataFrame) -> Tuple[float, float]:
    x = df[["age", "severity", "prior_usage"]].to_numpy()
    a = df["treatment"].to_numpy(dtype=float)
    y = df["outcome"].to_numpy(dtype=float)

    e_hat, mu1_hat, mu0_hat = cross_fitted_nuisance_predictions(x, a, y)

    # Trimming is a pragmatic variance-control decision, not a cure for positivity failure.
    e_hat = np.clip(e_hat, 0.02, 0.98)

    ipw_score = a * y / e_hat - (1.0 - a) * y / (1.0 - e_hat)
    ipw = float(np.mean(ipw_score))

    aipw_score = (
        mu1_hat
        - mu0_hat
        + a * (y - mu1_hat) / e_hat
        - (1.0 - a) * (y - mu0_hat) / (1.0 - e_hat)
    )
    aipw = float(np.mean(aipw_score))
    return ipw, aipw


def main() -> None:
    config = SimulationConfig()
    df = simulate_data(config)

    naive = naive_difference(df)
    adjusted = regression_adjustment(df)
    ipw, aipw = ipw_and_aipw(df)

    print(f"True ATE:                {config.true_ate: .3f}")
    print(f"Naive difference:        {naive: .3f}")
    print(f"Regression adjustment:   {adjusted: .3f}")
    print(f"Cross-fitted IPW:         {ipw: .3f}")
    print(f"Cross-fitted AIPW:        {aipw: .3f}")


if __name__ == "__main__":
    main()

The exact numbers vary by random seed. The important pattern is that the naive estimate is biased because treatment assignment depends on severity and other outcome causes. Adjustment methods move toward the known effect when their models and overlap are adequate.

5.8 Diagnostics before celebrating an estimate

Overlap

Plot or summarize propensity distributions by treatment group. Ask whether comparable treated and untreated units exist.

Covariate balance

After weighting or matching, inspect standardized mean differences and distributional balance. Balance on measured covariates is a diagnostic for the adjustment procedure, not proof that hidden confounding is absent.

Effective sample size

Large or uneven weights can reduce the effective sample size far below the row count.

For normalized weights $w_i$:

$$ ESS=\frac{(\sum_i w_i)^2}{\sum_i w_i^2}. $$

Negative controls

A negative-control outcome should not plausibly be affected by treatment. A negative-control exposure should not plausibly affect the outcome. Detected relationships can reveal residual confounding or data leakage.

Placebo and refutation tests

Examples:

  • replace treatment with a random variable;
  • use a pre-treatment outcome as though it were post-treatment;
  • add a simulated unobserved confounder;
  • vary adjustment sets justified by alternative graphs;
  • test whether effects appear before treatment.

These tests can reveal problems. Passing them does not prove the causal assumptions.

5.9 Sensitivity analysis

Sensitivity analysis asks how strong an unobserved confounder or assumption violation would need to be to change the conclusion.

Useful outputs include:

  • effect estimates over a grid of confounding strengths;
  • bounds under treatment–outcome confounding parameters;
  • E-values in some epidemiological settings;
  • robustness values in omitted-variable analyses;
  • tipping-point plots;
  • estimates across plausible graph variants.

A causal AI interface should make sensitivity results visible rather than hiding them behind a single score.

5.10 Uncertainty and repeated analysis

Standard errors and confidence intervals need to match the estimator and data structure. Watch for:

  • clustered observations;
  • repeated measurements;
  • interference within groups;
  • adaptive experiments;
  • multiple outcomes and subgroup searches;
  • model selection on the same data;
  • temporal autocorrelation.

A narrow confidence interval can coexist with severe identification bias. Statistical precision is conditional on the model and identifying assumptions.

5.11 Software ecosystem

Common Python tools include:

  • DoWhy: explicit model–identify–estimate–refute workflow R22.
  • EconML: heterogeneous treatment effects, double ML, orthogonal forests, and related estimators R23.
  • causal-learn: constraint-, score-, and functional-model causal discovery methods R24.
  • DoubleML: double/debiased ML in Python and R.
  • Ananke and related graph libraries: identification and graphical causal models.

Libraries help operationalize assumptions. They do not determine whether those assumptions are credible.

5.12 Exercises

  1. Why can a propensity model with excellent classification accuracy be bad for causal estimation?
  2. Explain what “doubly robust” does and does not mean.
  3. Modify the Python lab so treatment effects vary with severity. Estimate the ATE and plot the true CATE.
  4. Give an example where trimming extreme propensities changes the target population.
  5. Design a negative-control test for a deployment-analysis dataset.

6. Heterogeneous Effects and Decisions

Average effects are often insufficient for action. A treatment may help some units, harm others, and have an average near zero.

6.1 CATE and individual treatment effects

The conditional average treatment effect is:

$$ \tau(x)=E[Y(1)-Y(0)\mid X=x]. $$

It describes an average among units with covariates $x$. It is not generally the deterministic individual treatment effect for a specific person or system instance.

Individual counterfactual effects require stronger structural assumptions because the joint distribution of $Y(1)$ and $Y(0)$ is not identified from separate marginal distributions in ordinary experiments.

Use language carefully:

  • “Estimated CATE for this covariate profile” is defensible.
  • “The treatment will improve this exact unit by 3.2” often overstates what was learned.

6.2 Meta-learners

Meta-learners turn supervised learners into CATE estimators.

S-learner

Fit one outcome model using treatment as a feature:

$$ \hat\mu(a,x). $$

Then:

$$ \hat\tau(x)=\hat\mu(1,x)-\hat\mu(0,x). $$

Simple, but a flexible learner may underuse the treatment feature when effects are weak.

T-learner

Fit separate outcome models in treated and control groups:

$$ \hat\tau(x)=\hat\mu_1(x)-\hat\mu_0(x). $$

Flexible but can be unstable when one treatment group is small.

X-learner

Imputes treatment effects within each group and combines them, often useful under treatment imbalance.

R-learner

Residualizes outcome and treatment and estimates effect heterogeneity through an orthogonal objective. It connects naturally to double machine learning.

6.3 Causal forests

Causal forests adapt random forests to estimate heterogeneous treatment effects. Splitting criteria seek treatment-effect differences rather than merely outcome prediction. Honest sample splitting separates tree construction from effect estimation to reduce bias.

Forests are useful when:

  • heterogeneity is nonlinear;
  • interactions are unknown;
  • sample size is adequate;
  • overlap is reasonably strong.

They do not repair confounding or poor treatment definition.

6.4 Uplift modeling versus causal modeling

Uplift models rank units by the predicted difference between treated and untreated outcomes. In randomized marketing experiments, uplift can be causally interpretable. In observational logs, it can inherit confounding from treatment assignment.

A ranking metric can also hide calibration error. A policy needs effect magnitude, uncertainty, cost, and capacity constraints, not only rank order.

6.5 From effects to policies

Let $\pi(x)\in{0,1}$ be a treatment policy and $c(x)$ the cost of treatment. A simple value objective is:

$$ V(\pi)=E[Y(\pi(X))-c(X)\pi(X)]. $$

If treatment is beneficial when higher $Y$ is better, a plug-in rule might treat when:

$$ \hat\tau(x)>c(x). $$

But production decisions may also need:

  • uncertainty penalties;
  • risk constraints;
  • fairness constraints;
  • budgets;
  • treatment capacity;
  • delayed and multiple outcomes;
  • exploration to maintain learning;
  • abstention when overlap is weak.

6.6 Decision-focused uncertainty

Suppose the estimated effect is $0.5$ with a wide interval $[-1.2,2.1]$. Whether to act depends on asymmetric losses.

A low-cost reversible UI experiment may be reasonable. A high-risk medical or infrastructure intervention may require stronger evidence.

One useful rule is to separate:

  • epistemic uncertainty: reducible with more evidence;
  • aleatoric uncertainty: inherent outcome variability;
  • model ambiguity: several causal models remain plausible;
  • decision risk: consequence of choosing wrongly.

The action threshold should reflect decision risk, not only statistical significance.

6.7 Policy learning creates feedback

Once a policy is deployed, treatment assignment changes. This has three implications:

  1. Historical overlap can disappear.
  2. The population receiving treatment changes, so average effects can shift.
  3. Future data are generated by the policy and cannot be treated as passively sampled.

Maintain exploration where ethically and operationally acceptable, log propensities, version policies, and monitor covariate and mechanism shifts.

6.8 Interference and network effects

When one unit’s treatment affects others, define exposure more broadly. Examples:

  • notifications create competition for attention;
  • marketplace incentives change prices for untreated participants;
  • restarting one service shifts load to another;
  • a treatment affects infection risk in a social network;
  • one agent’s action changes another agent’s observations.

Possible strategies include cluster randomization, graph-based exposure models, partial-interference assumptions, and explicit multi-agent or system-level SCMs.

6.9 A deployment decision example

Assume extended testing reduces incident risk more for large, dependency-heavy changes but costs engineering time.

A decision system should not simply predict incident risk. It should estimate:

$$ \tau(x)= P(incident=1\mid do(test=1),X=x)

P(incident=1\mid do(test=0),X=x). $$

Because lower incident probability is desirable, testing is valuable when the expected reduction in incident loss exceeds test cost and delay cost.

A robust policy might:

  • require extended testing when the lower confidence bound on avoided loss exceeds cost;
  • randomize a small fraction of uncertain, low-risk cases to preserve learning;
  • abstain and request review when the change lies outside historical support;
  • separately model spillovers from delayed deployments into traffic peaks or release queues.

6.10 Exercises

  1. Explain why CATE is not necessarily an individual causal effect.
  2. Compare the failure modes of S-, T-, and R-learners under highly imbalanced treatment.
  3. Write a policy-value objective for deciding whether to roll back a deployment.
  4. Describe how deploying a personalized treatment policy can destroy future overlap.
  5. Give an example where the highest estimated treatment effect should not receive treatment because of cost or risk.

Part III — Learning Structure Through Data and Action

7. Causal Discovery

Causal discovery asks what aspects of causal structure can be learned from data. It is attractive because manually specifying a graph is difficult. It is also one of the easiest areas to overclaim.

7.1 Why observational discovery is underdetermined

Consider three variables with the same skeleton:

$$ X-Y-Z. $$

The graphs:

$$ X\rightarrow Y\rightarrow Z, $$

$$ X\leftarrow Y\rightarrow Z, $$

and:

$$ X\leftarrow Y\leftarrow Z $$

all imply $X\perp Z\mid Y$ under ordinary Markov and faithfulness assumptions. Observational conditional independence alone cannot distinguish them.

They belong to the same Markov equivalence class. A class is characterized by the same skeleton and the same unshielded colliders. Constraint- and score-based algorithms often recover a CPDAG, which represents directed edges shared by every DAG in the class and undirected edges whose direction remains ambiguous.

With latent confounding, the corresponding output may be a PAG, representing a broader equivalence class with endpoint marks that encode uncertainty about ancestry and confounding.

A single fully directed graph from observational data usually reflects additional assumptions, priors, or arbitrary tie-breaking.

7.2 Families of discovery methods

flowchart TD
    CD["Causal discovery"] --> CB["Constraint based"]
    CD --> SB["Score based"]
    CD --> FCM["Functional causal models"]
    CD --> INV["Invariance / multiple environments"]
    CD --> INT["Interventional discovery"]
    CD --> TEMP["Temporal and dynamical"]
    CD --> AM["Amortized / foundation models"]

    CB --> PC["PC, FCI"]
    SB --> GES["GES, Bayesian scores"]
    SB --> CONT["NOTEARS-style optimization"]
    FCM --> LINGAM["LiNGAM"]
    FCM --> ANM["Additive-noise models"]
    INV --> ICP["Invariant prediction"]
    AM --> AVICI["Pretrained graph predictors"]

    classDef root fill:#FFF4D6,stroke:#D97706,color:#3B2F0B,stroke-width:2px;
    classDef family fill:#E8F1FF,stroke:#2563EB,color:#102A43,stroke-width:2px;
    classDef method fill:#E8F8EF,stroke:#15803D,color:#123524,stroke-width:2px;
    class CD root;
    class CB,SB,FCM,INV,INT,TEMP,AM family;
    class PC,GES,CONT,LINGAM,ANM,ICP,AVICI method;
Loading

7.3 Constraint-based methods

Constraint-based methods perform conditional independence tests and use graphical rules to remove and orient edges.

PC

Under causal sufficiency, acyclicity, Markov, faithfulness, and correct independence tests, PC recovers the correct CPDAG asymptotically.

A simplified view:

  1. Begin with a complete undirected graph.
  2. Remove $X-Y$ when a conditioning set $S$ makes $X\perp Y\mid S$.
  3. Record separating sets.
  4. Orient unshielded colliders.
  5. Propagate orientations without introducing new colliders or cycles.

PC is sensitive to:

  • conditional-independence test quality;
  • significance thresholds;
  • variable ordering in some implementations;
  • sample size relative to graph degree;
  • weak effects and near-unfaithfulness;
  • measurement error.

FCI

FCI allows latent confounding and selection bias under its assumptions. It returns a PAG rather than pretending to know every edge direction.

The price is more ambiguity and computational complexity. That ambiguity is epistemically meaningful.

7.4 Score-based methods

Score-based methods search for graphs that optimize a fit–complexity objective such as BIC or a Bayesian marginal likelihood.

GES

Greedy Equivalence Search moves through equivalence classes, adding and then deleting edges to improve a score. Under suitable assumptions and large samples, score-equivalent and consistent scores can recover the correct equivalence class.

Bayesian structure learning

A posterior over graphs can represent model uncertainty:

$$ P(G\mid D)\propto P(D\mid G)P(G). $$

Exact inference is generally intractable for all but small graphs, so methods use MCMC, variational approximations, dynamic programming under restrictions, or continuous relaxations.

7.5 Continuous optimization and NOTEARS

NOTEARS introduced a smooth, exact acyclicity characterization for a weighted adjacency matrix $W$:

$$ h(W)=\operatorname{tr}(e^{W\circ W})-d=0. $$

This turns DAG learning into a constrained continuous optimization problem R11. It made gradient-based causal discovery architectures much easier to build.

Important caveats:

  • A differentiable acyclicity constraint solves a search representation problem, not causal identifiability.
  • The result depends on the score, functional model, regularization, and optimization landscape.
  • Thresholding weighted edges can materially alter the graph.
  • Later work has identified optimization subtleties in this family.

Do not equate “neural DAG learner” with “assumption-free causal discovery.”

7.6 Functional causal models

Functional-model methods exploit asymmetries beyond conditional independence.

LiNGAM

LiNGAM assumes linear relations, non-Gaussian independent noise, acyclicity, and no hidden confounding in its basic form. Under those conditions, independent-component structure can identify the full causal ordering from observational data R12.

Additive-noise models

For two variables, suppose:

$$ Y=f(X)+N, \qquad N\perp X. $$

In many nonlinear settings, the reverse representation:

$$ X=g(Y)+\tilde N, \qquad \tilde N\perp Y $$

does not exist. This asymmetry can identify direction R13.

These methods are powerful when their functional assumptions are plausible. They can be confidently wrong when those assumptions fail.

7.7 Invariance across environments

Causal mechanisms are often expected to remain stable while distributions of causes change across environments.

Invariant Causal Prediction seeks predictor sets $S$ such that:

$$ P(Y\mid X_S, E=e) $$

is invariant across environments $e$, under appropriate assumptions. Changes across environments supply information unavailable in a single i.i.d. dataset R14.

Environments can arise from:

  • known interventions;
  • different organizations or sites;
  • policy regimes;
  • time periods;
  • natural perturbations;
  • experimental batches.

The environment variable must be interpreted carefully. If the mechanism for $Y$ itself changes, invariance can fail even for true parents.

7.8 Temporal causal discovery

Time provides ordering information but not automatic causal identification. Common families include:

  • Granger-style predictability methods;
  • vector autoregressive structural models;
  • PCMCI and conditional-independence methods for time series;
  • temporal LiNGAM variants;
  • neural temporal graph learners;
  • continuous-time causal models;
  • methods for nonstationary or regime-switching mechanisms.

Questions to settle before applying them:

  • What is the sampling interval relative to mechanism speed?
  • Are contemporaneous effects possible?
  • Are relevant lags measured?
  • Is the process stationary?
  • Are interventions or policy changes logged?
  • Does aggregation create apparent reverse causality?
  • Are there hidden common drivers?

“Cause precedes effect” is necessary in many systems but not sufficient.

7.9 Interventional discovery

Interventions break observational equivalence. Setting $X$ externally removes or changes its natural incoming mechanism, allowing edges around $X$ to be oriented.

Interventional data can vary in quality:

  • known perfect interventions;
  • known soft interventions;
  • unknown intervention targets;
  • off-target effects;
  • stochastic policies;
  • failed or partially executed interventions;
  • interventions whose mechanism differs by environment.

Discovery algorithms should model what actually happened rather than label every experiment as ideal do(X=x) data.

7.10 Latent confounding, cycles, and measurement error

Three realities frequently invalidate standard benchmarks.

Latent confounding

Most real datasets omit causes. Use methods and graph representations that allow hidden common causes, or perform sensitivity analysis.

Cycles

Feedback systems may require time unrolling, equilibrium causal models, cyclic SCMs, or mechanism-centered representations.

Measurement error

A noisy measurement can create or destroy apparent conditional independences. Derived metrics may share computation and therefore correlated error. In software telemetry, two dashboards can look causally related because they reuse the same sampled logs.

7.11 How to choose a method

Situation Reasonable starting family Main warning
i.i.d., no hidden confounding, modest variables PC or GES Assumptions and CI-test power
Hidden confounding plausible FCI/PAG methods Output will remain ambiguous
Linear, non-Gaussian mechanisms plausible LiNGAM Sensitive to model mismatch
Nonlinear additive noise plausible ANM methods Bivariate assumptions may not scale cleanly
Multiple environments or interventions ICP / interventional methods Environment semantics matter
Time series PCMCI or structural temporal methods Sampling, lags, and nonstationarity
Many related datasets and low-latency inference Amortized/foundation approach Synthetic-prior mismatch
Safety-critical domain Ensemble of formal methods plus expert review Never accept one discovered graph as ground truth

7.12 Evaluating causal discovery

Synthetic benchmarks provide ground truth, but they can reward methods whose assumptions match the simulator rather than reality.

Common graph metrics include:

  • structural Hamming distance;
  • adjacency precision and recall;
  • orientation precision and recall;
  • SID-like intervention-distance metrics;
  • calibration of edge probabilities.

More decision-relevant evaluations include:

  • accuracy of unseen interventional distributions;
  • error on target causal effects;
  • success of interventions selected from the graph;
  • stability across environments and resamples;
  • ability to abstain outside supported regimes;
  • recovery of equivalence classes rather than one arbitrary DAG.

A graph can score poorly by edge count yet answer a particular intervention query correctly. The reverse can also happen. Evaluation should match the intended use.

7.13 A practical discovery workflow

flowchart TD
    D["Define variables, time scale, and intervention semantics"] --> K["Encode hard domain constraints"]
    K --> P["Profile missingness, measurement, environments, and selection"]
    P --> M["Run multiple appropriate discovery families"]
    M --> S["Assess stability across samples, thresholds, and methods"]
    S --> E["Represent ambiguity as CPDAG, PAG, or graph posterior"]
    E --> Q["Test target queries and implied independences"]
    Q --> X["Select informative, safe interventions"]
    X --> U["Update model and version assumptions"]
    U --> S

    classDef prep fill:#FFF4D6,stroke:#D97706,color:#3B2F0B,stroke-width:2px;
    classDef learn fill:#E8F1FF,stroke:#2563EB,color:#102A43,stroke-width:2px;
    classDef validate fill:#E8F8EF,stroke:#15803D,color:#123524,stroke-width:2px;
    classDef act fill:#FCE8F3,stroke:#BE185D,color:#4A102A,stroke-width:2px;
    class D,K,P prep;
    class M,E learn;
    class S,Q,U validate;
    class X act;
Loading

7.14 The role of LLMs

An LLM can help:

  • normalize variable descriptions;
  • retrieve domain documentation;
  • propose candidate omitted variables;
  • explain assumptions of algorithms;
  • generate code and diagnostics;
  • compare outputs from multiple methods;
  • help experts record edge constraints and rationales.

An LLM should not silently turn semantic plausibility into causal evidence. A 2026 position paper from researchers including Peter Spirtes and Kun Zhang argues for agent-assisted workflows in which causal claims remain grounded in data, explicit assumptions, formal algorithms, diagnostics, and expert decisions R35.

A safe design records the provenance of every edge:

  • experimentally established;
  • implied by a formal algorithm under named assumptions;
  • supplied by a domain expert;
  • hypothesized by an LLM;
  • unresolved.

7.15 Exercises

  1. List three DAGs in one Markov equivalence class and state the shared conditional independences.
  2. Explain what extra assumption allows LiNGAM to orient edges that a Gaussian linear model cannot.
  3. Design a resampling-stability test for a discovered deployment graph.
  4. Why can a method perform well on synthetic DAGs yet fail on real data?
  5. For one causal edge proposed by an LLM, define an experiment that could falsify it.

8. Active Causal Learning and Experimental Design

Passive causal analysis asks what can be learned from existing data. Active causal learning asks which intervention should be performed next.

This changes causal AI from a reporting system into an experimental system.

8.1 Why active learning matters

Observational data often leave many causal models compatible with the evidence. A small number of targeted interventions can be more informative than a much larger passive dataset.

Suppose two hypotheses remain:

  • $H_1$: service load causes latency;
  • $H_2$: a hidden dependency event causes both load and latency.

Collecting more ordinary traffic may preserve the ambiguity. A controlled load shift, dependency isolation, or routing intervention may separate the hypotheses quickly.

8.2 Intervention types

Perfect intervention

Replaces the natural assignment for $X$ with $X:=x$ and removes incoming causal influences.

Soft intervention

Changes the distribution or parameters of the mechanism for $X$ without fully fixing its value.

Shift intervention

Adds or multiplies a mechanism output, such as increasing a dosage or traffic allocation by a chosen amount.

Stochastic intervention

Sets a treatment distribution or policy rather than a fixed value.

Mechanism intervention

Replaces a specific equation or component while possibly producing the same observed variable value as another intervention.

Compound intervention

Changes multiple variables or mechanisms together.

Real experiments are often soft, stochastic, compound, and imperfect. Logging only the target variable discards important causal information.

8.3 What should an experiment optimize?

The “best” experiment depends on the goal.

Learn the whole graph

Select interventions expected to reduce uncertainty over graph structure.

Answer a target query

Select interventions that reduce uncertainty about a particular effect, counterfactual, or policy value. This can require far fewer experiments than recovering every edge.

Find an optimal intervention

Search directly for actions that achieve a target state, without fully identifying the graph.

Falsify a model

Choose an intervention on which plausible models make maximally different predictions.

Improve control

Choose experiments whose information has high expected decision value after accounting for risk and cost.

Do not spend experiments learning irrelevant portions of a graph when the decision concerns one pathway.

8.4 Expected information gain

Let $M$ represent uncertain causal models, $a$ a candidate intervention, and $Y_a$ its possible result. A common objective is:

$$ EIG(a)= E_{Y_a} \left[ KL\big(P(M\mid D,Y_a,a);||;P(M\mid D)\big) \right]. $$

This measures the expected reduction in model uncertainty.

A cost-aware objective might be:

$$ a^*=\arg\max_a EIG(a)-\lambda C(a)-\rho R(a), $$

where $C(a)$ is operational cost and $R(a)$ is safety risk.

Information gain is not the only criterion. It can favor dramatic but decision-irrelevant experiments. Query-focused or value-of-information objectives are often better.

8.5 Model disagreement as an experimental signal

A practical approximation is to maintain an ensemble of plausible models. For each candidate intervention:

  1. Predict the outcome distribution under every model.
  2. Measure disagreement.
  3. Reject unsafe or infeasible actions.
  4. Select a high-disagreement action.
  5. Observe results and downweight inconsistent models.
flowchart LR
    H["Hypothesis set"] --> P["Predict each intervention"]
    P --> D["Measure disagreement"]
    D --> S["Apply safety and cost constraints"]
    S --> A["Run selected intervention"]
    A --> O["Observe outcome"]
    O --> F["Falsify or reweight models"]
    F --> H

    classDef belief fill:#E8F1FF,stroke:#2563EB,color:#102A43,stroke-width:2px;
    classDef evaluate fill:#FFF4D6,stroke:#D97706,color:#3B2F0B,stroke-width:2px;
    classDef action fill:#FCE8F3,stroke:#BE185D,color:#4A102A,stroke-width:2px;
    classDef evidence fill:#E8F8EF,stroke:#15803D,color:#123524,stroke-width:2px;
    class H,P belief;
    class D,S evaluate;
    class A action;
    class O,F evidence;
Loading

The 2026 Adversarial Causal Intervention Falsification proposal formalizes a related game: a structural generator proposes observational and interventional distributions, while an experimentalist chooses interventions intended to maximally falsify it R40. The key lesson is that observational fit, interventional equivalence, and point identification are different guarantees.

8.6 Adaptive experiments

In adaptive experiments, later actions depend on earlier outcomes. This improves efficiency but complicates inference.

Risks include:

  • optional stopping;
  • changing assignment probabilities;
  • winner’s curse;
  • repeated peeking;
  • feedback from outcomes into measurement;
  • nonstationarity during the experiment.

Log the exact action policy and probability of each action. Use estimators and confidence procedures appropriate for adaptive data.

8.7 Pure exploration versus acting for reward

An experimental agent faces an exploration–exploitation tradeoff.

  • Pure exploration: choose actions to learn mechanisms.
  • Exploitation: choose the currently best action.
  • Dual control: choose actions that both perform well and improve the model.

Causal experimentation is distinct from ordinary bandits when actions alter multiple variables, model structure is uncertain, or counterfactual transfer across interventions matters.

8.8 CausaLab and current agent limitations

CausaLab, released in 2026, places LLM agents in synthetic laboratories governed by hidden SCMs. Agents inspect observations, intervene, predict a held-out outcome, and record a mechanism hypothesis. The benchmark separates task success from faithfulness of the recovered graph and equations R33.

The important finding is qualitative: an agent can predict correctly while recovering the wrong mechanism. Mixed observation–intervention strategies improve structural recovery, but strong agents still struggle to choose informative interventions and may stop experimenting too early.

CausalGame adds challenges such as hidden confounding, selection bias, and measurement error to interactive game settings R37. These benchmarks shift evaluation from answering causal vocabulary questions toward conducting experiments.

8.9 A bounded causal-scientist architecture

flowchart TD
    U["User goal or anomaly"] --> L["LLM proposes explicit hypotheses"]
    L --> V["Formal validator checks variables, graph syntax, and assumptions"]
    V --> B["Bayesian or ensemble causal model"]
    B --> X["Experiment designer scores interventions"]
    X --> G["Safety gate and authorization"]
    G --> R["Simulator, test environment, or real system"]
    R --> O["Structured evidence and provenance"]
    O --> T["Statistical tests and likelihood update"]
    T --> B
    T --> L

    classDef llm fill:#F2EAFE,stroke:#7C3AED,color:#2E1065,stroke-width:2px;
    classDef formal fill:#E8F1FF,stroke:#2563EB,color:#102A43,stroke-width:2px;
    classDef safety fill:#FCE8F3,stroke:#BE185D,color:#4A102A,stroke-width:2px;
    classDef evidence fill:#E8F8EF,stroke:#15803D,color:#123524,stroke-width:2px;
    class U,L llm;
    class V,B,X,T formal;
    class G safety;
    class R,O evidence;
Loading

The LLM is useful for hypothesis generation and interface work. It is not the arbiter of whether evidence supports a causal edge.

8.10 Case study: diagnosing deployment failures

Assume three competing hypotheses:

  • $H_1$: application change causes database saturation;
  • $H_2$: regional network degradation causes both database latency and application errors;
  • $H_3$: autoscaling delay causes application errors under traffic spikes.

Candidate interventions might include:

  • replaying the change under controlled traffic;
  • routing a small traffic slice to a healthy region;
  • pinning replica count before replay;
  • substituting a database mock;
  • changing one timeout without changing load.

A sensible agent would:

  1. Encode predicted observables under each hypothesis.
  2. Remove interventions that exceed a risk budget.
  3. Prefer an intervention whose predicted outcomes differ sharply across hypotheses.
  4. Run it in the least costly environment that preserves the relevant mechanism.
  5. Update confidence and record the intervention’s exact implementation.
  6. Stop when the decision is robust, not when every graph edge is known.

8.11 Experimental validity across simulators and reality

A simulator intervention identifies effects in the simulator. Transferring them to production requires assumptions that the relevant mechanisms are shared.

Maintain separate concepts:

  • simulation validity: did the experiment answer the query inside the model?
  • transport validity: do the mechanisms transfer to the target environment?
  • implementation validity: did the real intervention match its formal specification?

A digital twin that reproduces ordinary logs may still be wrong under novel interventions.

8.12 Exercises

  1. For three competing causal graphs, design one intervention that best distinguishes them.
  2. Explain why maximizing graph-wide information gain can waste experiments.
  3. Give an example of two interventions that set the same observed variable value but replace different mechanisms.
  4. Design a stopping rule for a causal diagnosis agent.
  5. List the provenance fields you would log for every automated intervention.

Part IV — Modern Causal AI

9. Causal Representation Learning

Classical causal analysis usually assumes the variables are already defined: treatment, outcome, confounders, mediators, and so on. Intelligent systems often receive pixels, audio, text, traces, and high-dimensional telemetry instead.

Causal representation learning asks whether a model can recover useful high-level variables with causal meaning from those observations.

9.1 The latent causal model

Let $Z=(Z_1,\dots,Z_k)$ be latent causal variables generated by an SCM, and let observations be produced by a mixing function:

$$ X=g(Z). $$

Examples:

  • $Z$ contains object identity, pose, velocity, and friction; $X$ is video.
  • $Z$ contains service health, dependency state, traffic regime, and rollout phase; $X$ is logs and metrics.
  • $Z$ contains disease processes; $X$ is images, notes, and laboratory measurements.
  • $Z$ contains goals, beliefs, and environment state; $X$ is an agent trajectory.

The aim is not merely to compress $X$. It is to recover a representation whose components correspond, up to stated equivalence, to variables that participate in stable mechanisms and support intervention reasoning.

flowchart LR
    U["Exogenous factors U"] --> Z1["Latent cause Z1"]
    U --> Z2["Latent cause Z2"]
    Z1 --> Z2
    Z1 --> G["Observation process g"]
    Z2 --> G
    G --> X["Pixels, text, or telemetry X"]
    X --> E["Learned encoder"]
    E --> H["Recovered causal state H"]

    classDef hidden fill:#F2EAFE,stroke:#7C3AED,color:#2E1065,stroke-width:2px;
    classDef observe fill:#E8F1FF,stroke:#2563EB,color:#102A43,stroke-width:2px;
    classDef learned fill:#E8F8EF,stroke:#15803D,color:#123524,stroke-width:2px;
    class U,Z1,Z2 hidden;
    class G,X observe;
    class E,H learned;
Loading

9.2 Why ordinary representation learning is not enough

A predictive representation preserves information useful for a training objective. It can entangle causal factors in any invertible way and exploit unstable shortcuts.

Suppose a visual model represents:

$$ H_1 = Z_1 + Z_2, \qquad H_2 = Z_1 - Z_2. $$

This may reconstruct observations perfectly while making interventions on $Z_1$ or $Z_2$ awkward. A causal representation aims for components that correspond to modular, independently changeable aspects of the system.

Desirable properties include:

  • intervention semantics: changing one component has a stable meaning;
  • modularity: mechanisms can change independently;
  • compositionality: known mechanisms combine to predict new situations;
  • sparsity: interventions affect a small subset of mechanisms;
  • temporal persistence: entities and state remain coherent over time;
  • transportability: relevant mechanisms survive environment changes;
  • interpretability: components can be mapped to domain concepts or tested experimentally.

9.3 The identifiability problem

Unsupervised latent representations are generally not unique. If $g$ is invertible, many transformations of $Z$ can reproduce the same observed distribution. Observations alone usually do not reveal which coordinates are the “true” causal variables.

Causal representation learning therefore depends on additional structure. Common sources include:

  • multiple environments;
  • known or unknown interventions;
  • temporal ordering and independent innovations;
  • paired observations before and after sparse changes;
  • multi-view or multimodal observations;
  • object-centric structure;
  • weak labels or grouped observations;
  • assumptions about independent causal mechanisms;
  • known portions of the graph;
  • restricted mixing functions, such as linearity.

A representation theorem is only as meaningful as the equivalence class it identifies. Recovery “up to permutation and component-wise invertible transformations” is much stronger than arbitrary invertible recovery, but it still does not assign human-readable names without grounding.

9.4 Interventions as representation supervision

Suppose an intervention changes one latent mechanism while leaving others unchanged. Paired observations before and after the intervention reveal which variation should be localized in representation space.

flowchart LR
    X0["Observation before"] --> ENC["Shared encoder"]
    X1["Observation after intervention"] --> ENC
    ENC --> Z["Latent causal state"]
    I["Intervention metadata"] --> L["Sparse-change / mechanism loss"]
    Z --> L
    L --> ENC

    classDef data fill:#E8F1FF,stroke:#2563EB,color:#102A43,stroke-width:2px;
    classDef model fill:#F2EAFE,stroke:#7C3AED,color:#2E1065,stroke-width:2px;
    classDef supervision fill:#FFF4D6,stroke:#D97706,color:#3B2F0B,stroke-width:2px;
    class X0,X1 data;
    class ENC,Z model;
    class I,L supervision;
Loading

Intervention targets do not always need to be known. Some methods jointly infer which latent components changed and what representation generated them.

A 2026 result moves the theory beyond asymptotic identifiability: under a high-dimensional linear mixing setting, causal representations, the latent graph, and unknown intervention targets can be consistently recovered with a logarithmic number of unknown multi-node intervention environments, with explicit finite-sample guarantees R32. The assumptions remain specialized, but the shift toward sample complexity is important.

9.5 Multiple environments and invariance

If the same latent mechanism operates across environments while upstream distributions change, a representation can be trained to expose stable conditional relationships.

For example, deployment telemetry from different regions may vary in traffic and hardware while sharing an application-error mechanism. A representation that isolates the shared mechanism may transfer better than one that uses region-specific shortcuts.

But invariance is not automatically causal:

  • a stable shortcut can persist across all training environments;
  • the chosen environments may not vary the relevant causes;
  • a true mechanism may change due to version differences;
  • enforcing too much invariance can erase useful causal heterogeneity.

The environments must be sufficiently diverse and causally informative.

9.6 Temporal and object-centric structure

Time provides several useful biases:

  • latent causes at $t$ precede effects at $t+1$;
  • entities persist;
  • actions create localized state changes;
  • independent innovations enter specific mechanisms;
  • repeated transitions expose stable dynamics.

Object-centric models add another bias: scenes are composed of entities with properties and relations. This can align representation components with intervention targets such as “move object,” “change friction,” or “block service route.”

Neither temporal prediction nor object slots guarantee causal semantics. The test is whether the representation supports correct predictions under interventions and mechanism changes.

9.7 Multimodal causal representations

Different modalities can reveal complementary aspects of a latent state:

  • video shows motion;
  • audio reveals contact or failure modes;
  • text describes goals or interventions;
  • logs record discrete events;
  • metrics provide continuous trajectories.

Multi-view structure can aid identifiability when views share causal factors but have conditionally independent noise or distinct observation functions.

For a software system, combine:

  • source and configuration diffs;
  • deployment events;
  • traces and logs;
  • topology and ownership metadata;
  • test outcomes;
  • incidents and remediation actions.

A causal representation should distinguish a mechanism-changing event from a correlated symptom, even when both appear in text and telemetry.

9.8 Evaluation

Reconstruction loss is not a causal metric. Useful evaluations include:

  1. Latent recovery: correlation or equivalence-aware recovery of known factors in synthetic data.
  2. Graph recovery: accuracy of causal relations among latent components.
  3. Intervention localization: whether an intervention changes the correct components.
  4. Unseen intervention prediction: accuracy under interventions not represented in training.
  5. Mechanism shift: transfer when one mechanism changes.
  6. Counterfactual consistency: whether factual and counterfactual worlds share inferred background state appropriately.
  7. Downstream control: sample efficiency and robustness of policies using the representation.
  8. Abstention: whether the model detects unsupported environments.

CausalVerse, introduced in 2025, is an example of a benchmark designed to combine realistic visual complexity with access to ground-truth generating processes across images, physical simulations, robotics, and traffic scenes R31.

9.9 A Proofline-style infrastructure representation

Raw observations might include hundreds of metrics and thousands of events. A useful causal state could contain:

  • deployment version;
  • rollout fraction;
  • traffic regime;
  • dependency availability;
  • resource saturation;
  • schema compatibility;
  • credential validity;
  • queue pressure;
  • active failure mode;
  • remediation state.

The representation should support interventions such as:

  • roll back version;
  • reroute traffic;
  • increase replicas;
  • rotate credentials;
  • disable a feature flag;
  • isolate a dependency;
  • replay a request set.

Training signals can come from historical interventions, canary rollouts, chaos experiments, test environments, and paired before/after traces. The model should not infer causal state solely from incident labels, which are heavily selected and often incomplete.

9.10 Practical design principles

  • Start with variables needed for a small set of interventions, not universal ontology discovery.
  • Use known action and environment metadata aggressively.
  • Separate observation encoders from the explicit causal state model.
  • Preserve uncertainty over latent state and mappings.
  • Test sparse intervention effects.
  • Version representations when mechanisms or instrumentation change.
  • Provide human-readable probes, but do not equate probe accuracy with causal identification.
  • Evaluate on held-out interventions, not only held-out rows.

9.11 Exercises

  1. Give two observationally equivalent latent representations that differ in intervention usefulness.
  2. What additional supervision would help recover causal variables from deployment logs?
  3. Design an unseen-intervention evaluation for a video world model.
  4. Explain why object-centric slots can help causal representation learning but do not guarantee it.
  5. Define the equivalence class you would accept when learning latent infrastructure state.

10. Causal Foundation Models

Foundation models for causality aim to amortize causal inference or discovery across many datasets. Instead of fitting a new model from scratch for every study, a model is pretrained on a distribution of causal tasks and performs in-context inference on a new dataset.

This is one of the fastest-moving areas of causal AI as of August 2026.

10.1 Prior-data fitted networks

A prior-data fitted network is trained on synthetic datasets sampled from a prior over data-generating processes.

For each training task:

  1. Sample a causal model $M\sim P(M)$.
  2. Sample a dataset $D\sim P(D\mid M)$.
  3. Sample a causal query $q$.
  4. Compute or simulate the target answer $\theta(M,q)$.
  5. Train a transformer to predict $\theta$ from $(D,q)$.

With enough coverage and optimization, the network approximates Bayesian inference under the training prior:

$$ f_\phi(D,q)\approx P(\theta\mid D,q). $$

flowchart LR
    P["Prior over SCMs"] --> S["Sample graphs, mechanisms, and noise"]
    S --> D["Generate observational and interventional datasets"]
    D --> T["Train transformer on causal queries"]
    T --> N["New dataset in context"]
    N --> O["Zero-shot effect, graph, or bounds"]

    classDef prior fill:#F2EAFE,stroke:#7C3AED,color:#2E1065,stroke-width:2px;
    classDef synth fill:#FFF4D6,stroke:#D97706,color:#3B2F0B,stroke-width:2px;
    classDef train fill:#E8F1FF,stroke:#2563EB,color:#102A43,stroke-width:2px;
    classDef infer fill:#E8F8EF,stroke:#15803D,color:#123524,stroke-width:2px;
    class P prior;
    class S,D synth;
    class T,N train;
    class O infer;
Loading

The model’s causal competence is therefore inseparable from its synthetic prior.

10.2 Why the approach is attractive

Traditional causal workflows are dataset-specific and computationally repeated. A causal foundation model can potentially provide:

  • zero-shot or few-shot inference;
  • fast amortized computation;
  • uncertainty outputs;
  • a unified interface across dataset sizes and mechanisms;
  • reusable inductive biases learned from many SCMs;
  • rapid hypothesis screening before more expensive analysis.

The analogy to language-model pretraining is imperfect. Causal tasks have hard identifiability limits, and synthetic SCM families explicitly encode assumptions about how worlds work.

10.3 Foundation models for effect estimation

CausalFM trains PFN-style models on SCM-based priors for causal effect estimation across several settings, including heterogeneous effects and instrumental-variable or frontdoor-like structures. At inference time, it can estimate effects for a new dataset without task-specific gradient updates R25.

Do-PFN similarly targets in-context estimation of interventional outcomes R26. Other 2026 work extends the recipe to time-series, continuous-time, panel, and longitudinal settings R47, R48, R49.

The central design problem is the prior:

  • Which graph families?
  • Which functional mechanisms?
  • Which noise distributions?
  • Which confounding patterns?
  • Which intervention types?
  • Which sample sizes and variable counts?
  • Which missingness and measurement processes?

A model trained on clean acyclic synthetic SCMs may be badly calibrated on selected, cyclic, noisy production data.

10.4 Foundation models for causal discovery

Arrow

Arrow factorizes a DAG into an undirected skeleton and a topological order, which guarantees acyclicity by construction. A transformer processes a new tabular dataset and predicts skeleton probabilities and node-order scores. It is pretrained across diverse synthetic graph and mechanism families for zero-shot discovery R27.

CDFM

CDFM treats unknown causal mechanisms as latent variables and uses a variational decomposition to organize a general-purpose discovery architecture. Its theoretical framing emphasizes that causal prior mechanisms are indispensable when observational identifiability is absent R28.

DAG-FM

DAG-FM decomposes discovery into autoregressive leaf-node and parent-node prediction, with a mixture-of-experts mechanism intended to adapt across heterogeneous functional causal model families R29.

These systems promise speed and reuse. Their strongest experimental evidence is still dominated by synthetic, semi-synthetic, and relatively small real benchmark graphs. Claims of general-purpose discovery should be read as a research direction, not solved causality.

10.5 Partial graphs and structural knowledge

A practical causal foundation model should accept partial knowledge:

  • required edges;
  • forbidden edges;
  • temporal ordering;
  • known intervention targets;
  • uncertain latent confounding;
  • equivalence-class constraints;
  • mechanism types for some variables.

Causal Foundation Models with Partial Graphs explores conditioning a model on incomplete graph information rather than assuming either a blank slate or a complete known structure R30. This is important in real systems, where domain knowledge is substantial but incomplete.

10.6 Partial identification as a first-class output

A neural model should not manufacture a point estimate when the target is only bounded. Foundation Models for Partial Causal Identification proposes training over a canonical prior with broad support on discrete SCMs so the model learns distributions over causal queries compatible with observations and structural assumptions R38.

Conceptually, the output should distinguish:

  • posterior concentration caused by data;
  • concentration caused by the prior;
  • the identified set implied by assumptions;
  • uncertainty caused by finite samples.

This is a more credible direction than presenting a single confident number for every query.

10.7 Causal foundation model versus LLM

A causal foundation model need not process natural language. It may be a transformer over rows, columns, intervention masks, graphs, and query tokens.

An LLM can sit around it:

  • translate a natural-language question into a formal query;
  • retrieve variable definitions;
  • propose assumptions for review;
  • invoke the causal model;
  • explain results and limitations.

But the causal foundation model’s outputs come from its formal training distribution and inputs, not from textual plausibility.

flowchart LR
    NL["Natural-language question"] --> LLM["LLM interface"]
    LLM --> SPEC["Typed causal query and assumptions"]
    DATA["Dataset + intervention metadata"] --> CFM["Causal foundation model"]
    SPEC --> CFM
    CFM --> RES["Effects, graph posterior, or bounds"]
    RES --> CHECK["Diagnostics and sensitivity layer"]
    CHECK --> LLM
    LLM --> REPORT["Explanation with provenance"]

    classDef language fill:#F2EAFE,stroke:#7C3AED,color:#2E1065,stroke-width:2px;
    classDef formal fill:#E8F1FF,stroke:#2563EB,color:#102A43,stroke-width:2px;
    classDef evidence fill:#E8F8EF,stroke:#15803D,color:#123524,stroke-width:2px;
    class NL,LLM,REPORT language;
    class SPEC,CFM,CHECK formal;
    class DATA,RES evidence;
Loading

10.8 Prior mismatch

Because pretraining is Bayesian inference under an engineered task distribution, prior mismatch is the central risk.

Examples:

  • graphs are denser than training graphs;
  • real mechanisms contain discontinuities or thresholds absent from the simulator;
  • noise is dependent or heavy-tailed;
  • measurement error creates false dependencies;
  • variables are selected after treatment;
  • treatment assignment is adaptive;
  • the system contains feedback;
  • interventions are soft or off-target;
  • sample units are not i.i.d.;
  • semantic variable selection changes the target graph.

A model can be fast, accurate on benchmark priors, and systematically wrong in a domain whose mechanisms lie outside them.

10.9 Evaluation checklist

A causal foundation model should be evaluated on:

  1. In-prior accuracy: Does it solve tasks drawn from its training family?
  2. Mechanism OOD: What happens under unseen functional forms and noise?
  3. Graph OOD: What happens with different density, motifs, and variable counts?
  4. Intervention OOD: Can it handle new targets, strengths, and intervention types?
  5. Measurement realism: Missingness, selection, aggregation, and noisy variables.
  6. Calibration: Do probabilities and intervals match empirical frequencies?
  7. Equivalence awareness: Does it avoid arbitrary directions when data are insufficient?
  8. Partial identification: Does it return bounds or abstain when appropriate?
  9. Downstream query accuracy: Are intervention predictions correct?
  10. Real-data falsification: Do proposed edges survive targeted experiments?

10.10 How to build a domain causal foundation model

For a bounded domain such as deployment systems:

  1. Define a typed variable and mechanism vocabulary.
  2. Build generative SCM families from real architectural motifs.
  3. Include hidden variables, cycles through time, failures, missing data, and selection.
  4. Sample observational and intervention logs with realistic instrumentation.
  5. Train on multiple outputs: graph equivalence class, effect queries, intervention predictions, and uncertainty.
  6. Condition on partial topology and domain constraints.
  7. Calibrate against historical experiments and test environments.
  8. Detect prior mismatch and abstain.
  9. Use active experiments to update or fine-tune the model.
  10. Preserve a formal, inspectable model outside the neural network.

This is more promising than pretraining on arbitrary random DAGs and expecting universal transfer.

10.11 Exercises

  1. Explain why a PFN can be interpreted as amortized Bayesian inference.
  2. Design a synthetic prior for a causal foundation model of microservice deployments.
  3. Give three examples of prior mismatch that ordinary graph benchmarks would miss.
  4. Why should a causal foundation model output an equivalence class or bounds rather than always a point graph and point effect?
  5. Separate the roles of an LLM and a causal foundation model in a production architecture.

11. LLMs, Agents, and Causal Reasoning

LLMs can fluently discuss causality. The harder question is whether their conclusions are grounded in formal structure and evidence rather than semantic associations.

11.1 Five distinct uses of LLMs around causality

1. Causal knowledge retrieval

The model recalls plausible mechanisms from text. This is useful prior knowledge, not proof.

2. Formal causal reasoning

The model is given a graph, probabilities, or equations and computes an intervention or counterfactual answer.

3. Causal data science

The model inspects files, proposes an estimand, selects methods, writes code, checks diagnostics, and explains uncertainty.

4. Active causal scientist

The model maintains hypotheses, selects experiments, observes outcomes, and revises a causal model.

5. Causal analysis of the LLM system

Causal methods estimate effects of pretraining data, alignment choices, prompts, tools, routing policies, judges, or agent components on outcomes.

These uses have different evidence requirements. Good performance on formal graph questions does not establish competence at discovery from messy data or experimental design.

11.2 Semantic plausibility is a dangerous shortcut

Language contains abundant causal statements and stereotypes. An LLM may orient “smoking” and “cancer” correctly because it remembers world knowledge. It may also invent a plausible direction where the data and assumptions are ambiguous.

CausalFlip constructs semantically similar question pairs with opposite causal answers, using confounder, chain, and collider structures. Its results show that explicit chain-of-thought can still follow spurious semantic cues; a training method that internalizes causal computation performs better on the benchmark R36.

The lesson is not that hidden reasoning is inherently causal. It is that verbal explanations can be faithful-looking while following the wrong signal.

11.3 Post-training causal reasoning

CauGym provides training data across several formal causal tasks and evaluates supervised and reinforcement-learning post-training methods. The work reports that targeted post-training can make smaller models competitive with larger ones on causal benchmarks, with GRPO performing strongly in the studied setup R34.

This suggests formal causal skills can be trained. It does not show that the model can infer causal structure from arbitrary observational data or operate safely as an autonomous scientist.

A useful distinction is:

  • executing a supplied causal calculus;
  • choosing the correct causal model and assumptions;
  • collecting evidence that distinguishes models.

The latter two are much harder.

11.4 Benchmarks are becoming more realistic

InterveneBench

Evaluates end-to-end study-design reasoning around policy interventions rather than only local graph questions. Reported state-of-the-art models struggle, and a theory-guided multi-agent scaffold improves performance R42.

CausaLab

Evaluates interactive discovery inside hidden SCM laboratories and scores both task success and mechanism fidelity R33.

CausalDS

Evaluates causal reasoning in file-backed data-science scenes, complementing symbolic benchmarks and interactive laboratories R41.

CausalGame

Adds interactive experimental design with hidden confounding, selection bias, and measurement error R37.

ReplaySCM

Pushes toward executable mechanism recovery rather than isolated answers R50.

The trend is from “Can the model answer a causal multiple-choice question?” toward “Can it build, test, and execute a coherent causal model?”

11.5 Why passive LLMs face a structural limit

If two causal models generate the same observational distribution, text-only or observational fine-tuning cannot distinguish them merely by fitting that distribution. A 2026 theoretical paper frames this as a limitation of passive LLM causal discovery and argues that interventional agents can escape the ambiguity by acting on the environment R43.

World knowledge can supply a prior, but then the answer depends on the prior. That can be valuable, provided it is labeled and tested.

11.6 A reliable division of labor

Task LLM Formal causal/statistical system Human/domain owner
Interpret natural-language goal Lead Validate types Confirm intent
Retrieve domain context Lead Track provenance Judge relevance
Propose graph hypotheses Suggest only Store as uncertain Approve or reject constraints
Determine identifiability Explain Compute formally Review assumptions
Fit estimators Orchestrate/code Execute and test Approve target and data
Select intervention Generate candidates Score information, cost, risk Authorize high-impact actions
Update causal beliefs Summarize Likelihood/tests/posterior Resolve semantic disputes
Communicate result Lead Supply exact values and caveats Own decision

11.7 Typed causal tools

Do not ask an LLM to emit free-form causal conclusions when it can call typed tools.

A tool interface might require:

{
  "query_type": "average_treatment_effect",
  "treatment": "extended_test_suite",
  "treatment_values": [0, 1],
  "outcome": "incident_within_24h",
  "population": "production_deployments",
  "adjustment_set": ["change_complexity", "service_tier", "traffic_forecast"],
  "graph_version": "deployment-scm-v18",
  "assumption_ids": ["A12", "A19", "A27"]
}

The execution layer should reject invalid variables, descendants in an adjustment set, unsupported graph versions, or unidentified queries.

11.8 Hypothesis provenance

Every hypothesis should carry:

  • natural-language statement;
  • formal graph or equation change;
  • source: literature, expert, LLM, discovery algorithm, or experiment;
  • confidence and calibration method;
  • predicted observations under candidate interventions;
  • known counterevidence;
  • assumptions and scope;
  • status: proposed, testable, supported, contradicted, or unresolved.

This turns an agent trace into scientific state rather than a chat transcript.

11.9 Preventing premature stopping

Agents often accept the first coherent explanation. Countermeasures include:

  • require at least one competing hypothesis;
  • ask which evidence would be likely under both hypotheses;
  • require an intervention-disagreement score;
  • check every hypothesis against all accumulated observations;
  • reserve a falsification budget;
  • separate “enough to act” from “mechanism established”;
  • use a stop rule tied to decision robustness or posterior concentration.

CausaLab’s finding that agents may stop before gathering informative evidence makes this a first-class design concern R33.

11.10 Causal methods for building LLM systems

Causality is not only an ability to add to models. It is a methodology for developing them.

Questions include:

  • What is the effect of adding a data source during pretraining?
  • Does chain-of-thought supervision improve correctness or merely answer style?
  • What is the effect of a tool call on task success for prompts that trigger the tool?
  • Does routing to a larger model improve outcomes after accounting for prompt difficulty?
  • How does an LLM judge’s style preference affect rankings?
  • Which component caused an agent failure?

Logged model-development data are confounded by adaptive choices. For example, hard prompts are routed to stronger models, so raw outcome comparisons can make stronger models appear worse. Causal Methods for LLM Development and Evaluation maps identification and estimation problems across pretraining, alignment, routing, judging, and agentic workflows R44.

11.11 A causal audit of an agent

Represent the system:

flowchart LR
    Q["Prompt difficulty"] --> R["Model routing"]
    Q --> Y["Task success"]
    R --> M["Chosen model"]
    M --> T["Tool calls"]
    T --> Y
    M --> Y
    J["Judge bias"] --> S["Recorded score"]
    Y --> S

    classDef conf fill:#FFF4D6,stroke:#D97706,color:#3B2F0B,stroke-width:2px;
    classDef action fill:#E8F1FF,stroke:#2563EB,color:#102A43,stroke-width:2px;
    classDef out fill:#E8F8EF,stroke:#15803D,color:#123524,stroke-width:2px;
    classDef measure fill:#FCE8F3,stroke:#BE185D,color:#4A102A,stroke-width:2px;
    class Q,J conf;
    class R,M,T action;
    class Y out;
    class S measure;
Loading

The raw relationship between tool use and success is confounded by task difficulty. The recorded score is also affected by judge bias. A causal audit separates policy effects, component effects, and measurement effects.

11.12 Exercises

  1. Distinguish causal knowledge retrieval from causal discovery.
  2. Design a typed tool call for a frontdoor query.
  3. Give an example where an LLM’s world knowledge is a useful prior and an example where it is harmful.
  4. Draw a causal graph for evaluating whether an agent’s browser tool improves task success.
  5. Define a stopping rule that prevents an agent from confusing a plausible mechanism with an established one.

12. Causal World Models

A world model predicts aspects of an environment and supports planning. A causal world model additionally represents interventions and mechanisms well enough to predict how the environment changes when an agent acts, including under shifts not captured by ordinary observational correlation.

12.1 From predictive dynamics to causal dynamics

A conventional dynamics model learns something like:

$$ P(S_{t+1}\mid S_t,A_t). $$

If historical actions $A_t$ were chosen by a policy based on hidden state, this conditional can mix action effects with selection effects.

A causal target is closer to:

$$ P(S_{t+1}\mid do(A_t=a),S_t=s), $$

with explicit assumptions about which state is observed, how actions are implemented, and which mechanisms remain stable.

In a fully observed Markov decision process with randomized or sufficiently explored actions, the distinction may collapse operationally. In partially observed, confounded, multi-agent, or offline settings, it matters.

12.2 Components of a causal world model

flowchart LR
    O["Raw observations"] --> E["State encoder"]
    E --> S["Entity and causal state"]
    A["Action / intervention"] --> M["Mechanism model"]
    S --> M
    M --> N["Next-state and outcome distribution"]
    N --> P["Planner"]
    P --> A
    N --> C["Counterfactual evaluator"]
    S --> C

    classDef obs fill:#E8F1FF,stroke:#2563EB,color:#102A43,stroke-width:2px;
    classDef state fill:#F2EAFE,stroke:#7C3AED,color:#2E1065,stroke-width:2px;
    classDef mechanism fill:#FFF4D6,stroke:#D97706,color:#3B2F0B,stroke-width:2px;
    classDef action fill:#FCE8F3,stroke:#BE185D,color:#4A102A,stroke-width:2px;
    classDef outcome fill:#E8F8EF,stroke:#15803D,color:#123524,stroke-width:2px;
    class O,E obs;
    class S state;
    class M mechanism;
    class A,P action;
    class N,C outcome;
Loading

A useful model needs:

  • state: sufficient information for the relevant future and queries;
  • entities and relations: stable objects or components;
  • mechanisms: modular transition or structural functions;
  • action semantics: what each intervention replaces or changes;
  • uncertainty: over state, structure, parameters, and environment;
  • observation model: how hidden state produces sensors or text;
  • counterfactual coupling: how alternative actions share background circumstances;
  • scope and abstention: where the model should not be trusted.

12.3 Levels of abstraction

A causal world model can operate at different levels:

  1. Pixel or waveform level: detailed but expensive and often causally entangled.
  2. Latent dynamics level: compact state optimized for prediction or control.
  3. Object or entity level: persistent entities, properties, and relations.
  4. Mechanism level: modular equations describing how actions and entities interact.
  5. Conceptual level: human-meaningful states and rules used for explanation and planning.

The 2026 paper A Unifying Perspective on Causal World Models argues for connecting observations, representations, and causal structure, and defines usefulness relative to the tasks the model must support R45.

No single level is universally best. A robot may need pixel-level contact geometry and object-level causal planning. An infrastructure agent may need raw traces for detection and mechanism-level service state for intervention.

12.4 Observational fit is not causal validity

Two world models can predict logged trajectories equally well while disagreeing under a new action. This occurs when:

  • the behavior policy never explored key actions;
  • hidden variables influenced both action and outcome;
  • the models encode different causal directions;
  • multiple mechanisms yield the same equilibrium observations;
  • generated futures are visually plausible but not action-sensitive.

Validation must include held-out interventions and policy shifts.

12.5 Failed actions are valuable evidence

Datasets often contain successful demonstrations and filter failures. For causal learning, failures can be especially informative because they reveal which actions do not produce the intended transition and expose hidden preconditions.

A world-action model trained only on successful trajectories may learn to generate success-looking futures regardless of the action. Include:

  • failed grasps;
  • rejected deployments;
  • partial rollouts;
  • rollback attempts;
  • tool calls that returned irrelevant evidence;
  • off-target interventions;
  • recovery sequences.

The model should predict failure modes, not merely imitate desired endpoints.

12.6 Counterfactual planning

A planner evaluates alternative action sequences while keeping relevant background state coupled.

For a sequence $a_{t:t+H}$:

$$ P(S_{t+1:t+H}\mid do(A_{t:t+H}=a_{t:t+H}),E_t), $$

where $E_t$ is current evidence.

Counterfactual planning should distinguish:

  • uncertainty about current hidden state;
  • stochastic future noise;
  • uncertainty about mechanisms;
  • uncertainty caused by unsupported actions.

A plan that is optimal under one guessed model may be fragile. Robust planning evaluates a set or posterior of models and may prefer an action with slightly lower expected reward but lower worst-case loss.

12.7 Equilibrium systems and mechanism-specific interventions

In equilibrium systems, two actions can enforce the same variable value by replacing different mechanisms, producing different downstream consequences. Standard $do(X=x)$ notation can be too coarse.

Bipartite graphical causal models represent variable nodes and equation nodes separately. An intervention names the equation being replaced, target variable, and value R39. This is relevant to:

  • physical equilibrium;
  • economic markets;
  • biological regulation;
  • distributed systems;
  • control loops;
  • policy systems with multiple implementation routes.

For example, “latency is 100 ms” can result from throttling traffic, adding capacity, caching, or dropping requests. Those mechanisms are not interchangeable.

12.8 Causal abstraction

A high-level model is valid when interventions and outcomes at the lower level map consistently to the higher level for the queries of interest.

An abstraction need not reproduce every detail. It must preserve the causal relationships needed for decisions.

For a deployment system, a high-level node dependency_unavailable might summarize many packet, DNS, authentication, and regional-failure states. The abstraction is useful if interventions such as failover or isolation have predictable meanings at that level.

Causal abstraction is query-relative. A state sufficient for incident prevention may be insufficient for explaining tail latency.

12.9 World models and control

A causal graph can be structurally accurate yet fail to improve control if:

  • effect magnitudes are wrong;
  • relevant state is missing;
  • planning error accumulates;
  • the action space is misrepresented;
  • the controller already performs near optimally;
  • model use adds latency;
  • the model cannot detect when it is wrong.

A causal model should earn its place through decision outcomes. Reliability certificates, uncertainty thresholds, and fallback policies can be more important than small graph-metric gains.

12.10 Evaluation hierarchy

flowchart TD
    L1["1. Observational prediction"] --> L2["2. Held-out action prediction"]
    L2 --> L3["3. Novel intervention prediction"]
    L3 --> L4["4. Mechanism-shift transfer"]
    L4 --> L5["5. Counterfactual consistency"]
    L5 --> L6["6. Better decisions under cost and risk"]
    L6 --> L7["7. Calibrated abstention outside scope"]

    classDef low fill:#E8F1FF,stroke:#2563EB,color:#102A43,stroke-width:2px;
    classDef mid fill:#FFF4D6,stroke:#D97706,color:#3B2F0B,stroke-width:2px;
    classDef high fill:#E8F8EF,stroke:#15803D,color:#123524,stroke-width:2px;
    class L1,L2 low;
    class L3,L4,L5 mid;
    class L6,L7 high;
Loading

Passing level 1 does not imply level 3. A model intended for autonomous intervention should be evaluated near the top of the hierarchy.

12.11 Infrastructure world model example

A Proofline-style world model might combine:

  • a typed graph of services, resources, credentials, routes, environments, and deployments;
  • event-sourced state transitions;
  • structural mechanisms for rollout, routing, scaling, and failure propagation;
  • uncertain latent failure modes;
  • test and production evidence;
  • intervention operators such as deploy, rollback, reroute, isolate, rotate, and replay;
  • an LLM that proposes hypotheses and explains results;
  • formal validators and simulators that decide whether evidence supports them.

Its main output should not be a pretty causal graph. It should be defensible predictions such as:

Under this exact rollout intervention and current evidence, the probability of database saturation increases from a bounded baseline range to a higher range; the result depends primarily on assumptions A17 and A22, and a 5% traffic canary would most reduce the remaining ambiguity.

12.12 Exercises

  1. Give two world models with equal observational accuracy but different action predictions.
  2. Define intervention semantics for “restart service” at three levels of abstraction.
  3. Design a held-out mechanism-shift evaluation for a deployment world model.
  4. Why can a correct graph fail to improve control?
  5. Write a counterfactual query that requires coupling latent background state across two action worlds.

Part V — Building Causal AI Systems

13. Production Architecture and Engineering Discipline

Causal AI fails in production less often because someone forgot a formula than because assumptions, data lineage, intervention semantics, and model versions were not treated as software artifacts.

13.1 The production lifecycle

flowchart TD
    Q["1. Specify causal question"] --> G["2. Build or select causal model"]
    G --> D["3. Audit data and measurement"]
    D --> I["4. Identify estimand"]
    I --> E["5. Estimate and quantify uncertainty"]
    E --> V["6. Refute, stress-test, and review"]
    V --> P["7. Convert estimate into policy"]
    P --> X["8. Execute authorized intervention"]
    X --> M["9. Monitor outcomes and mechanism shift"]
    M --> G

    classDef spec fill:#FFF4D6,stroke:#D97706,color:#3B2F0B,stroke-width:2px;
    classDef model fill:#E8F1FF,stroke:#2563EB,color:#102A43,stroke-width:2px;
    classDef validate fill:#E8F8EF,stroke:#15803D,color:#123524,stroke-width:2px;
    classDef action fill:#FCE8F3,stroke:#BE185D,color:#4A102A,stroke-width:2px;
    class Q,G,D,I spec;
    class E model;
    class V,M validate;
    class P,X action;
Loading

Each step should produce versioned artifacts and machine-readable metadata.

13.2 A typed causal query specification

A query should include:

  • treatment or intervention operator;
  • treatment values or policy;
  • outcome definition;
  • time horizon;
  • target population;
  • estimand type;
  • graph or SCM version;
  • candidate adjustment set or identification strategy;
  • interference assumptions;
  • transport environment;
  • acceptable uncertainty and decision threshold.

Example:

query_id: test-depth-incident-ate-v3
estimand: average_treatment_effect
population:
  entity: production_deployment
  filters:
    service_tier: [critical, high]
    date_range: 2026-01-01/2026-08-01
intervention:
  variable: extended_test_suite
  from: false
  to: true
  semantics: run_suite_version_12_before_rollout
outcome:
  variable: sev1_or_sev2_incident
  horizon: 24h
graph_version: deployment-scm-18
identification:
  method: backdoor
  adjustment_set:
    - change_complexity
    - service_tier
    - dependency_risk
    - forecast_traffic
assumptions:
  - no_unmeasured_test_assignment_confounding_v5
  - no_cross_deployment_interference_within_24h_v2
  - positivity_in_declared_population_v4
decision:
  incident_cost_eur: 75000
  max_test_cost_eur: 3500
  require_lower_confidence_bound_positive: true

The point is not YAML. The point is that causal intent should not live only in a notebook or prompt.

13.3 Causal data contracts

A causal data contract extends an ordinary schema contract with semantics that affect identification.

For every variable, record:

  • entity and unit of analysis;
  • event time and observation time;
  • whether it is pre-treatment or post-treatment for named queries;
  • measurement procedure;
  • missingness mechanism and sentinel values;
  • aggregation window;
  • intervention status and target;
  • data source and transformations;
  • known descendants and shared measurement inputs;
  • valid environments and versions.

A variable called traffic_at_deploy is ambiguous unless you know whether it was forecast before treatment, measured during rollout, or aggregated after an incident. Those versions occupy different graph positions.

13.4 Event sourcing for interventions

Every action should be logged as an event with:

  • requested action;
  • authorized action;
  • actual action executed;
  • exact parameters;
  • start and completion times;
  • mechanism or component targeted;
  • partial failures and off-target effects;
  • actor or policy version;
  • assignment probability, when randomized or adaptive;
  • environment and system version;
  • observed immediate consequences.

This supports later analysis of implementation fidelity and unknown intervention targets.

13.5 Causal model registry

Treat graphs and SCMs like versioned code.

A registry entry should contain:

  • graph structure;
  • equation or mechanism definitions where available;
  • latent-variable declarations;
  • time scale;
  • supported intervention operators;
  • parameter posterior or fitted models;
  • assumptions and evidence;
  • unresolved edges and equivalence-class information;
  • tests and benchmark results;
  • owners and reviewers;
  • lineage from previous versions;
  • deprecation conditions.

Graph diffs should be reviewable:

+ dependency_health -> incident
- engineer_seniority -> incident
? traffic <-> dependency_health   # latent confounding unresolved
~ rollback mechanism v2 -> v3    # now distinguishes failover from binary rollback

13.6 Assumption registry

Assumptions need identifiers, owners, evidence, and scope.

Example:

assumption_id: no_unmeasured_test_assignment_confounding_v5
statement: >
  Conditional on change complexity, service tier, dependency risk,
  forecast traffic, and team release policy, no unmeasured pre-treatment
  variable jointly affects extended-test assignment and 24-hour incident risk.
status: disputed
supporting_evidence:
  - release-policy audit 2026-Q2
  - balance report test-depth-v7
counterevidence:
  - manual reviewer confidence is not logged
sensitivity_analysis: omitted-confounder-grid-v4
owner: reliability-science
review_after: 2026-10-01

This prevents an assumption from becoming permanent simply because it once appeared in code.

13.7 Testing causal software

Unit tests

  • intervention removes or replaces the correct mechanism;
  • graph traversal finds expected adjustment sets;
  • descendants are rejected from inappropriate control sets;
  • counterfactual worlds share exogenous state correctly;
  • effect sign and scale match analytic toy SCMs.

Property-based tests

Generate random SCMs satisfying known conditions and verify:

  • identified formulas agree with simulated interventions;
  • estimators converge with increasing sample size;
  • d-separation matches implied independences;
  • graph serialization preserves endpoint semantics;
  • bounds contain true effects under the test assumptions.

Metamorphic tests

  • renaming variables does not change formal results;
  • adding an irrelevant independent variable does not change the estimand;
  • duplicating rows affects uncertainty but not point logic;
  • rescaling units changes coefficients appropriately but not causal conclusions;
  • permuting dataset rows changes nothing.

Adversarial tests

  • introduce hidden confounding;
  • violate positivity;
  • add measurement error;
  • select on a collider;
  • change a mechanism between train and test;
  • execute an off-target intervention;
  • prompt the LLM with a semantically plausible but graph-inconsistent claim.

13.8 Refutation is a product feature

A production interface should make it easy to ask:

  • Which assumption, if removed, destroys identification?
  • Which omitted variable strength reverses the conclusion?
  • Which graph alternatives remain compatible with the data?
  • Which negative controls failed?
  • Which experiment would most reduce uncertainty?
  • Is the current unit outside overlap?
  • Did the mechanism change after the last deployment or policy update?

The result should be a defensible argument, not only an estimate.

13.9 Monitoring causal validity

Ordinary model monitoring tracks input drift and prediction error. Causal monitoring should also track:

  • treatment-policy drift;
  • overlap and propensity support;
  • changes in adjustment-variable distributions;
  • violations of previously stable conditional independences;
  • mechanism residual shifts;
  • intervention implementation drift;
  • outcome-definition changes;
  • selection and missingness changes;
  • disagreement among graph or mechanism models;
  • calibration under recent interventions.

A mechanism can change while marginal prediction remains acceptable, especially if compensating correlations arise.

13.10 Safety gates

Automated interventions should be tiered.

Tier Example Required control
0 Read-only analysis Provenance and reproducibility
1 Simulator or test environment Resource limits and reset
2 Low-risk reversible canary Automated rollback and bounded exposure
3 Production intervention with material impact Human authorization and live monitoring
4 High-stakes or irreversible intervention Formal approval, independent review, and domain governance

The causal model’s confidence should never be the only authorization signal.

13.11 LLM security and prompt boundaries

When an LLM orchestrates causal tools:

  • do not let retrieved text redefine intervention permissions;
  • separate descriptive data from executable instructions;
  • use typed schemas and allowlists;
  • require explicit authorization for writes;
  • record tool arguments and returned evidence;
  • prevent the model from editing its own assumption history silently;
  • treat external documents as untrusted evidence with provenance;
  • verify numerical claims from tool outputs rather than generated prose.

A causal agent that can act is also a privileged automation system.

13.12 Production anti-patterns

“Discover a graph, then declare it causal”

A graph learner returns a hypothesis under assumptions, not ground truth.

“Control for every available feature”

This can introduce collider and post-treatment bias.

“Use an LLM to fill missing edges”

Semantic plausibility becomes hidden prior evidence.

“Use the highest CATE as the policy”

Ignores uncertainty, cost, overlap, and interference.

“Validate on held-out rows”

This tests i.i.d. prediction, not intervention transfer.

“One graph for all time”

Mechanisms, instrumentation, and abstractions evolve.

“A confidence interval proves causality”

The interval is conditional on the identifying assumptions and estimator.

13.13 Exercises

  1. Write a causal data contract for one post-treatment metric that is often mistaken for a confounder.
  2. Design three property-based tests for a counterfactual engine.
  3. Define a mechanism-drift alert for a support intervention model.
  4. Propose authorization tiers for an agent that can alter cloud infrastructure.
  5. List five fields that must appear in a graph-version diff.

14. End-to-End Case Study: A Product Feature

Consider a SaaS product introducing an AI setup assistant. The assistant generates configuration suggestions during onboarding. The team asks:

Does enabling the assistant increase 30-day activation and 90-day retention?

14.1 Make the treatment concrete

“Using AI” is not a treatment. Define:

  • treatment: assistant panel enabled at the start of onboarding;
  • version: prompt, model, tool, and UI bundle assistant-v6;
  • exposure: panel shown and available, regardless of whether the user accepts suggestions;
  • population: newly created workspaces eligible for self-serve onboarding;
  • outcomes: activated within 30 days; retained activity during days 61–90.

This defines an intention-to-treat effect. The effect of actually accepting a suggestion is a different, post-treatment question.

14.2 Initial causal graph

flowchart LR
    S["Workspace sophistication"] --> E["Eligible for experiment"]
    S --> A["Activation"]
    S --> R["Retention"]
    E --> T["Assistant enabled"]
    T --> U["Assistant usage"]
    U --> Q["Configuration quality"]
    Q --> A
    A --> R
    T --> F["Onboarding friction"]
    F --> A
    M["Marketing source"] --> S
    M --> R

    classDef conf fill:#FFF4D6,stroke:#D97706,color:#3B2F0B,stroke-width:2px;
    classDef treatment fill:#E8F1FF,stroke:#2563EB,color:#102A43,stroke-width:2px;
    classDef mediator fill:#F2EAFE,stroke:#7C3AED,color:#2E1065,stroke-width:2px;
    classDef outcome fill:#E8F8EF,stroke:#15803D,color:#123524,stroke-width:2px;
    class S,E,M conf;
    class T treatment;
    class U,Q,F mediator;
    class A,R outcome;
Loading

If treatment is randomized among eligible workspaces, the intention-to-treat effect is identified by randomization. Do not control for usage, configuration quality, friction, or activation when estimating the total effect on retention; they are post-treatment mediators.

14.3 Experiment design

Randomize at the workspace level. Before launch:

  • define eligibility from pre-treatment data;
  • validate assignment logging;
  • predefine primary and secondary outcomes;
  • set a minimum detectable effect and stopping policy;
  • check whether users belong to multiple workspaces, creating interference;
  • freeze the assistant version or treat version changes as separate interventions;
  • log failures and response latency, not only successful suggestions.

14.4 Estimands

Intention-to-treat activation effect

$$ E[A(1)-A(0)]. $$

Intention-to-treat retention effect

$$ E[R(1)-R(0)]. $$

Effect among compliers

If assignment affects actual usage imperfectly, assignment may serve as an instrument for usage under exclusion and monotonicity assumptions. But exclusion may fail because merely showing the assistant changes friction even when no suggestion is accepted.

Mediated effects

The team may ask how much of the retention effect flows through activation or configuration quality. This requires stronger mediation assumptions and should not replace the primary total-effect analysis.

14.5 Heterogeneity

Pre-treatment modifiers may include:

  • company size;
  • technical sophistication;
  • use case;
  • integration count;
  • acquisition channel;
  • region and language.

Explore heterogeneity with sample splitting or pre-specified subgroups. Avoid selecting subgroups because they have the most favorable noisy estimate.

A policy question may be:

Enable the assistant only where expected retention gain exceeds inference cost and support burden.

This requires calibrated CATE estimates and cost modeling, not just subgroup p-values.

14.6 Failure analysis with causal structure

Suppose assistant-enabled users submit more support tickets. Possible explanations:

  1. The assistant creates configuration errors.
  2. The assistant increases activation, exposing users to later complexity.
  3. The assistant makes support easier to discover.
  4. A model version change increases both latency and ticketing.

Ticket count is not automatically a harm outcome. Analyze downstream mechanisms and define user-value outcomes.

14.7 Long-term rollout and feedback

After rollout, assignment is no longer random if the team targets likely responders. Future observational data become policy-selected.

Maintain:

  • logged assignment propensities;
  • a randomized exploration slice where acceptable;
  • versioned treatment definitions;
  • holdout or switchback designs for system-wide effects;
  • monitoring for changes in user composition and model behavior.

14.8 Counterfactual product debugging

For a specific churned workspace, the question “Would it have retained without the assistant?” is not answered by the average experiment alone. A unit-level counterfactual needs a model of latent workspace state and outcome coupling. Present such answers as uncertain and distinguish them from experimentally identified population effects.

14.9 Decision memo structure

A defensible launch decision should state:

  • the exact estimand;
  • randomization and implementation checks;
  • point estimate and uncertainty;
  • outcome and subgroup multiplicity;
  • adverse effects and mediation evidence;
  • external-validity limits;
  • expected value under costs;
  • monitoring and rollback thresholds;
  • which questions remain unidentified.

14.10 Exercises

  1. Why is assistant usage a bad control for the total effect of enabling the assistant?
  2. Under what conditions could random assignment be an instrument for actual usage?
  3. Draw an interference graph for consultants who administer several customer workspaces.
  4. Define a long-term policy-learning plan that preserves some overlap.
  5. Give one product metric that might be a collider after launch.

15. End-to-End Case Study: A Causal Deployment World Model

This chapter develops a concrete architecture for a deployment-assurance system. The goal is not to infer a universal graph of all infrastructure. The goal is to answer a bounded family of decisions with evidence.

15.1 Target decisions

Start with queries such as:

  • Will this rollout increase incident risk relative to the current version?
  • Which pre-deployment test would most reduce uncertainty?
  • Is an observed failure more consistent with the application change, dependency degradation, or traffic conditions?
  • Would rollback, failover, or scaling reduce expected recovery time?
  • Which evidence supports the release decision, and which assumptions remain unresolved?

These queries determine the model’s abstraction.

15.2 Typed entity model

Entities:

  • repository and commit;
  • build artifact;
  • deployment;
  • environment;
  • service;
  • dependency;
  • resource;
  • route;
  • credential;
  • test;
  • alert;
  • incident;
  • remediation action.

Relations:

  • artifact produced by commit;
  • deployment installs artifact into environment;
  • service calls dependency;
  • route sends traffic to service;
  • credential authorizes call;
  • test exercises route or mechanism;
  • alert measures state;
  • action changes component or mechanism.

The entity graph is not automatically a causal graph. It defines possible mechanism structure and intervention targets.

15.3 State variables and observation variables

Separate latent or causal state from measurements.

Causal state Possible observations
Dependency availability errors, timeouts, health checks
Resource saturation CPU, queue depth, scheduler events
Schema compatibility migration state, parse errors, version metadata
Credential validity auth failures, rotation events
Rollout fraction deployment controller events, instance versions
Active failure mode combination of traces, logs, symptoms
Traffic regime request mix, arrival rate, geographic distribution

A dashboard metric is an observation, not necessarily a causal variable.

15.4 Mechanisms

A structural model might include:

$$ \begin{aligned} Load_t &amp;:= f_{load}(Traffic_t,Route_t,U^{load}_t) \\ Saturation_t &amp;:= f_{sat}(Load_t,Capacity_t,U^{sat}_t) \\ ErrorRate_t &amp;:= f_{err}(Version_t,Saturation_t,Dependency_t,Schema_t,U^{err}_t) \\ Alert_t &amp;:= f_{alert}(ErrorRate_t,Instrumentation_t,U^{alert}_t) \\ Rollback_t &amp;:= f_{policy}(Alert_t,Operator_t,ReleasePolicy_t,U^{policy}_t) \\ Recovery_{t+1} &amp;:= f_{recover}(Rollback_t,FailureMode_t,Traffic_t,U^{recover}_{t+1}). \end{aligned} $$

The model makes clear that rollback is selected in response to alerts and latent failure severity. A raw comparison of rolled-back and non-rolled-back incidents is confounded.

15.5 Intervention operators

Define operators with mechanism semantics:

operators:
  deploy:
    changes: [version, rollout_fraction]
    preserves: [traffic_generation, dependency_mechanisms]
  rollback:
    variants:
      - restore_previous_artifact
      - route_to_previous_replica_pool
      - disable_feature_flag
  reroute:
    changes: [route_distribution]
    possible_side_effects: [cross_region_latency, downstream_load]
  scale:
    changes: [capacity_assignment]
    lag_model: required
  isolate_dependency:
    changes: [dependency_call_mechanism]
    replacement: mock_or_cached_response
  replay:
    environment: test_only
    couples_background_request_trace: true

“Rollback” is not one intervention if its implementations replace different mechanisms.

15.6 Evidence layers

flowchart TB
    C["Code, config, and topology evidence"] --> H["Hypothesis generator"]
    T["Tests and simulations"] --> H
    O["Production observations"] --> H
    I["Historical interventions"] --> H
    H --> SCM["Versioned causal model and hypothesis set"]
    SCM --> Q["Query engine"]
    Q --> EX["Experiment selector"]
    EX --> SAFE["Risk and authorization gate"]
    SAFE --> ENV["Test, canary, or production environment"]
    ENV --> EV["New evidence with provenance"]
    EV --> SCM

    classDef evidence fill:#E8F1FF,stroke:#2563EB,color:#102A43,stroke-width:2px;
    classDef reasoning fill:#F2EAFE,stroke:#7C3AED,color:#2E1065,stroke-width:2px;
    classDef formal fill:#FFF4D6,stroke:#D97706,color:#3B2F0B,stroke-width:2px;
    classDef safety fill:#FCE8F3,stroke:#BE185D,color:#4A102A,stroke-width:2px;
    classDef result fill:#E8F8EF,stroke:#15803D,color:#123524,stroke-width:2px;
    class C,T,O,I evidence;
    class H reasoning;
    class SCM,Q,EX formal;
    class SAFE safety;
    class ENV,EV result;
Loading

15.7 Hypothesis generation

An LLM can translate diffs and telemetry into candidate mechanisms:

  • H1: new retry policy amplifies database load during partial failures;
  • H2: credential rotation invalidates a subset of long-lived workers;
  • H3: rollout coincides with an unrelated regional dependency event;
  • H4: alert increase is instrumentation change, not error-rate change.

Each hypothesis must produce:

  • a graph or mechanism delta;
  • predicted observations;
  • predicted response to candidate interventions;
  • required assumptions;
  • a falsification plan.

15.8 Identification before prediction

For rollout risk, historical deployments are selected by release policy. The system should model why some changes received more testing, smaller canaries, or off-peak scheduling.

Potential adjustment variables:

  • pre-treatment change complexity;
  • service criticality;
  • forecast traffic;
  • team release policy;
  • known dependency risk;
  • prior incident history.

Bad controls include:

  • bugs found by the selected test suite;
  • alerts triggered during rollout;
  • rollback decision;
  • post-deployment traffic rerouting.

When hidden reviewer judgment affects both precautions and incident risk, point identification may fail. Use sensitivity analysis, partial identification, or randomized policy variation.

15.9 Query-focused experiments

Suppose the release decision depends on whether retries cause saturation. Candidate tests:

  • replay fixed traffic with retries on versus off;
  • isolate database with a controlled latency injector;
  • hold replica count constant;
  • vary only retry budget;
  • compare resource trajectories and errors.

This experiment need not identify every service edge. It targets the decision-relevant mechanism.

15.10 Combining static and dynamic evidence

Static evidence:

  • code path and call graph;
  • configuration dependency;
  • schema compatibility;
  • infrastructure topology;
  • policy and ownership metadata.

Dynamic evidence:

  • traces;
  • metrics over time;
  • intervention outcomes;
  • canary comparisons;
  • fault-injection tests.

Static evidence constrains possible paths. Dynamic evidence estimates whether and how mechanisms activate. Neither alone is sufficient.

15.11 Confidence as a structured object

Instead of one confidence score, return:

conclusion:
  statement: retry policy materially increases saturation risk under partial database latency
  query: incident_risk_under_rollout_v7
identification:
  status: partially_identified
  lower_effect_pp: 3.1
  upper_effect_pp: 11.8
model_support:
  posterior_mass: 0.78
  alternative_hypotheses:
    regional_dependency_event: 0.15
    instrumentation_only: 0.07
assumption_risk:
  highest:
    - replay_environment_preserves_retry_mechanism
    - no_unlogged_manual_release_selection
recommended_test:
  action: 5_percent_canary_with_retry_budget_1
  expected_information_gain: high
  risk: low

The exact numerical framework can vary. The decomposition should not.

15.12 Deployment gate

A gate can use causal evidence without claiming certainty:

  • Pass: expected benefit positive across plausible models and risk below threshold.
  • Canary: decision sensitive to one testable uncertainty; run bounded intervention.
  • Block: credible harmful effect or unsupported high-risk state.
  • Review: identification depends on disputed assumptions or the change is out of domain.

The gate should expose the reason and counterfactual: what additional evidence would change the decision?

15.13 Learning loop

After every deployment:

  1. Record exact intervention and policy propensity.
  2. Compare predicted and observed mechanism-level outcomes.
  3. Update parameter and graph uncertainty.
  4. Detect instrumentation and mechanism drift.
  5. Store failures and near misses, not only successful rollouts.
  6. Re-evaluate past assumptions.
  7. Select future tests that improve decision value.

This creates a causal evidence system rather than a static risk classifier.

15.14 Exercises

  1. Define a causal query for whether a canary reduces incident impact.
  2. Identify two hidden confounders in historical rollback data.
  3. Write mechanism semantics for three different rollback implementations.
  4. Design a query-focused test that distinguishes retry amplification from raw traffic overload.
  5. Define an abstention condition for a deployment world model.

16. The Research Frontier and a Learning Roadmap

The frontier is not one grand causal model. It is the convergence of causal inference, representation learning, pretrained inference, active experimentation, agents, and world models.

16.1 Major research directions as of August 2026

Causal foundation models

CausalFM, Do-PFN, Arrow, CDFM, DAG-FM, partial-graph models, and temporal variants attempt to amortize effect estimation or discovery across datasets R25, R26, R27, R28, R29, R30, R47, R48, R49. The key open problem is reliable generalization under prior mismatch.

Partial identification by default

Foundation Models for Partial Causal Identification treats non-uniqueness as an output rather than forcing a point answer R38. This direction could make neural causal systems more honest about hidden confounding and structural ambiguity.

Finite-sample causal representation learning

Recent theory is moving from “identifiable with infinite data” toward sample complexity, unknown intervention targets, and realistic numbers of environments R32. The gap to nonlinear, high-dimensional real systems remains large.

Interactive causal agents

CausaLab, CausalGame, CausalDS, and related benchmarks evaluate hypothesis formation, experiment design, mechanism recovery, and data analysis rather than isolated causal answers R33, R37, R41. Current agents can solve tasks while holding incorrect mechanisms and often choose weak experiments.

Mechanism-level interventions

Bipartite graphical causal models address settings where the same variable value can be enforced by replacing different equations R39. This is particularly relevant for equilibrium systems and engineered environments.

Adversarial falsification

ACIF frames learning as a game between a causal generator and an experimentalist selecting interventions that expose incorrect structure R40. This connects generative modeling with scientific falsification.

Causal world models

New work is clarifying that world models should be judged by tasks, representations, entities, mechanisms, and intervention transfer, not only generative quality R45, R46.

Unstructured outcomes

Causal Inference with Unstructured Outcomes asks how to define treatment effects when outcomes are text or images and subtraction is meaningless, proposing learned features that expose the strongest causal contrast R51.

Actual causes for model predictions

Work on Halpern–Pearl actual causation is being adapted to neural predictions with structured causal inputs, aiming to avoid explanations that treat dependent features as independent R52.

Causal data infrastructure

A 2026 proposal for a persistent, queryable causal world system argues that agents need an explicit causal layer over heterogeneous data sources R53. The systems problem—lineage, versioning, permissions, and shared causal semantics—is becoming part of the research agenda.

16.2 The hardest open problems

Discovering the right variables

Once variables and interventions are correctly defined, many causal tools are mature. Automatically finding stable, manipulable abstractions in complex data remains difficult.

Identifiability under realistic assumptions

Real systems contain hidden confounding, feedback, selection, measurement error, interference, and nonstationarity simultaneously. Most theory handles a subset.

Bridging simulator and reality

Synthetic SCMs make pretraining and evaluation possible, but causal validity depends on whether real mechanisms lie within the simulator’s support.

Causal model evaluation without full ground truth

Real causal graphs are rarely known. The field needs stronger evaluation through interventions, invariance, decision value, and falsification.

Safe active experimentation

The most informative intervention may be expensive or dangerous. Experiment design must combine epistemic value with operational constraints and authorization.

Counterfactuals for individuals and trajectories

Population effects are easier to identify than unit-level alternative histories. Agents and medicine increasingly demand trajectory counterfactuals, raising strong assumptions and fairness concerns.

Causal abstraction across scales

A model must connect low-level dynamics to high-level concepts without losing intervention meaning. This is central to robotics, biology, and software systems.

Human–agent epistemology

Who owns assumptions? How are disputes represented? When can an agent act? How are prior beliefs separated from data evidence? These are technical interface and governance questions, not only philosophy.

16.3 Research principles worth keeping

  1. Interventions outrank stories. Plausibility generates hypotheses; controlled changes test them.
  2. Identification outranks model capacity. An unidentified target stays unidentified in a larger network.
  3. Mechanisms outrank graph aesthetics. A graph matters when it supports correct intervention predictions.
  4. Ambiguity should be represented. Use equivalence classes, bounds, ensembles, and abstention.
  5. Evaluation should match action. Test unseen interventions and decisions, not only i.i.d. prediction.
  6. Failure data are valuable. Failed actions reveal constraints and mechanisms.
  7. The environment is part of the dataset. Active systems create their own future evidence.
  8. Provenance is part of reasoning. Every edge, assumption, and result needs a source.

16.4 A 12-week learning plan

Weeks 1–2: Questions and graphs

Study Chapters 1–3. Practice chains, forks, colliders, d-separation, potential outcomes, and SCM interventions.

Deliverable: write three causal questions from your domain, each with a graph and estimand.

Weeks 3–4: Identification

Study Chapter 4. Work through backdoor, frontdoor, instruments, mediation, longitudinal confounding, and partial identification.

Deliverable: for each question, state whether it is identified and what additional evidence would help.

Weeks 5–6: Estimation

Study Chapters 5–6. Run the Python lab. Implement standardization, IPW, AIPW, and one CATE estimator.

Deliverable: a reproducible report with overlap, balance, sensitivity, and uncertainty diagnostics.

Weeks 7–8: Discovery and experiments

Study Chapters 7–8. Run PC/FCI or GES on synthetic data. Perturb assumptions. Design query-focused interventions.

Deliverable: compare at least two discovery methods and show how one intervention resolves an ambiguity.

Weeks 9–10: Modern causal AI

Study Chapters 9–12. Read one CRL paper, one causal foundation model paper, one agent benchmark, and one world-model paper.

Deliverable: architecture for an agent whose causal claims are formally grounded.

Weeks 11–12: Production system

Study Chapters 13–15. Create typed query, data, assumption, graph, and intervention schemas.

Deliverable: a small end-to-end causal evidence system for one bounded decision.

16.5 A paper-reading sequence

Foundation

  1. Rubin, Estimating Causal Effects of Treatments in Randomized and Nonrandomized Studies R1.
  2. Pearl, Causal Diagrams for Empirical Research R2.
  3. Peters, Janzing, and Schölkopf, Elements of Causal Inference R4.
  4. Hernán and Robins, Causal Inference: What If R5.

Discovery and invariance

  1. Shimizu et al., LiNGAM R12.
  2. Hoyer et al., additive-noise models R13.
  3. Peters et al., invariant causal prediction R14.
  4. Zheng et al., NOTEARS R11.

Causal machine learning

  1. Chernozhukov et al., double/debiased ML R15.
  2. Künzel et al., meta-learners R17.
  3. Oprescu et al., orthogonal random forests R16.
  4. Applied Causal Inference Powered by ML and AI R9.

Representation and foundation models

  1. Schölkopf et al., Towards Causal Representation Learning R6.
  2. Lee, Jin, and Aragam, finite-sample CRL R32.
  3. Ma et al., CausalFM R25.
  4. Arrow, CDFM, and DAG-FM R27, R28, R29.
  5. Bellot and Dhir, partial causal identification R38.

Agents and world models

  1. CausaLab R33.
  2. Causal Discovery in the Era of Agents R35.
  3. CausalFlip and CauGym R34, R36.
  4. A Unifying Perspective on Causal World Models R45.
  5. Bipartite Graphical Causal Models R39.

16.6 A final mental model

Causal AI is best understood as a loop:

flowchart LR
    OBS["Observe"] --> ABS["Form causal abstractions"]
    ABS --> HYP["Maintain competing hypotheses"]
    HYP --> ID["Identify answerable queries"]
    ID --> EXP["Choose safe, informative interventions"]
    EXP --> ACT["Act"]
    ACT --> EVI["Collect structured evidence"]
    EVI --> FAL["Falsify, update, or bound"]
    FAL --> HYP
    FAL --> DEC["Make a decision with uncertainty"]

    classDef observe fill:#E8F1FF,stroke:#2563EB,color:#102A43,stroke-width:2px;
    classDef reason fill:#F2EAFE,stroke:#7C3AED,color:#2E1065,stroke-width:2px;
    classDef formal fill:#FFF4D6,stroke:#D97706,color:#3B2F0B,stroke-width:2px;
    classDef action fill:#FCE8F3,stroke:#BE185D,color:#4A102A,stroke-width:2px;
    classDef evidence fill:#E8F8EF,stroke:#15803D,color:#123524,stroke-width:2px;
    class OBS,ABS observe;
    class HYP reason;
    class ID,EXP formal;
    class ACT,DEC action;
    class EVI,FAL evidence;
Loading

The goal is not to make an AI system sound causal. The goal is to make its decisions traceable to interventions, assumptions, evidence, and mechanisms that can be challenged.

16.7 Exercises

  1. Choose one frontier direction and state the key identifiability obstacle.
  2. Design a benchmark that distinguishes observational prediction from causal world modeling.
  3. What would a causal foundation model need to know to abstain responsibly?
  4. Define a 12-week project in your own domain using the roadmap.
  5. Write one causal claim that your system should refuse to make from current evidence.

Appendix A — Mathematical Toolkit

This appendix collects the pieces of probability and statistics that recur throughout causal analysis. It is not a substitute for a probability course. It is a compact reference for translating causal assumptions into estimands and estimators.

A.1 Random variables, distributions, and expectations

A random variable $X$ maps outcomes of an underlying experiment to values. Its distribution can be described by a probability mass function, density, cumulative distribution, or samples.

The expectation of a function $g(X)$ is:

$$ E[g(X)] = \int g(x)p(x),dx $$

for a continuous variable, or the corresponding sum for a discrete variable.

Conditional expectation is a function of the conditioning variable:

$$ m(x)=E[Y\mid X=x]. $$

Many causal estimators are built from conditional expectations. Under conditional exchangeability, the causal response under treatment $a$ can be recovered by averaging an outcome regression over the target population:

$$ E[Y(a)] = E_X[E[Y\mid A=a,X]]. $$

The outer expectation matters. A conditional response surface is not yet a population effect until it is averaged over a defined target population.

A.2 Marginalization and the law of total expectation

Marginalization removes variables from a joint distribution:

$$ p(y)=\int p(y,x),dx. $$

The law of total expectation says:

$$ E[Y] = E_X[E[Y\mid X]]. $$

The adjustment formula is a causal use of this identity after the causal assumptions justify replacing an intervention distribution with observational conditionals.

For a valid adjustment set $X$:

$$ p(y\mid do(a)) = \int p(y\mid a,x)p(x),dx. $$

The algebra is simple. The hard step is justifying that $X$ is a valid adjustment set.

A.3 Bayes' rule

Bayes' rule updates beliefs about a hypothesis $H$ after observing evidence $D$:

$$ p(H\mid D)=\frac{p(D\mid H)p(H)}{p(D)}. $$

In Bayesian causal discovery, $H$ may be a graph, an SCM, an intervention target, or a causal effect. A posterior over causal models is:

$$ p(M\mid D)\propto p(D\mid M)p(M). $$

An intervention result $D_a$ updates the posterior as:

$$ p(M\mid D,D_a,a)\propto p(D_a\mid M,a)p(M\mid D). $$

This makes the semantics of active experimentation explicit: choose $a$, observe $D_a$, and reweight models according to how well they predicted the result.

A.4 Variance, covariance, and correlation

Variance measures dispersion:

$$ Var(X)=E[(X-E[X])^2]. $$

Covariance measures linear co-movement:

$$ Cov(X,Y)=E[(X-E[X])(Y-E[Y])]. $$

Correlation standardizes covariance:

$$ Corr(X,Y)=\frac{Cov(X,Y)}{\sqrt{Var(X)Var(Y)}}. $$

None of these quantities is inherently causal. A causal relationship can produce zero correlation, and a strong correlation can arise without direct or indirect causation.

A.5 Regression as a conditional model

A regression function estimates a conditional expectation:

$$ \mu(a,x)=E[Y\mid A=a,X=x]. $$

In a linear model:

$$ E[Y\mid A,X]=\beta_0+\tau A+\beta^TX. $$

The coefficient $\tau$ is not automatically a causal effect. It acquires a causal interpretation only when the model, adjustment set, study design, and target estimand justify that interpretation.

Flexible machine learning can improve estimation of $\mu(a,x)$ while worsening extrapolation or uncertainty if overlap is poor. Prediction quality is useful but not sufficient evidence for causal validity.

A.6 Propensity scores and weights

The propensity score is:

$$ e(x)=P(A=1\mid X=x). $$

For binary treatment, inverse-probability weights are:

$$ w_i=\frac{A_i}{e(X_i)}+\frac{1-A_i}{1-e(X_i)}. $$

These weights create a pseudo-population in which measured pre-treatment covariates are balanced in expectation. Extreme weights signal practical positivity problems and can dominate the estimate.

Stabilized treated and untreated weights may be written as:

$$ w_i^{(1)}=\frac{P(A=1)}{e(X_i)}, \qquad w_i^{(0)}=\frac{P(A=0)}{1-e(X_i)}. $$

Weight trimming changes the effective target population. It should be reported as an estimand decision, not hidden as a numerical cleanup step.

A.7 Orthogonality and nuisance functions

Modern causal machine learning often separates:

  • the low-dimensional causal target $\theta$;
  • high-dimensional nuisance functions such as outcome regressions and propensity scores.

An estimating equation is Neyman-orthogonal when small nuisance-estimation errors have only second-order influence on the target around the truth. Informally:

$$ \left.\frac{\partial}{\partial \eta} E[\psi(W;\theta_0,\eta)]\right|_{\eta=\eta_0}=0. $$

Cross-fitting estimates nuisance models on one fold and evaluates scores on another. This limits overfitting-induced bias and supports valid inference under appropriate rate and regularity conditions.

Orthogonality does not fix confounding, positivity failure, invalid instruments, or bad measurement. It protects estimation from a narrower class of nuisance-model errors.

A.8 Information theory

The Kullback–Leibler divergence between distributions $P$ and $Q$ is:

$$ KL(P|Q)=E_P\left[\log\frac{p(X)}{q(X)}\right]. $$

It is nonnegative but asymmetric. Active causal design often chooses interventions expected to maximize the divergence between a posterior and prior over causal hypotheses, or the disagreement among predicted interventional distributions.

Entropy is:

$$ H(M)=-E[\log p(M)]. $$

Expected information gain can be written as expected posterior entropy reduction:

$$ EIG(a)=H(M\mid D)-E_{Y_a}[H(M\mid D,Y_a,a)]. $$

This objective learns what distinguishes models. It does not automatically prioritize safety, cost, or decision relevance.

A.9 Bootstrap and repeated sampling

The nonparametric bootstrap repeatedly resamples observed units with replacement, recomputes an estimate, and uses the empirical distribution of estimates to approximate sampling uncertainty.

A basic workflow is:

  1. sample $n$ rows with replacement;
  2. refit every nuisance and target model;
  3. recompute the complete causal estimate;
  4. repeat many times;
  5. summarize quantiles or standard errors.

Resampling only the final-stage regression while freezing learned nuisance functions usually understates uncertainty. Clustered, longitudinal, networked, or adaptive data require resampling schemes that preserve the dependence structure.

A.10 Matrix notation for linear SCMs

A linear acyclic SCM can be written as:

$$ X = B^T X + U, $$

where $B_{ij}$ represents the direct effect from $X_i$ to $X_j$. Under an ordering compatible with the DAG, $B$ can be permuted into a triangular matrix. Solving gives:

$$ X=(I-B^T)^{-1}U. $$

A hard intervention on variable $X_j$ replaces its structural equation. In matrix terms, one should not merely condition on $X_j=x$; one removes the incoming mechanism for the intervened variable and solves the modified system.

A.11 Monte Carlo intervention simulation

For a known or sampled SCM, a Monte Carlo estimate of an intervention query follows a simple pattern:

  1. sample exogenous variables $U^{(b)}$;
  2. replace the target mechanism with the intervention;
  3. execute the remaining structural equations in causal order;
  4. record $Y^{(b)}$;
  5. average over draws.

For the average treatment effect:

$$ \widehat{ATE}_{MC}

\frac{1}{B}\sum_{b=1}^B \left(Y^{(b)}{do(A=1)}-Y^{(b)}{do(A=0)}\right). $$

Using the same exogenous draw in both counterfactual worlds can reduce Monte Carlo noise and gives an SCM interpretation to unit-level contrasts.

A.12 Uncertainty has multiple sources

Do not compress all uncertainty into one confidence interval. Distinguish at least:

  • sampling uncertainty: finite observed data;
  • nuisance-estimation uncertainty: fitted conditional models;
  • structural uncertainty: graph or mechanism ambiguity;
  • identification uncertainty: multiple causal values compatible with assumptions;
  • measurement uncertainty: noisy or proxy variables;
  • transport uncertainty: mismatch between source and target populations;
  • intervention uncertainty: incomplete execution, compliance, or off-target effects;
  • prior mismatch: synthetic pretraining distribution differs from the real system.

A production causal system should expose which of these sources are represented and which are omitted.


Appendix B — Causal Project Checklist

Use this before approving a causal estimate or allowing an agent to recommend an intervention.

B.1 Question and estimand

  • The decision that the analysis will support is written down.
  • Treatment or action is defined operationally, including timing and dose.
  • Outcome is defined, including measurement window and censoring.
  • Unit of analysis is explicit.
  • Target population is explicit.
  • Estimand is named: ATE, ATT, CATE, policy value, mediation effect, counterfactual, bound, or another query.
  • Interference assumptions are stated.
  • The intervention is feasible and corresponds to a meaningful real action.

B.2 Time and data-generating process

  • Every variable has a timestamp or a clearly defined temporal role.
  • Pre-treatment covariates are separated from post-treatment variables.
  • Data inclusion and selection mechanisms are documented.
  • Missingness mechanisms have been considered.
  • Measurement definitions and changes over time are versioned.
  • Derived metrics that share raw inputs are identified.
  • Repeated observations, clustering, networks, and spillovers are represented.
  • Treatment assignment or logging policies are preserved where available.

B.3 Causal assumptions

  • A DAG, SCM, potential-outcome assumption set, or equivalent causal model is recorded.
  • Candidate confounders are justified by temporal and domain knowledge.
  • Mediators are not accidentally included in a total-effect adjustment set.
  • Colliders and descendants of treatment are not adjusted for without a specific reason.
  • Hidden confounding is considered rather than silently ruled out.
  • Positivity or overlap is plausible in the target population.
  • Consistency and intervention versions are addressed.
  • Transport assumptions are explicit when generalizing across environments.
  • Every discovered edge has provenance and confidence.

B.4 Identification

  • The identification argument is written before estimator selection.
  • A valid adjustment set or alternative design is shown.
  • Instrument relevance, exclusion, independence, and monotonicity are separately discussed when using IVs.
  • Parallel trends and anticipation assumptions are diagnosed for difference-in-differences.
  • Continuity and manipulation assumptions are diagnosed for regression discontinuity.
  • Longitudinal treatment–confounder feedback is handled with appropriate methods.
  • Partial identification is reported when the query is not point identified.
  • The analysis abstains when no defensible identification strategy exists.

B.5 Estimation

  • At least one transparent baseline estimator is included.
  • Flexible nuisance models are evaluated out of sample.
  • Cross-fitting is used where required.
  • Propensity overlap and weight distributions are plotted.
  • Covariate balance is checked after weighting or matching.
  • Standard errors reflect clustering, repeated measures, or adaptive assignment.
  • Model choices are not selected solely because they produce a desired effect.
  • Heterogeneous effects are evaluated out of sample and adjusted for search.
  • Policy evaluation is separate from policy learning.

B.6 Robustness and falsification

  • Placebo outcomes or negative controls are used where meaningful.
  • Placebo treatments or impossible timing tests are considered.
  • Results are tested across defensible adjustment sets and specifications.
  • Unobserved-confounding sensitivity is quantified.
  • Graph conclusions are stable across resamples and method families.
  • Predictions under held-out interventions are evaluated where possible.
  • Failed or off-target interventions are retained as evidence rather than discarded.
  • Contradictory evidence remains visible.

B.7 Communication and governance

  • The causal claim is expressed in intervention language.
  • Assumptions are presented next to the result, not buried in an appendix.
  • Effect size, uncertainty, and practical relevance are distinguished.
  • The effective target population after exclusions or trimming is reported.
  • Data, graph, code, estimator, and report versions are linked.
  • Human authorization boundaries for experiments are explicit.
  • The system records who or what supplied each causal claim.
  • There is a rollback or containment plan for harmful interventions.
  • Monitoring detects drift in assignment, measurement, mechanisms, and overlap.

Appendix C — Graph Pattern Cards

C.1 Confounder

flowchart LR
    Z["Z: common cause"] --> A["A: treatment"]
    Z --> Y["Y: outcome"]
    A --> Y

    classDef cause fill:#FFF4D6,stroke:#D97706,color:#3B2F0B,stroke-width:2px;
    classDef treatment fill:#F2EAFE,stroke:#7C3AED,color:#2E1065,stroke-width:2px;
    classDef outcome fill:#E8F8EF,stroke:#15803D,color:#123524,stroke-width:2px;
    class Z cause;
    class A treatment;
    class Y outcome;
Loading

Default lesson: adjust for an adequate set of pre-treatment common causes to block noncausal backdoor paths.

Failure mode: the measured $Z$ may be a weak proxy for the real common cause.

C.2 Mediator

flowchart LR
    A["A: treatment"] --> M["M: mediator"] --> Y["Y: outcome"]
    A --> Y

    classDef treatment fill:#F2EAFE,stroke:#7C3AED,color:#2E1065,stroke-width:2px;
    classDef mediator fill:#FFF4D6,stroke:#D97706,color:#3B2F0B,stroke-width:2px;
    classDef outcome fill:#E8F8EF,stroke:#15803D,color:#123524,stroke-width:2px;
    class A treatment;
    class M mediator;
    class Y outcome;
Loading

Default lesson: do not adjust for $M$ when estimating the total effect of $A$ on $Y$.

Failure mode: controlled and natural direct/indirect effects require additional assumptions and carefully defined interventions on the mediator.

C.3 Collider

flowchart LR
    A["A"] --> C["C: collider"] <-- Y["Y"]

    classDef source fill:#E8F1FF,stroke:#2563EB,color:#102A43,stroke-width:2px;
    classDef collider fill:#FCE8F3,stroke:#BE185D,color:#4A102A,stroke-width:2px;
    class A,Y source;
    class C collider;
Loading

Default lesson: the path is blocked until one conditions on $C$ or a descendant of $C$.

Failure mode: dataset inclusion is often a collider. Studying only admitted patients, active users, successful deployments, or approved loans can create associations that do not exist in the source population.

C.4 Instrument

flowchart LR
    Z["Z: instrument"] --> A["A: treatment"] --> Y["Y: outcome"]
    U["U: hidden confounder"] --> A
    U --> Y

    classDef instrument fill:#FFF4D6,stroke:#D97706,color:#3B2F0B,stroke-width:2px;
    classDef treatment fill:#F2EAFE,stroke:#7C3AED,color:#2E1065,stroke-width:2px;
    classDef hidden fill:#FCE8F3,stroke:#BE185D,color:#4A102A,stroke-width:2px,stroke-dasharray:5 5;
    classDef outcome fill:#E8F8EF,stroke:#15803D,color:#123524,stroke-width:2px;
    class Z instrument;
    class A treatment;
    class U hidden;
    class Y outcome;
Loading

Default lesson: an instrument changes treatment, has no relevant path to outcome except through treatment, and is independent of unobserved causes under the design.

Failure mode: relevance is testable; exclusion and independence are generally not fully testable from the observed data.

C.5 Selection

flowchart LR
    A["A: exposure"] --> S["S: selected into data"] <-- Y["Y: outcome"]
    S --> D["Observed dataset"]

    classDef source fill:#E8F1FF,stroke:#2563EB,color:#102A43,stroke-width:2px;
    classDef selected fill:#FCE8F3,stroke:#BE185D,color:#4A102A,stroke-width:2px;
    classDef data fill:#FFF4D6,stroke:#D97706,color:#3B2F0B,stroke-width:2px;
    class A,Y source;
    class S selected;
    class D data;
Loading

Default lesson: conditioning on being observed can induce bias even before any statistical model is fitted.

Failure mode: filtering invalid rows, requiring complete follow-up, or analyzing only users who remained active may implement selection without making it explicit.

C.6 Time-varying confounding affected by prior treatment

flowchart LR
    L0["L0"] --> A0["A0"] --> L1["L1"] --> A1["A1"] --> Y["Y"]
    L0 --> L1
    L0 --> Y
    A0 --> Y
    L1 --> Y

    classDef cov fill:#FFF4D6,stroke:#D97706,color:#3B2F0B,stroke-width:2px;
    classDef treatment fill:#F2EAFE,stroke:#7C3AED,color:#2E1065,stroke-width:2px;
    classDef outcome fill:#E8F8EF,stroke:#15803D,color:#123524,stroke-width:2px;
    class L0,L1 cov;
    class A0,A1 treatment;
    class Y outcome;
Loading

Default lesson: ordinary regression adjustment for $L_1$ can block part of earlier treatment effects while being needed to control confounding of later treatment. G-methods are designed for this structure.

C.7 Feedback and equilibrium

flowchart LR
    X["X"] --> Y["Y"]
    Y --> X
    I1["Intervention on mechanism fX"] -.-> X
    I2["Intervention on mechanism fY"] -.-> Y

    classDef state fill:#E8F1FF,stroke:#2563EB,color:#102A43,stroke-width:2px;
    classDef intervention fill:#FCE8F3,stroke:#BE185D,color:#4A102A,stroke-width:2px;
    class X,Y state;
    class I1,I2 intervention;
Loading

Default lesson: time-unroll the system when dynamics are observed, or use a causal formalism that gives explicit semantics to equilibrium and mechanism replacement.

Failure mode: two actions that force the same measured value can have different downstream effects because they replace different mechanisms.


Appendix D — Glossary

Abduction: Inferring latent disturbances or hidden state for a particular observed case. In an SCM counterfactual, abduction is followed by intervention and prediction.

Adjustment set: Variables conditioned on or standardized over to block noncausal paths between treatment and outcome.

ATE: Average treatment effect, $E[Y(1)-Y(0)]$, for a defined target population.

ATT: Average treatment effect among the treated, $E[Y(1)-Y(0)\mid A=1]$.

Backdoor path: A path from treatment to outcome that begins with an arrow entering treatment. Such paths can transmit confounding association.

CATE: Conditional average treatment effect, $E[Y(1)-Y(0)\mid X=x]$.

Causal discovery: Inferring aspects of causal structure from observational, interventional, temporal, multi-environment, or mixed data under explicit assumptions.

Causal effect: A contrast between outcomes or distributions under different interventions.

Causal foundation model: A pretrained model intended to amortize causal estimation, discovery, or reasoning across datasets sampled from a broad prior over causal tasks.

Causal Markov condition: Each variable is independent of its nondescendants given its direct causes in the causal graph.

Causal representation learning: Learning latent variables and mechanisms with causal semantics from high-dimensional observations.

Causal sufficiency: The assumption that the modeled variables include all common causes relevant to the causal relations under study.

Causal world model: A model of entities, state, mechanisms, actions, and environment dynamics intended to support intervention transfer, counterfactuals, planning, or scientific reasoning.

Collider: A node on a path with two arrowheads pointing into it. Conditioning on a collider or its descendant can open a path.

Conditional exchangeability: $Y(a)\perp A\mid X$. Given $X$, treatment assignment carries no further information about the potential outcome under treatment $a$.

Confounder: A cause of treatment and outcome, or more generally a variable participating in a noncausal open path that must be handled for the target query.

Consistency: When a unit actually receives treatment $a$, its observed outcome equals its potential outcome $Y(a)$, assuming treatment versions are sufficiently well defined.

Counterfactual: A query about what would have happened to a particular unit or world under an intervention different from the factual one.

CPDAG: Completed partially directed acyclic graph representing a Markov equivalence class of DAGs under causal sufficiency.

DAG: Directed acyclic graph.

d-separation: A graphical criterion for determining conditional independences implied by a DAG.

Do-calculus: Rules for transforming interventional distributions using graphical conditions, enabling identification beyond ordinary adjustment.

Doubly robust estimator: In common treatment-effect settings, an estimator consistent when either the treatment model or outcome model is correctly specified, under the remaining identifying assumptions.

Effect modification: Variation in a causal effect across values of baseline variables.

Estimand: The population-level causal quantity the study aims to learn.

Estimator: A procedure computed from finite data to approximate an estimand.

Faithfulness: The assumption that observed conditional independences arise from graph separation rather than exact cancellation of causal effects.

FCI: Fast Causal Inference, a constraint-based discovery family that allows latent confounding and returns a PAG.

Frontdoor adjustment: Identification through a mediator under a specific set of graphical conditions even when treatment and outcome are confounded.

Identifiability: Whether a causal query is uniquely determined by the observed distribution plus stated assumptions.

Instrumental variable: A variable that shifts treatment while satisfying relevance, exclusion, independence, and any additional assumptions required for the target effect.

Interference: One unit's treatment affects another unit's outcome.

Intervention: An external change to a variable assignment, policy, mechanism, distribution, or environment.

Invariant causal prediction: A family of methods that uses stability of a target mechanism across environments to infer causal predictors under assumptions.

LiNGAM: Linear non-Gaussian acyclic model, which exploits non-Gaussian independent noise to identify direction in its basic setting.

Markov equivalence: Multiple DAGs imply the same set of observational conditional independences.

Mediator: A variable on a causal path from treatment to outcome.

Mechanism: A structural assignment or conditional process generating a variable from its causes and exogenous influences.

Nuisance function: A function required for estimation but not itself the primary causal target, such as a propensity score or outcome regression.

Orthogonality: Local insensitivity of an estimating equation to small errors in nuisance functions.

Overlap / positivity: Every treatment relevant to the estimand has positive probability for units in the target covariate support.

PAG: Partial ancestral graph representing causal ambiguity when hidden confounding or selection may be present.

Partial identification: Assumptions determine a set or bounds for a causal quantity rather than one point value.

Potential outcome: The outcome that would be observed for a unit under a specified treatment or intervention.

Prior-data fitted network: A neural model pretrained on synthetic tasks sampled from a prior so that inference for a new dataset occurs in context without ordinary task-specific fitting.

Propensity score: Conditional treatment probability given covariates.

SCM: Structural causal model consisting of variables, structural assignments, and exogenous disturbances.

Selection bias: Bias introduced because inclusion or observation depends on variables related to the exposure and outcome or their causes.

Soft intervention: An intervention that changes a mechanism or its distribution without fully fixing the variable.

Structural equation: An assignment describing how a variable is generated from its parents and exogenous noise.

Target trial: The hypothetical randomized experiment whose protocol an observational analysis attempts to emulate.

Transportability: Conditions under which causal information learned in one population or environment can be transferred to another.

Unobserved confounding: A hidden variable causes both treatment and outcome or otherwise creates an unblocked noncausal path.


Appendix E — Selected References

The references emphasize primary papers, open textbooks, and official project publications. Papers dated 2026 should be treated as current research rather than settled doctrine. The edition date of this book is August 24, 2026.

Foundations and applied inference

R1 Donald B. Rubin. “Estimating Causal Effects of Treatments in Randomized and Nonrandomized Studies.” Journal of Educational Psychology 66(5), 1974. https://doi.org/10.1037/h0037350

R2 Judea Pearl. “Causal Diagrams for Empirical Research.” Biometrika 82(4), 1995. https://doi.org/10.1093/biomet/82.4.669

R3 Judea Pearl. Causality: Models, Reasoning, and Inference, second edition. Cambridge University Press, 2009. https://doi.org/10.1017/CBO9780511803161

R4 Jonas Peters, Dominik Janzing, and Bernhard Schölkopf. Elements of Causal Inference: Foundations and Learning Algorithms. MIT Press, 2017. https://mitpress.mit.edu/9780262037310/elements-of-causal-inference/

R5 Miguel A. Hernán and James M. Robins. Causal Inference: What If. Chapman & Hall/CRC, 2020; continuously maintained online edition. https://www.hsph.harvard.edu/miguel-hernan/causal-inference-book/

R6 Bernhard Schölkopf, Francesco Locatello, Stefan Bauer, Nan Rosemary Ke, Nal Kalchbrenner, Anirudh Goyal, and Yoshua Bengio. “Toward Causal Representation Learning.” Proceedings of the IEEE 109(5), 2021. https://arxiv.org/abs/2102.11107

R7 Guido W. Imbens and Donald B. Rubin. Causal Inference for Statistics, Social, and Biomedical Sciences. Cambridge University Press, 2015. https://doi.org/10.1017/CBO9781139025751

R8 Scott Cunningham. Causal Inference: The Mixtape. Yale University Press, 2021. https://mixtape.scunning.com/

R9 Victor Chernozhukov, Christian Hansen, Nathan Kallus, Martin Spindler, and Vasilis Syrgkanis. Applied Causal Inference Powered by ML and AI. Living online book, 2024–2026. https://arxiv.org/abs/2403.02467

R10 Peter Spirtes, Clark N. Glymour, Richard Scheines, and David Heckerman. Causation, Prediction, and Search, second edition. MIT Press, 2000. https://doi.org/10.7551/mitpress/1754.001.0001

Discovery, estimation, and active learning

R11 Xun Zheng, Bryon Aragam, Pradeep K. Ravikumar, and Eric P. Xing. “DAGs with NO TEARS: Continuous Optimization for Structure Learning.” NeurIPS, 2018. https://arxiv.org/abs/1803.01422

R12 Shohei Shimizu, Patrik O. Hoyer, Aapo Hyvärinen, and Antti Kerminen. “A Linear Non-Gaussian Acyclic Model for Causal Discovery.” Journal of Machine Learning Research 7, 2006. https://www.jmlr.org/papers/v7/shimizu06a.html

R13 Patrik O. Hoyer, Dominik Janzing, Joris M. Mooij, Jonas Peters, and Bernhard Schölkopf. “Nonlinear Causal Discovery with Additive Noise Models.” NeurIPS, 2008. https://proceedings.neurips.cc/paper/2008/hash/f7664060cc52bc6f3d620bcedc94a4b6-Abstract.html

R14 Jonas Peters, Peter Bühlmann, and Nicolai Meinshausen. “Causal Inference by Using Invariant Prediction: Identification and Confidence Intervals.” Journal of the Royal Statistical Society: Series B 78(5), 2016. https://arxiv.org/abs/1501.01332

R15 Victor Chernozhukov, Denis Chetverikov, Mert Demirer, Esther Duflo, Christian Hansen, Whitney Newey, and James Robins. “Double/Debiased Machine Learning for Treatment and Structural Parameters.” The Econometrics Journal 21(1), 2018. https://arxiv.org/abs/1608.00060

R16 Miruna Oprescu, Vasilis Syrgkanis, and Zhiwei Steven Wu. “Orthogonal Random Forest for Causal Inference.” ICML, 2019. https://arxiv.org/abs/1806.03467

R17 Sören R. Künzel, Jasjeet S. Sekhon, Peter J. Bickel, and Bin Yu. “Metalearners for Estimating Heterogeneous Treatment Effects Using Machine Learning.” PNAS 116(10), 2019. https://doi.org/10.1073/pnas.1804597116

R18 Heejung Bang and James M. Robins. “Doubly Robust Estimation in Missing Data and Causal Inference Models.” Biometrics 61(4), 2005. https://doi.org/10.1111/j.1541-0420.2005.00377.x

R19 Christian Toth, Lars Lorch, Christian Knoll, Andreas Krause, Franz Pernkopf, Robert Peharz, and Julius von Kügelgen. “Active Bayesian Causal Inference.” NeurIPS, 2022. https://arxiv.org/abs/2206.02063

R20 Chandler Squires, Sara Magliacane, Kristjan Greenewald, Dmitriy Katz, Murat Kocaoglu, and Karthikeyan Shanmugam. “Active Structure Learning of Causal DAGs via Directed Clique Trees.” NeurIPS, 2020. https://arxiv.org/abs/2011.00641

R21 Yashas Annadani, Panagiotis Tigas, Stefan Bauer, and Adam Foster. “Amortized Active Causal Induction with Deep Reinforcement Learning.” NeurIPS, 2024. https://arxiv.org/abs/2405.16718

Software

R22 Amit Sharma and Emre Kiciman. “DoWhy: An End-to-End Library for Causal Inference.” 2020. https://arxiv.org/abs/2011.04216

R23 Keith Battocchi, Eleanor Dillon, Maggie Hei, Greg Lewis, Paul Oka, Miruna Oprescu, and Vasilis Syrgkanis. “EconML: A Python Package for ML-Based Heterogeneous Treatment Effects Estimation.” 2019. https://github.com/py-why/EconML

R24 Yujia Zheng, Biwei Huang, Wei Chen, Joseph Ramsey, Mingming Gong, Ruichu Cai, Shohei Shimizu, Peter Spirtes, and Kun Zhang. “Causal-learn: Causal Discovery in Python.” Journal of Machine Learning Research 25(60), 2024. https://arxiv.org/abs/2307.16405

Causal representation and foundation models

R25 Yuchen Ma, Dennis Frauen, Emil Javurek, and Stefan Feuerriegel. “Foundation Models for Causal Inference via Prior-Data Fitted Networks.” ICLR 2026. https://arxiv.org/abs/2506.10914

R26 Jake Robertson, Arik Reuter, Siyuan Guo, Noah Hollmann, Frank Hutter, and Bernhard Schölkopf. “Do-PFN: In-Context Learning for Causal Effect Estimation.” NeurIPS 2025. https://arxiv.org/abs/2506.06039

R27 Ryan Thompson, He Zhao, Daniel M. Steinberg, and Edwin V. Bonilla. “Arrow: A Foundation Model for Causal Discovery.” 2026. https://arxiv.org/abs/2605.07204

R28 Jie Qiao, Ruichu Cai, Zijian Li, Weilin Chen, Pengfei Hua, Boyan Xu, Zhengming Chen, Zhifeng Hao, and Peng Cui. “CDFM: Towards a General-Purpose Causal Discovery Foundation Model.” 2026. https://arxiv.org/abs/2607.11508

R29 Yikang Chen, Zhengkang Guan, Haoyuan Qian, Peng Cui, Yi Yang, and Kun Kuang. “DAG-FM: A Foundation Model for Causal Discovery under Heterogeneous Causal Mechanisms.” 2026. https://arxiv.org/abs/2607.11510

R30 Arik Reuter, Anish Dhir, Cristiana Diaconu, Jake Robertson, Ole Ossen, Frank Hutter, Adrian Weller, Mark van der Wilk, and Bernhard Schölkopf. “Use What You Know: Causal Foundation Models with Partial Graphs.” 2026. https://arxiv.org/abs/2602.14972

R31 Guangyi Chen, Yunlong Deng, Peiyuan Zhu, Yan Li, Yifan Sheng, Zijian Li, and Kun Zhang. “CausalVerse: Benchmarking Causal Representation Learning with Configurable High-Fidelity Simulations.” 2025. https://arxiv.org/abs/2510.14049

R32 Inbeom Lee, Tongtong Jin, and Bryon Aragam. “Beyond Identifiability: Learning Causal Representations with Few Environments and Finite Samples.” 2026. https://arxiv.org/abs/2603.25796

LLM agents, interactive discovery, and evaluation

R33 Junlin Yang, Dylan Zhang, Xiangchen Song, Qirun Dai, Xiao Liu, Yuen Chen, Aniket Vashishtha, Jing Shi, Chenhao Tan, and Hao Peng. “CausaLab: A Scalable Environment for Interactive Causal Discovery Toward AI Scientists.” 2026. https://arxiv.org/abs/2605.26029

R34 Junqi Chen, Sirui Chen, and Chaochao Lu. “Can Post-Training Transform LLMs into Causal Reasoners?” 2026. https://arxiv.org/abs/2602.06337

R35 Yujia Zheng, Vishal Verma, Mantej Gill, Haoyue Dai, Peter Spirtes, and Kun Zhang. “Causal Discovery in the Era of Agents.” 2026. https://arxiv.org/abs/2606.23608

R36 Yuzhe Wang, Yaochen Zhu, and Jundong Li. “CausalFlip: A Benchmark for LLM Causal Judgment Beyond Semantic Matching.” 2026. https://arxiv.org/abs/2602.20094

R37 Zhenhao Chen, Yongqiang Chen, Chenxi Liu, Junchi Yu, Xiangchen Song, Zijian Li, Jialin Li, Philip Torr, Bo Han, and Kun Zhang. “CausalGame: Benchmarking Causal Thinking of LLM Agents in Games.” 2026. https://arxiv.org/abs/2607.04293

R38 Alexis Bellot and Anish Dhir. “Foundation Models for Partial Causal Identification.” 2026. https://arxiv.org/abs/2608.20841

R39 Joris M. Mooij. “Causal Reasoning with Bipartite Graphical Causal Models.” UAI 2026. https://arxiv.org/abs/2608.19831

R40 Mojtaba Eslami. “Adversarial Causal Intervention Falsification.” 2026. https://arxiv.org/abs/2608.06427

R41 Andrej Leban and Yuekai Sun. “CausalDS: Benchmarking Causal Reasoning in Data-Science Agents.” 2026. https://arxiv.org/abs/2607.08093

R42 Shaojie Shi, Zhengyu Shi, Lingran Zheng, Xinyu Su, Anna Xie, Bohao Lv, Rui Xu, Zijian Chen, Zhichao Chen, Guolei Liu, Naifu Zhang, Mingjian Dong, Zhuo Quan, Bohao Chen, Teqi Hao, Yuan Qi, Yinghui Xu, and Libo Wu. “InterveneBench: Benchmarking LLMs for Intervention Reasoning and Causal Study Design in Real Social Systems.” 2026. https://arxiv.org/abs/2603.15542

R43 Amartya Roy and Sonali Parbhoo. “Why LLMs Fail at Causal Discovery and How Interventional Agents Escape.” 2026. https://arxiv.org/abs/2605.27567

R44 Dennis Frauen, Marie Brockschmidt, Konstantin Hess, Haorui Ma, Yuchen Ma, Abdurahman Maarouf, Maresa Schröder, Jonas Schweisthal, Yuxin Wang, Athiya Deviyani, Sonali Parbhoo, Rahul G. Krishnan, and Stefan Feuerriegel. “Causal Methods for LLM Development and Evaluation.” 2026. https://arxiv.org/abs/2605.25998

World models and emerging directions

R45 Avinash Kori and Fabrizio Russo. “A Unifying Perspective on Causal World Models: From Observations to Representations to Structure.” 2026. https://arxiv.org/abs/2608.13456

R46 Xinyuan Chen, Haoyu Guo, Shi Guo, Bingqi Jiang, Chunhua Shen, Xing Shen, Tianfan Xue, Yufei Xue, Mulin Yu, Weinan Zhang, Bin Zhao, Bowen Zhou, and Ming Zhou. “A Definition and Roadmap for World Models.” 2026. https://arxiv.org/abs/2607.06401

R47 Dennis Thumm and Ying Chen. “Interventional Time Series Priors for Causal Foundation Models.” 2026. https://arxiv.org/abs/2603.11090

R48 Dennis Thumm, Ruben Wiedemann, and Ying Chen. “Towards Continuous-time Causal Foundation Models.” 2026. https://arxiv.org/abs/2605.28880

R49 Shravan Talupula and Saurabh Sharma. “Temporal Causal Prior-Data Fitted Networks for Panel Data with Learned Reliability Signals.” 2026. https://arxiv.org/abs/2606.20889

R50 Serafim Batzoglou. “ReplaySCM: A Benchmark for Executable Causal Mechanism Induction from Interventions.” 2026. https://arxiv.org/abs/2605.08197

R51 Kevin Christian Wibisono and Yixin Wang. “Causal Inference with Unstructured Outcomes.” 2026. https://arxiv.org/abs/2608.03085

R52 Jannick Strobel, Muqsit Azeem, and Stefan Leue. “Computing Actual Causes for Neural Network Predictions under Structured Causal Inputs.” 2026. https://arxiv.org/abs/2608.03772

R53 Dazhuo Qiu, Yingli Zhou, Amedeo Pachera, Angela Bonifati, and Andrea Mauri. “Toward a Causal Data Management Ecosystem for Decision Making and Agentic AI.” 2026. https://arxiv.org/abs/2608.07214


Closing Note

Causal AI is not a shortcut around experimental science. It is a way to make the logic of intervention, uncertainty, and mechanism explicit enough that machines can participate in analysis without hiding the assumptions that make an answer possible.

The strongest systems will combine:

  • precise causal questions;
  • typed, versioned causal models;
  • formal identification;
  • robust estimation;
  • active but bounded experimentation;
  • learned representations and reusable priors;
  • explicit abstention under ambiguity;
  • evidence that can falsify the system's own beliefs.

The practical standard is not whether a system can draw a graph or use causal vocabulary. It is whether the system can state what would change under an intervention, explain why the claim is identified, quantify what remains unknown, and revise its model when the world disagrees.

Contributors

mohsen1

Issues