Mental model
The G-Formula
A mathematical formula that identifies causal effects from observational data by adjusting for confounders, enabling estimation of what would happen under different treatment scenarios.
Discover
To estimate the true effect of a medical treatment from observational data, you need to follow a specific sequence of steps. Which step comes first?
Understanding causal identification
Understanding this sequence unlocks causal insights from everyday data.
Understand
Understand
Imagine you want to know if a new diet causes weight loss, but people who choose the diet are also more likely to exercise. The G-Formula builds a model to predict what would happen to each person under both scenarios (diet vs. no diet), then compares these predictions to estimate the diet's actual effect. Without accounting for confounding factors like exercise, any comparison might confuse correlation with causation. Try this: When you read about a study claiming X causes Y, ask yourself what other factors might explain both.
Full explanation
Full explanation
The G-Formula (also called G-computation) works through a structured three-step process that transforms observational data into causal estimates. First, you identify and measure confounders—variables that affect both treatment assignment and outcomes. These might include age, income, or prior health status in a medical study. Second, you build an outcome model that predicts what would happen to each person under different treatment scenarios, holding confounders constant. Third, you compare these predicted outcomes across the population to estimate the causal effect. This approach, pioneered by James Robins in the 1980s, is mathematically equivalent to the backdoor adjustment formula from causal graph theory.
In healthcare, the G-Formula has been used to estimate how a specific drug affects patient survival when randomization isn't ethical. Researchers might model survival probability as a function of both treatment and patient characteristics, then predict what would have happened if everyone had received the drug versus no one. The difference reveals the treatment's true effect, accounting for the fact that healthier patients might have been more likely to receive it in the first place.
In education policy, analysts use the G-Formula to evaluate programs like after-school tutoring. They model test scores based on participation, prior achievement, and demographic factors. By predicting what each student's score would have been with and without the program, then averaging these predictions, they can estimate the program's causal impact rather than just observing that participants tend to have higher scores.
The power of the G-Formula comes from its flexibility—it handles both binary treatments (take the drug or not) and continuous treatments (dosage levels), and it extends to time-varying scenarios where treatments and confounders change over time. However, it requires strong assumptions: you must have measured all relevant confounders, your outcome model must be correctly specified, and there must be sufficient overlap in characteristics between treated and untreated groups. When these assumptions hold, the G-Formula transforms correlation into causation.
Research
Research
The G-Formula provides a mathematical identification strategy for causal effects in the potential outcomes framework. Under the assumptions of consistency, positivity, and conditional exchangeability (no unmeasured confounding given measured covariates L), the mean potential outcome under treatment A=a is identified by summing conditional expectations across the covariate distribution: E[Y^a] = Σ_l E[Y | A=a, L=l] × P(L=l). This expression, known as the nonparametric identification formula or adjustment formula, shows that causal effects are identified when we can block all backdoor paths between treatment and outcome [1].
The G-Formula has multiple algebraically equivalent representations: the non-iterated conditional expectation form (standardization), the iterated conditional expectation form (nested prediction), and the inverse probability weighted form [2]. All three identify the same quantity but may differ in finite samples due to model misspecification. For time-varying treatments with time-varying confounders, the G-Formula extends naturally through sequential modeling at each time point, addressing limitations of simpler methods that cannot handle confounder-treatment feedback [3].
- Hernán and Robins (2020): Provide the authoritative textbook treatment of G-computation, showing how the formula identifies causal effects under the core assumptions and demonstrating its equivalence to standardization in epidemiology [1].
- Pearl (2009): Establishes the connection between the G-Formula and the do-calculus, proving that the backdoor adjustment criterion provides sufficient conditions for causal identification via this formula [2].
- Robins (1986): Introduces the original G-computation algorithm for estimating causal effects from complex longitudinal data with time-varying treatments and confounders [3].
Limitations
Limitations
The G-Formula requires the untestable assumption of no unmeasured confounding—any omitted variable that affects both treatment and outcome will bias estimates. Model misspecification is equally critical: if the outcome model is wrong (incorrect functional form, missing interactions, or omitted variables), causal estimates will be biased even when identification assumptions hold. Positivity violations occur when certain covariate combinations never receive a specific treatment, making extrapolation necessary and estimates unstable. The formula also struggles with rare outcomes and high-dimensional covariate spaces, where accurate modeling becomes difficult. Computational complexity increases sharply for time-varying treatments, requiring careful specification of sequential models. Alternative approaches like inverse probability weighting and targeted maximum likelihood estimation can offer robustness to some forms of model misspecification through double-robustness properties.
Try it
Synthesize
Choose a pattern from the guide, then pick an action to try with it.
Which pattern stands out?
What will you try?
Choose a pattern above to select an action.
Sources
Sources
- [1] Causal Inference: What IfMiguel A. Hernán and James M. Robins - 2020
- [2] Causality: Models, Reasoning, and InferenceJudea Pearl - 2009
- [3] A New Approach to Causal Inference in Mortality Studies with Sustained Exposure: Application to Control of the Healthy Worker Survivor EffectJames M. Robins - 1986
- [4] Stanford Encyclopedia of Philosophy: Causal Models - Do-CalculusStanford University - 2024
- [5] The Many Faces of the G-FormulaChristopher B. Boyer - 2024
Try it
Check your understanding
A researcher studies the effect of a training program on employee productivity. She uses the G-Formula, first identifying confounders like prior experience and department, then modeling productivity based on these factors and program participation. What is the next step in her analysis?
Show the guide's explanation
Answer: Compare predicted outcomes under treatment versus no treatment
After modeling outcomes conditional on treatment and confounders, the G-Formula's final step is comparing predicted outcomes across treatment scenarios. For each person, the model predicts what would happen with and without the program; averaging these predictions and comparing them reveals the causal effect. Collecting more data, running an RCT, or removing outliers aren't part of the G-Formula sequence.
In a study of how remote work affects employee well-being, researchers apply the G-Formula but discover that employees with children are never assigned to in-office work. Which core assumption is violated?
Show the guide's explanation
Answer: Positivity
The positivity assumption requires that every group has a non-zero probability of receiving each treatment level. If employees with children never work in-office, there's zero probability of that treatment for them, violating positivity. This makes causal identification impossible for that subgroup without strong extrapolation assumptions. Consistency refers to well-defined treatments, exchangeability to no unmeasured confounding, and SUTVA to no interference between units.
A city evaluates a new crime prevention program by comparing neighborhoods that did and didn't adopt it. Analysts use the G-Formula to adjust for poverty rates, population density, and prior crime levels. True or False: If they've measured all confounders, the G-Formula will give the exact causal effect regardless of their modeling choices.
Show the guide's explanation
Answer: False
Even with all confounders measured, the G-Formula requires a correctly specified outcome model. If analysts choose the wrong functional form, miss interactions, or omit relevant covariates from the model itself, estimates will be biased. The G-Formula identifies the causal effect mathematically, but practical estimation depends on modeling accuracy. This is why sensitivity analysis and robust modeling practices are essential in real applications.
Keep exploring
Find another idea for the decision in front of you.
The complete Reframo library is free to read. Explore another guide whenever you are ready.