Mental model
Common DAG Mistakes
Avoidable pitfalls when drawing and interpreting causal diagrams that can lead to biased conclusions.
Discover
When drawing a causal diagram to decide which variables to adjust for in your analysis, you need to follow the right steps to avoid introducing bias accidentally.
Which step comes FIRST?
Understanding the right sequence reveals how DAGs prevent bias.
Understand
Understand
A Directed Acyclic Graph (DAG) is a map of your causal assumptions that shows which variables influence which others. The most common mistake is adjusting for variables that look like confounders but actually introduce bias—particularly colliders, which are variables caused by two other variables. When you control for a collider, you create a false association between its causes, leading to collider bias. Another frequent error is controlling for all pre-exposure variables, which can accidentally adjust for colliders. The correct approach is to first identify all variables that are causes of the exposure, outcome, or selection into your study based on subject matter knowledge—then use the DAG to determine which variables block biasing paths without opening new ones. Try this: Before including a variable in your adjustment set, ask whether it could be a common effect of two other causes in your system.
Full explanation
Full explanation
DAG mistakes typically arise from three sources: misunderstanding causal structure, relying on data instead of assumptions, and using oversimplified rules for confounder selection.
The most dangerous mistake is conditioning on colliders. A collider is a variable with two or more arrows pointing into it—it is caused by multiple other variables. When you control for a collider through stratification, matching, or regression adjustment, you open a backdoor path between its causes, creating a spurious association where none exists. This is called M-bias when the collider sits on a path between exposure and outcome via two unmeasured variables. Even if a collider occurs before your exposure in time, adjusting for it can still introduce bias.
A second major error is controlling for all pre-exposure variables—a rule sometimes taught as "adjust for everything measured before treatment." This approach fails because it can inadvertently adjust for colliders. Instead, use the disjunctive cause criterion: control for variables that are causes of the exposure, causes of the outcome, or causes of both—but exclude variables known to be instrumental variables (causes of exposure with no other connection to the outcome), as these can amplify bias from unmeasured confounders.
A third mistake is drawing DAGs based on statistical associations rather than causal assumptions. Arrows in DAGs represent causal relationships that would exist if you intervened, not correlations in your data. Two variables can be statistically independent even when a causal arrow exists between them if effects cancel out in the population. Conversely, associations in data can be misleading guides for causal structure.
Finally, many DAGs omit important nodes like selection into the study or unmeasured variables. Including a selection node clarifies how people entered your analysis and whether conditioning on being included (such as restricting to survivors) introduces bias. Unmeasured variables should be drawn explicitly to make hidden assumptions visible.
Research
Research
Research on DAG pitfalls has identified systematic errors that recur across applied research. Suzuki et al. (2019) document ten common pitfalls, including the failure to distinguish between nodes as random variables versus their realized values, misunderstanding that omitted arrows represent strong assumptions of no effect, and neglecting that DAGs describe confounding in expectation rather than realized confounding from random allocation [1]. Greenland (2003) demonstrated that collider-stratification bias can be as severe as classical confounding, yet is often overlooked because traditional confounder selection criteria (associated with exposure, associated with outcome, not on the causal pathway) do not protect against adjusting for colliders [2]. VanderWeele (2019) formalized the disjunctive cause criterion for confounder selection, showing that controlling for causes of exposure or outcome (excluding instrumental variables) provides protection against both M-bias and insufficient adjustment, while the common-cause criterion is too conservative and the pre-treatment criterion is too liberal [3]. Tennant et al. (2021) conducted a review of studies using DAGs and found widespread variation in reporting practices, with many studies failing to report key information such as unmeasured variables, selection nodes, or minimally sufficient adjustment sets, limiting the transparency and utility of the diagrams [4].
Limitations
Limitations
DAGs themselves have inherent limitations that can be mistaken for user error. DAGs are nonparametric—they indicate the presence or absence of causal effects but not their magnitude, functional form, or whether effects vary across individuals. They also cannot distinguish between confounding in expectation (systematic bias) versus realized confounding (chance imbalances from randomization or sampling). DAGs require complete knowledge of the causal structure, which is rarely available; incorrect assumptions encoded in the DAG will lead to incorrect conclusions about confounding control. Some biases, such as selection bias from conditioning on a common effect, cannot be clearly distinguished from confounding bias even in DAG notation. Additionally, DAGs assume static causal relationships and require separate nodes for each time point when modeling longitudinal processes with feedback or time-varying exposures.
Try it
Synthesize
Choose a pattern from the guide, then pick an action to try with it.
Which pattern stands out?
What will you try?
Choose a pattern above to select an action.
Sources
Sources
- [1] Causal Diagrams: Pitfalls and TipsEtsuji Suzuki, Toshiharu Mitsunhashi, Tetsutaro Tsuda, Eiji Yamamoto - 2019
- [2] Quantifying biases in causal models: classical confounding vs collider-stratification biasSander Greenland - 2003
- [3] Principles of confounder selectionTyler J. VanderWeele - 2019
- [4] How to use directed acyclic graphs: guide for clinical researchersTimothy Feeney, Fernando Pires Hartwig, Neil M. Davies - 2025
- [5] Causal Inference: What IfMiguel A. Hernán, James M. Robins - 2020
Try it
Check your understanding
A researcher studies the effect of childhood trauma on adult depression. They consider adjusting for "social support" because people with trauma may have less support and less support may increase depression. However, both trauma and depression also influence whether someone seeks therapy, and therapy affects social support. What is the risk of adjusting for social support?
Show the guide's explanation
Answer: Social support is likely a mediator and adjusting for it would block part of the effect
Social support appears to be on the causal pathway: trauma affects social support (potentially through therapy-seeking), and social support affects depression. Adjusting for a mediator blocks indirect effects, preventing you from estimating the total effect of trauma on depression. A DAG would reveal this mediator structure and warn against conditioning on social support.
You are studying whether exercise improves heart health. You have data on income, and you know income affects both exercise habits and heart health. You also have data on gym membership, which is influenced by both income and exercise. Should you adjust for gym membership in your analysis?
Show the guide's explanation
Answer: No—gym membership is a collider (affected by both income and exercise) and adjusting for it could introduce bias
Gym membership sits at the intersection of two arrows: from income to gym membership (wealthier people can afford gyms) and from exercise to gym membership (exercisers are more likely to join). This makes it a collider. When you condition on gym membership (by including it in your regression), you open a spurious association between income and exercise, potentially introducing collider bias that distorts the true relationship between exercise and heart health.
When building a DAG for a study, you first identify all variables measured before your exposure. Your colleague suggests this is sufficient to determine your adjustment set. Why is this approach potentially problematic?
Show the guide's explanation
Answer: Some pre-exposure variables may be colliders that introduce bias when adjusted for
The 'adjust for all pre-exposure variables' rule (sometimes called the pre-treatment criterion) fails because some variables occurring before exposure are colliders—common effects of two causes. M-bias is the classic example: an unmeasured cause of exposure and an unmeasured cause of outcome both affect a pre-exposure variable. Adjusting for this collider opens a biasing backdoor path. The disjunctive cause criterion is safer: adjust for causes of exposure or outcome, but not for variables that are merely correlated pre-exposure measurements.
Keep exploring
Find another idea for the decision in front of you.
The complete Reframo library is free to read. Explore another guide whenever you are ready.