Mental model

Interventions and the Do-Operator

A tool for distinguishing between seeing a relationship (correlation) and predicting the effect of making a change (causation).

Discover

A manager notices teams using new software are 20% more productive. To determine if the software *causes* this boost, which question is the most useful to ask?

Select the best question for causal analysis:

Let's explore the right way to think about cause and effect.

Understand

Understand

To find a true cause, you must ask what would happen if you forced a change, not just observe what's already happening. The 'do-operator' is a conceptual tool for thinking about this forced change, separating correlation from causation. Think of it as the difference between noticing that people with gym memberships are healthier (seeing a correlation) and predicting what would happen to public health if you gave everyone a free membership (simulating an intervention).

Ask this: 'Am I just observing this relationship, or am I modeling what would happen if I forced a change?'

Full explanation

Full explanation

The do-operator provides a clear language to distinguish between two different kinds of questions: 'seeing' versus 'doing'. 'Seeing' involves conditioning on an observation—for example, looking at the success rate of patients who happen to take a drug. This is passive observation and can be misleading because other factors might be at play.

'Doing', on the other hand, involves an intervention. This concept simulates forcing a variable to a certain state for the entire population. It mentally 'severs' the normal reasons why that variable might take on a certain value, allowing you to isolate its true downstream effects. It's the mathematical equivalent of running a perfect, randomized controlled trial.

Consider a city government that observes lower crime rates in neighborhoods with more police patrols. Is it the patrols causing the drop, or are more patrols assigned to historically safe areas? An intervention forces us to imagine the outcome if every neighborhood received high patrols, regardless of its history, thereby isolating the patrols' actual causal impact on crime.

A similar logic applies in personal decisions. You might notice your friends who wake up at 5 AM are very successful. Before you set your alarm, causal reasoning encourages you to ask: What would happen if I forced myself to wake up at 5 AM? This separates the potential effect of the early wake-up from the underlying traits (like discipline and ambition) that might cause both the early rising and the success.

Research

Research

The do-operator is the cornerstone of Judea Pearl's Structural Causal Model (SCM) framework, which provides a formal language for expressing and estimating causal effects from data. It enables a 'calculus of doing,' allowing researchers to determine if a causal quantity can be estimated from observational data alone, a process called identification.

  • Pearl (2009) formally defines an intervention, do(X=x), as a surgical operation on a causal graph that removes all incoming arrows to the variable X and fixes its value to x. This represents an external force that sets X, independent of its usual causes. [1]
  • Spirtes, Glymour, and Scheines (2000) established foundational principles, such as the Causal Markov Condition, showing that under specific assumptions about the causal graph, the effects of some interventions are identifiable from non-experimental data. [2]
  • Bareinboim & Pearl (2012) introduced z-identifiability, a framework for estimating causal effects by combining observational data with data from a 'surrogate experiment'—an intervention on a different variable than the one of direct interest. [3]
  • Bareinboim & Pearl (2016) synthesized years of research on 'data fusion,' providing formal criteria for transportability: the ability to generalize causal findings from an experimental study to a different population where only observational data is available. [4]

Limitations

Limitations

The power of the do-operator depends entirely on the correctness of the underlying causal model. If the assumed causal graph is wrong (e.g., it omits a key confounding variable or misrepresents a relationship), the resulting causal estimates will be biased. Furthermore, not all causal effects are 'identifiable' from observational data; in many cases, do-calculus will prove that a randomized experiment is necessary. Finally, the standard framework assumes that one unit's treatment doesn't affect another's outcome (the 'Stable Unit Treatment Value Assumption'), which can be violated in systems with network effects.

Try it

Synthesize

Choose a pattern from the guide, then pick an action to try with it.

Which pattern stands out?

What will you try?

Choose a pattern above to select an action.

Sources

Sources

Try it

Check your understanding

A marketing team observes that customers who open promotional emails buy more. To estimate the *causal* effect of the emails, what question should they be trying to answer?

Show the guide's explanation

Answer: What would the purchase rate be if we *forced* everyone to see the email's content?

The 'do-operator' helps us distinguish observation from intervention. Asking what would happen if we *forced* the action (seeing the email) isolates the causal effect of the email's content from the pre-existing characteristics of people who already choose to open them.

Which of the following scenarios is the best physical example of an intervention to estimate a causal effect?

Show the guide's explanation

Answer: A/B testing a website change, where users are randomly shown either the old or new version.

Randomized experiments like A/B tests are a physical implementation of the do-operator concept. They actively force some participants into the 'treatment' group and others into 'control', breaking any self-selection bias and isolating the change's true effect.

What is the primary danger of confusing 'seeing' (observation) with 'doing' (intervention)?

Show the guide's explanation

Answer: A policy or decision might have no effect, or even a negative effect, because it was based on a spurious correlation.

If a manager promotes a software because the teams who chose it were successful (ignoring that they were already the best teams), a company-wide rollout based on this observation may yield disappointing results because the correlation was not due to a causal link.

Keep exploring

Find another idea for the decision in front of you.

The complete Reframo library is free to read. Explore another guide whenever you are ready.