Mental model
Correlation vs. Causation
Just because two things happen together doesn't mean one causes the other—this critical distinction helps you avoid being misled by coincidences in data.
Discover
Ice cream sales and drowning deaths both spike in summer. If you wanted to reduce drownings, which action would you tackle first?
Choose the most logical first step
Understanding this pattern helps you spot misleading claims everywhere.
Understand
Understand
Correlation means two things change together, like ice cream sales and drowning deaths both rising in summer. Causation means one thing actually makes the other happen, like a virus causing a fever. Just because two patterns show up together doesn't prove one caused the other—there might be a hidden factor behind both, like hot weather driving both ice cream sales and swimming activity. Ask this: When you see two things linked together, what else might connect them?
Full explanation
Full explanation
How it works
Correlation means two variables change together. When one goes up, the other tends to go up or down. Causation goes deeper: changing one variable directly produces a change in the other. The challenge is that patterns can arise from three sources: direct cause, a hidden third factor (a confounder), or pure coincidence.
Why it matters
Misreading correlation as causation leads to costly mistakes. In business, a company might launch features that correlate with high user engagement but don't actually cause it—wasting resources on decorative changes. In health, observational studies have shown associations that may reflect confounding rather than causation. For example, early studies found coffee drinkers had higher heart disease rates, but later research suggested smoking was a key confounder driving both habits.
When to suspect correlation
Strong warning signs include: the relationship appears in observational data without an experiment, the mechanism is unclear or implausible, and reversing cause and effect also makes sense. For example, successful companies might have fancy offices, but building fancy offices doesn't cause success—the direction of influence matters. Similarly, sleep problems and depression are correlated, but determining causality requires careful study—sleep loss might worsen depression, depression might disrupt sleep, or both might share underlying causes.
How to investigate further
The strongest test is experimentation: if you change one variable and the other reliably shifts, you've found causation. When experiments aren't possible, look for natural experiments where circumstances randomly vary, or triangulate using multiple independent sources of evidence. Correlations are valuable starting points—they flag patterns worth investigating—but treating them as conclusions risks acting on coincidence while missing the true drivers.
Research
Research
Correlation versus causation is a foundational distinction across statistics, epidemiology, and machine learning. Observational data can only reveal associations. Establishing causation requires controlled experimentation or special designs that rule out alternative explanations. Modern causal inference provides formal frameworks for distinguishing genuine effects from spurious patterns, using tools like randomized controlled trials, natural experiments, and graphical causal models to clarify what interventions would actually change outcomes.
Limitations
Limitations
- Correlation doesn't imply causation, but causation almost always produces correlation—meaning correlations are valuable clues, just not conclusions on their own. Some correlations are robust yet non-causal, while others are causal but too noisy to detect reliably. Causal inference methods rely on assumptions that can't always be verified, and real-world interventions often have multiple interacting effects. Not all causal claims require the same standard of evidence—decisions under time pressure may rely on weaker correlations than life-altering medical or policy choices. There's also ongoing debate about how much causal language is appropriate for observational machine learning models, which capture associations but can't guarantee their underlying mechanisms.
Try it
Synthesize
Choose a pattern from the guide, then pick an action to try with it.
Which pattern stands out?
What will you try?
Choose a pattern above to select an action.
Sources
Sources
- [1] Causal Inference in Statistics: An OverviewJudea Pearl - 2009
- [2] Causal Inference: What IfMiguel Hernán and James Robins - 2020
- [3] Correlation does not imply causationStanford Encyclopedia of Philosophy - 2024
- [4] The Book of Why: The New Science of Cause and EffectJudea Pearl and Dana Mackenzie - 2018
- [5] Association vs. CausationKaiser Permanente - 2023
Try it
Check your understanding
A study finds that people who meditate daily report lower stress levels. A meditation app company claims this proves meditation causes reduced stress. What's the strongest counterargument?
Show the guide's explanation
Answer: People with lower stress might choose to meditate more often
This is the reverse causality concern: perhaps feeling less motivated to meditate when stressed, or having time/resources to meditate, drives both the meditation habit and lower stress. The correlation could be real without meditation causing the stress reduction—it might be that low-stress lifestyles make daily meditation easier to maintain. An experiment where people are randomly assigned to meditate or not would help distinguish whether meditation actually causes stress reduction.
Returning to the ice cream and drowning example: what's the most likely hidden factor (confounder) that explains why both increase in summer?
Show the guide's explanation
Answer: Hot weather driving both activities
Hot weather is the confounder: high temperatures increase both ice cream sales (people want cold treats) and swimming activity (people want to cool off), which in turn increases drowning risk. Neither ice cream nor drownings cause each other—they're both effects of the same underlying cause. Recognizing this pattern helps you look for third factors whenever you see surprising correlations.
You're evaluating a job training program. You find that graduates earn higher salaries than non-graduates. Which design would best help establish whether the training actually causes higher earnings?
Show the guide's explanation
Answer: Compare graduates to people who applied but were randomly selected
Random selection creates a fair comparison between otherwise similar people—some get the training, some don't—so any subsequent earnings difference is likely caused by the program itself. The other options have weaknesses: surveys measure perception not outcomes; before/after comparisons can't separate program effects from natural career progression; and adjusting for observed factors can't account for unobserved differences like motivation or career ambition that might drive both who applies and who succeeds.
Keep exploring
Find another idea for the decision in front of you.
The complete Reframo library is free to read. Explore another guide whenever you are ready.