Mental model

Difference-in-Differences

A method for estimating causal effects by comparing changes over time between a treatment group and a control group.

Discover

A city bans smoking in restaurants, and neighboring restaurants across the state border see no change in policy. How would you figure out whether the ban actually reduced restaurant revenue?

Put these steps in the right order

You'll see why the order matters in Step 2.

Understand

Understand

Difference-in-differences compares how things change over time between two groups: one that experiences something new and one that doesn't. By looking at the change in each group rather than just final outcomes, it cancels out factors that were already different between them. For example, to test whether a smoking ban hurt restaurant sales, you'd compare how sales changed in restaurants with the ban versus how they changed in similar restaurants without it—because the ban might not be the only thing affecting sales. Try this: The next time you see a policy change in one place but not another, ask what would have happened without it.

Full explanation

Full explanation

How It Works

First, identify a treatment group that experiences an intervention and a control group that doesn't. Then measure your outcome of interest in both groups before and after the intervention happens. The key is to calculate the change over time in each group separately, then compare those changes. This removes pre-existing differences between groups (because they're captured in the before-period) and removes overall trends affecting everyone (because the control group experiences them too).

The Critical Parallel Trend Assumption

The method only works if, in the absence of treatment, both groups would have followed similar paths over time. This is called the parallel trend assumption. You can't prove it holds, but you can check whether groups had similar trends before the intervention. If they were already moving in different directions, the method will give misleading results.

Real-World Examples

Economists used this method to study minimum wage increases by comparing fast-food employment in New Jersey (which raised its minimum wage) to neighboring Pennsylvania (which didn't). Rather than just comparing employment levels after the increase, they compared how employment changed in each state, finding that the wage hike didn't reduce employment as predicted. Public health researchers have applied the same approach to study everything from Medicaid expansions to vaccination policies, always relying on a comparison group that didn't experience the policy change but likely followed a similar trend otherwise.

Practical Considerations

The approach works best when the treatment happens at a clear, known time and when you have multiple observations before and after to check trends. Be wary if other major events happen around the same time, or if the treatment is rolled out gradually to different groups at different times—both situations can complicate the analysis.

Research

Research

Difference-in-differences (DiD) is a quasi-experimental design that estimates causal effects by comparing changes in outcomes between treatment and control groups over time. The method assumes that without treatment, groups would have parallel trends—a weaker assumption than the exchangeability required for naive comparisons. Card and Krueger's (1994) seminal study of New Jersey's minimum wage increase demonstrated DiD's power: comparing fast-food employment changes in New Jersey versus Pennsylvania found no significant employment loss, challenging conventional economic predictions [1]. The method was formalized in econometrics decades later.

  • Wing, Coady, Jason M. Fletcher, and Riley E. Dunleavey (2023): Recent research shows that staggered treatment adoption (where units receive treatment at different times) creates bias in traditional DiD estimators, because later-treated groups serve as controls for earlier-treated ones after their own treatment begins [2].
  • Goodman-Bacon (2019): Demonstrated that two-way fixed effects DiD with staggered timing produces weighted averages of different 2×2 comparisons across groups, potentially comparing treated units to already-treated controls rather than true untreated units [3].
  • Callaway and Sant'Anna (2021): Proposed alternative estimators that explicitly separate group-time average treatment effects, addressing biases from staggered adoption and heterogeneous effects across groups and time [4].

Limitations

Limitations

The parallel trend assumption cannot be definitively tested—you can only examine pre-trends and argue plausibility. Violation produces biased estimates, and there's no universal solution when groups were already diverging. The method also assumes no spillovers: treatment of one unit shouldn't affect outcomes for control units. With staggered treatment timing, newer research shows traditional estimators can be severely biased unless effects are homogeneous across groups and time. Additionally, DiD cannot handle situations where treatment assignment itself is determined by expected outcomes (e.g., targeting struggling regions for help), as this breaks parallel trends by design.

Try it

Synthesize

Choose a pattern from the guide, then pick an action to try with it.

Which pattern stands out?

What will you try?

Choose a pattern above to select an action.

Sources

Sources

Try it

Check your understanding

A city implements a new paid family leave policy, while a neighboring city does not. Researchers compare employment rates before and after in both cities. What is the SECOND step in calculating the difference-in-differences estimate?

Show the guide's explanation

Answer: Calculate the change in employment for the treatment city

After collecting baseline data from both cities (step 1), the second step is to calculate how much the treatment city's employment changed from before to after the policy. Only after calculating changes for BOTH cities can you compare them (step 4) to isolate the policy effect.

A researcher uses DiD to study whether banning sugary drinks in schools reduces childhood obesity. She compares weight changes in schools with the ban to schools without it. What would most threaten the validity of this comparison?

Show the guide's explanation

Answer: Schools with the ban were already showing declining obesity rates before the policy

This violates the parallel trend assumption—the most critical requirement for DiD. If treatment schools were already improving faster than control schools before the ban, you can't attribute post-ban differences to the policy itself. The other factors (wealth differences, staggered timing, existing programs) are problematic too, but divergent pre-trends fundamentally breaks the core logic of the method.

True or False: Difference-in-differences requires that treatment and control groups start at similar levels of the outcome.

Show the guide's explanation

Answer: False

DiD focuses on changes over time, not absolute levels. Groups can start at very different levels—the method subtracts out those pre-existing differences by comparing changes, not final outcomes. What matters is that they follow similar *trends* before treatment, not that they start at the same point. This makes DiD more flexible than methods that require exact baseline matching.

Keep exploring

Find another idea for the decision in front of you.

The complete Reframo library is free to read. Explore another guide whenever you are ready.