Mental model
Attrition & Missing Data
When participants drop out of studies or data goes missing, the remaining group may no longer represent the original sample, potentially distorting conclusions about cause and effect.
Discover
A weight loss study starts with 100 people in each group. After 12 weeks, only 60 people in the treatment group and 80 in the control group complete the study. The treatment group shows significantly better results. What's the most important question to ask?
Select the critical diagnostic question
Understanding attrition helps you spot when study results might be misleading.
Understand
Understand
When participants drop out of an experiment or researchers can't collect all the data they planned, the people who remain might be different from those who left. This difference can create a biased picture that doesn't represent the true effect. For example, if a workplace training program loses the least motivated employees before the final evaluation, the results will look better than they really are. Check this: Whenever you see study results, ask what proportion of participants completed the study and whether those who dropped out differed from those who stayed.
Full explanation
Full explanation
First, random assignment initially creates comparable groups at the start of an experiment. Then, as the study progresses, some participants leave, skip measurements, or provide incomplete data. The key question is whether this attrition relates to both the treatment and the outcome. If the people who drop out differ systematically from those who remain—such as when sicker patients leave a medical study while healthier patients stay—the final comparison becomes biased because the groups are no longer comparable.
Understanding the mechanism behind missing data helps diagnose the problem. When data is missing completely at random (like a misplaced survey), the remaining sample still represents the original group. When data is missing based on observed characteristics (like younger participants being harder to track), researchers can sometimes adjust statistically. But when the missingness depends on the unmeasured outcome itself—such as depressed people being less likely to respond to depression surveys—the bias cannot be fixed without strong assumptions.
This pattern appears across many domains. In education research, students who struggle are more likely to transfer schools, potentially making a teaching method appear more effective than it truly is. In clinical trials, patients experiencing side effects may drop out of the treatment group while those doing well remain, exaggerating the benefits. In workplace studies, highly engaged employees participate more fully in surveys, skewing satisfaction measures upward. The validity threat is most severe when dropout rates differ between treatment and control groups, or when the reason for dropping out relates to the outcome being measured.
The most reliable safeguard is tracking all participants regardless of completion (intention-to-treat analysis), comparing the characteristics of dropouts versus completers, and conducting sensitivity analyses to test how different assumptions about missing data would change the conclusions.
Research
Research
Attrition bias threatens validity when the probability of missing outcome data depends on both group assignment and the true outcome value. The Cochrane RoB 2 tool assesses this through three signaling questions: whether outcome data are available for nearly all participants, whether evidence suggests the result is unbiased, and whether missingness depends on the true value [1]. Research on missing data mechanisms distinguishes MCAR (missing completely at random), MAR (missing at random given observed data), and MNAR (missing not at random), where only MCAR andMAR permit valid inference without strong assumptions [2]. Differential attrition between experimental arms often indicates MNAR mechanisms that distort effect estimates [3].
- Higgins (2019): Attrition bias is one of five core domains in the RoB 2 tool; bias occurs when missingness in the outcome depends on both intervention group and the true outcome value, making complete-case analysis unreliable [1].
- Little & Rubin (2020): Missing data are classified as MCAR (probability of missingness unrelated to any data), MAR (missingness depends only on observed variables), or MNAR (missingness depends on unobserved values); modern methods like multiple imputation require MAR for valid inference [2].
- Salthouse (2019): In a longitudinal cognitive study of nearly 5,000 adults, participants over 65 who dropped out had significantly lower baseline cognitive functioning than those who continued, confirming selective attrition that restricts generalizability to higher-functioning individuals [3].
Limitations
Limitations
Missing data classifications rely on untestable assumptions about the unobserved values—researchers cannot definitively distinguish MAR from MNAR using only the observed data. Even when attrition rates appear similar across groups, the underlying reasons may differ systematically. Sensitivity analyses help but cannot fully resolve the fundamental problem. Additionally, intention-to-treat analysis preserves randomization benefits but cannot eliminate bias from outcome-dependent missingness. Multiple imputation methods, while sophisticated, still require the MAR assumption and correctly specified models.
Try it
Synthesize
Choose a pattern from the guide, then pick an action to try with it.
Which pattern stands out?
What will you try?
Choose a pattern above to select an action.
Sources
Sources
- [1] Cochrane Handbook for Systematic Reviews of Interventions, Chapter 8: Assessing risk of bias in a randomized trialJulian PT Higgins, Jelena Savović, Matthew J Page, et al. - 2019
- [2] Statistical Analysis with Missing DataRoderick J. A. Little and Donald B. Rubin - 2020
- [3] Attrition in longitudinal data is primarily selective with respect to level rather than rate of changeTimothy A. Salthouse - 2019
- [4] Flexible Imputation of Missing DataStef van Buuren - 2018
- [5] Multiple Imputation for Nonresponse in SurveysDonald B. Rubin - 2004
Try it
Check your understanding
A workplace wellness program study reports that 70% of participants in the treatment group completed the 6-month survey, compared to 90% in the control group. The treatment group shows significantly higher job satisfaction. What's the most critical validity concern?
Show the guide's explanation
Answer: Differential attrition may have biased the comparison
When treatment and control groups have different dropout rates, the people who remain may systematically differ. If dissatisfied employees were more likely to quit the treatment group (perhaps due to program demands), the remaining group would appear more satisfied than they truly are compared to the control group.
True or False: If a randomized trial reports that 'missing data were handled using last observation carried forward,' you can trust that attrition bias has been properly addressed.
Show the guide's explanation
Answer: False
Last observation carried forward (LOCF) is a problematic method that assumes no change occurs after dropout, which is often unrealistic. Modern guidance recommends multiple imputation or maximum likelihood methods under MAR assumptions, or sensitivity analyses to test MNAR scenarios. The key is understanding the missing data mechanism, not just applying a technical fix.
You're evaluating a study of a smoking cessation program where 40% of participants in the treatment group dropped out versus 25% in the control group. The researchers analyzed only participants who completed the program. What's the first step in assessing whether attrition bias threatens the conclusions?
Show the guide's explanation
Answer: Compare baseline characteristics of dropouts versus completers in each group
The most informative diagnostic is comparing who dropped out from each group. If treatment-group dropouts were heavier smokers or had more quit attempts than those who remained (while control dropouts did not differ systematically), then the attrition is outcome-dependent and likely biasing results. This comparison reveals whether the missing data mechanism is MCAR/MAR versus MNAR.
Keep exploring
Find another idea for the decision in front of you.
The complete Reframo library is free to read. Explore another guide whenever you are ready.