Mental model
Unit of Randomization & Clustering
Understand why choosing what to randomize—individuals or groups—is a critical decision that can make or break an experiment.
Discover
You're testing a new teaching method in ten schools. To ensure your results are valid, what's the first critical decision you must make about how you run the experiment?
Select the most crucial first step:
Let's explore why this choice is so important.
Understand
Understand
The first decision is choosing your unit of randomization—the fundamental 'thing' you randomly assign to either a treatment or control group. This could be an individual person, a classroom, a clinic, or even a whole village. When you randomize groups, like classrooms, the individuals within them are often more similar to each other than to people in other groups, a phenomenon called clustering. For example, when testing a new workplace wellness program, randomizing by office is often better than randomizing by employee, as colleagues in the same office influence each other.
Ask this: In this experiment, are the participants influencing each other?
Full explanation
Full explanation
In any experiment, the process begins by defining the unit of randomization. This choice determines whether you are running a simple randomized controlled trial (RCT) or a more complex cluster randomized trial (cRCT).
If the treatment can 'spill over' from one person to another, randomizing individuals is a bad idea. For example, if you give a new communication training to half the employees in an office, they will likely use their new skills on the other half, contaminating your control group. In this case, the office itself should be the unit of randomization; one office gets the training, another doesn't.
This choice has a major statistical consequence called clustering. People within a cluster (an office, a school, a family) tend to be more alike in ways that affect the outcome. Students in the same classroom share a teacher and social environment; villagers share a local economy and culture. This non-independence violates a key assumption of many simple statistical tests.
Ignoring clustering is a common mistake that can lead to a false sense of precision. You might think you have a sample of 1,000 students, but if they are in only 20 classrooms, your 'effective' sample size is much smaller. This can cause you to conclude a program is effective when the results were just due to chance—a false positive that wastes resources on ineffective interventions.
Research
Research
The choice of randomization unit is a foundational design decision that dictates the appropriate methods for sample size calculation and statistical analysis. When the unit is a group (a cluster), the correlation among individuals within that group must be accounted for. This correlation is measured by the Intra-Class Correlation Coefficient (ICC), a value from 0 to 1 indicating how much of the total variation in the outcome is due to variation between clusters.
- A higher ICC means stronger clustering, which inflates the variance of the treatment effect estimate. This requires a larger sample size of clusters to achieve the same statistical power as an individually randomized trial [1].
- Hayes & Moulton (2017): Failing to account for clustering in analysis can dramatically increase the Type I error rate (false positives), leading researchers to incorrectly claim an intervention is effective when it is not [2].
- The World Bank's DIME Wiki provides practical guidance, noting that cluster randomization is often necessary for logistical feasibility or to assess group-level effects, despite the statistical complexities it introduces [3].
Limitations
Limitations
While statistically powerful, individual randomization is often impractical or unethical. You can't assign one member of a household to a clean water program and not another. Furthermore, choosing large clusters (like cities or districts) drastically reduces the number of units you can afford to include, which can make it impossible to detect a true effect (low statistical power). Finally, accurately estimating the ICC before a study to calculate the required sample size can be very difficult, often relying on data from previous, potentially dissimilar studies.
Try it
Synthesize
Choose a pattern from the guide, then pick an action to try with it.
Which pattern stands out?
What will you try?
Choose a pattern above to select an action.
Sources
Sources
- [1] Design and analysis of cluster randomization trials in health researchDonner, A., & Klar, N. - 2000
- [2] Cluster Randomised TrialsHayes, R. J., & Moulton, L. H. - 2017
- [3] Cluster Randomized Control TrialsThe World Bank (DIME Wiki)
Try it
Check your understanding
A health organization wants to test a new diabetes management app. They recruit patients from 10 different clinics. Why might they choose to randomize the *clinics* instead of the individual *patients*?
Show the guide's explanation
Answer: To prevent patients in the same clinic from sharing information about the app, which would contaminate the control group.
This is the primary reason for choosing cluster randomization in this context. If treated and untreated patients interact, the effect of the treatment can 'spill over,' making it impossible to measure the true impact.
A researcher randomizes 20 classrooms to a new math program and finds a positive result. They analyze the data as 500 individual students. This approach is flawed because:
Show the guide's explanation
Answer: It ignores that students within the same class are more similar (clustered), which can create a false positive.
The key error is treating non-independent observations as independent. The shared classroom environment creates clustering, and failing to account for it can make random noise look like a real effect.
Which of the following experiments is LEAST likely to have a problem with clustering?
Show the guide's explanation
Answer: Randomly showing two different website layouts to individual, anonymous visitors.
Anonymous website visitors are typically considered independent. They don't interact or share a common environment in the same way that farm plots (shared soil), families (shared genetics/environment), or households in a village (shared community) do.
Keep exploring
Find another idea for the decision in front of you.
The complete Reframo library is free to read. Explore another guide whenever you are ready.