Mental model
Testing & Evaluation in Behavioral Policy
How to test behavioral interventions end-to-end: diagnose the behavior, run fair comparisons, measure what matters, and decide whether to scale.
Discover
Which step should come first when testing a behavioral policy idea?
Quick self-check (fits this process topic).
Next: see the full test sequence.
Understand
Understand
Testing and evaluation in behavioral policy is a step-by-step way to see if a small change truly shifts real-world behavior. The first step is to define the target outcome and the specific behavior you want to change, then test versions fairly and measure what matters. For example, a tax office might compare two reminder letters and judge success by payments within 30 days, not email opens.
Check this: Write your primary outcome and how you'll measure it before drafting any message.
Full explanation
Full explanation
The process is Input → Process → Output. First, clarify the policy goal, define a concrete outcome, and diagnose why the behavior isn’t happening (e.g., friction, timing, misunderstanding).
Then design feasible variants (messages, defaults, reminders) and pretest them for clarity. Keep the change minimal so you can attribute effects to the tweak.
Choose an evaluation design: ideally a randomized test (A/B), or when not possible, a robust quasi-experiment. Set sample size, randomization, ethics, and a short measurement window.
Measure a primary outcome tied to the goal, plus guardrails (cost, complaints) and subgroups for equity. Monitor implementation so every group gets what was planned.
Analyze using pre‑specified rules, look for heterogeneity that matters, and estimate cost per additional desired action. Decide: adopt, adapt, or drop.
Examples: late-taxpaying households receive different letters; hospitals test badge prompts for hand hygiene; utilities compare energy reports to reduce usage. Stronger results come from clear outcomes, good diagnostics, and faithful delivery.
Research
Research
Evidence shows public-sector experiments can be practical and impactful, but average effects are modest and context-dependent. Robust design, clear outcomes, and transparency (pre-registration, reporting) improve reliability and scaling.
- UK Cabinet Office (2012): Outlines a practical, stepwise RCT approach for government and when to use it. [1]
- Hallsworth et al. (2017): Simple tax-letter tweaks increased timely payments in large-scale field trials. [2]
- Allcott (2011): Social norm energy reports reduced household electricity use by a few percent at scale. [3]
- Mertens et al. (2022): Meta-analysis finds nudges have small-to-moderate effects, highlighting context and design quality. [4]
- Gertler et al. (2016): Standard playbook for causal evaluation, threats to validity, and cost-effectiveness. [5]
Limitations
Limitations
- Small average effects; meaningful only when outcomes are high-value or cheap to implement.
- External validity: results may fade or flip when scaled or moved to new contexts.
- Constraints on randomization, ethics, and equity; some behaviors require structural policy, not nudges.
- Measurement pitfalls: using clicks instead of real outcomes; short follow-ups miss persistence.
- Publication bias and hype; pre-registration and full reporting help.
Try it
Synthesize
Choose a pattern from the guide, then pick an action to try with it.
Which pattern stands out?
What will you try?
Choose a pattern above to select an action.
Sources
Sources
- [1] Test, Learn, Adapt: Developing Public Policy with Randomised Controlled TrialsUK Cabinet Office - 2012
- [2] The Behavioralist as Tax Collector: Using Natural Field Experiments to Enhance Tax ComplianceMichael Hallsworth, John A. List, Robert D. Metcalfe, Ivo Vlaev - 2017
- [3] Social Norms and Energy ConservationHunt Allcott - 2011
- [4] The effectiveness of nudging: A meta-analysis of choice architecture interventions across behavioral domainsStefan Mertens, Laura Herberz, Ulf J. J. Hahnel, Tobias Brosch - 2022
- [5] Impact Evaluation in Practice (2nd ed.)Paul J. Gertler, Sebastian Martinez, Patrick Premand, Laura B. Rawlings, Christel M. J. Vermeersch - 2016
Try it
Check your understanding
A city wants to test a new letter to reduce unpaid parking fines. What should they do first?
Show the guide's explanation
Answer: Define the target outcome and diagnose barriers
Testing starts with a clear outcome and understanding why the behavior lags. This informs design, measurement, and whether randomization will credibly answer the policy question.
A clinic's email nudge saw high opens but no change in vaccination completion. What’s the best next step?
Show the guide's explanation
Answer: Analyze completion as the primary outcome and redesign to tackle drop‑off
Behavioral evaluation focuses on the real outcome (vaccination), not intermediate clicks. Diagnose where people drop off and test a stronger lever (e.g., easy scheduling).
Which example best demonstrates evaluation in behavioral policy?
Show the guide's explanation
Answer: Randomly assigning two reminder letters and comparing repayment rates
A randomized comparison isolates the effect of the behavioral change on the target outcome, the core of testing and evaluation in behavioral policy.
Keep exploring
Find another idea for the decision in front of you.
The complete Reframo library is free to read. Explore another guide whenever you are ready.