Claims that a policy worked or failed rest on evaluation methods, and the strength of the claim depends entirely on which method was used.
The fundamental problem
Establishing what would have happened without the policy.
Which cannot be observed, since the same people cannot both receive and not receive an intervention.
Every method is an attempt to construct a credible comparison.
Before and after
Comparing outcomes before and after implementation.
Which is the weakest approach, since anything else changing over the same period is confounded with the policy.
Economic conditions, other policies and long-term trends all move simultaneously.
It is nonetheless the most commonly cited evidence in political argument.
Randomised trials
Assigning eligible participants randomly to receive or not receive an intervention.
Which produces the strongest causal evidence, and it is possible in social policy more often than assumed.
Ethical and political objections exist — withholding a beneficial intervention — and are frequently answerable where resources are limited anyway, since random allocation is arguably fairer than alternatives.
Large-scale trials have been conducted in employment, education, health and welfare policy.
Difference in differences
Comparing change in a treated group against change in a comparable untreated group.
Which controls for shared trends, and it requires the assumption that both groups would have changed similarly absent the policy.
Testing that assumption using pre-policy trends is standard practice and is frequently omitted.
Regression discontinuity
Comparing outcomes just above and just below an eligibility threshold.
Which produces credible comparison, since people either side of a cutoff are otherwise similar.
It answers a narrow question — the effect at the threshold — which may not generalise across the whole eligible population.
Natural experiments
Using variation created by circumstance — policy changes in some jurisdictions and not others, timing differences, arbitrary rules.
Which has produced a substantial body of credible evidence in policy areas where trials are impossible.
The credibility depends entirely on whether the variation was genuinely unrelated to the outcome, which requires argument rather than assertion.
What outcomes are measured
Frequently determined by what data exists rather than by what matters.
Which means administrative outcomes — programme participation, employment records, benefit receipt — dominate over wellbeing, which is harder to measure.
Short-term outcomes are measured more than long-term ones, since evaluation funding rarely extends for years.
Cost effectiveness
An effect that is real may not justify the cost, and comparing interventions requires common units.
Which is why cost per outcome measures are used, and comparing across policy areas requires valuing different outcomes against each other.
That valuation involves judgement that is frequently hidden inside a methodology.
Publication and replication
Evaluations producing null results are published less than positive ones, which distorts the evidence base.
Pre-registration of evaluation plans addresses this partially and has been adopted increasingly.
Replication in different contexts is what establishes whether a finding generalises, and it is funded far less than novel evaluation.
Implementation
A policy that works in trial conditions may not work when scaled.
Which is a documented pattern — effects frequently shrink when programmes move from carefully run pilots to routine delivery.
Implementation fidelity, whether the programme delivered is the one that was evaluated, is a substantial factor.
Heterogeneous effects
Average effects conceal variation, and a programme can help some people and harm others while showing a modest average benefit.
Which means subgroup analysis matters, and it must be pre-specified to be credible since testing many subgroups produces spurious findings.
Unintended consequences
Frequently the most consequential outcomes and the least measured.
Which requires anticipating what else might change and measuring it, and evaluation designs that measure only intended outcomes cannot detect them.
Evidence use
Policy decisions incorporate evidence alongside values, resources and politics.
Which means evidence rarely determines a decision, and expecting it to misunderstands the process.
Data infrastructure
Linked administrative data allows evaluation that surveys cannot support.
Which several countries have built, with governance arrangements addressing privacy.
The availability of such data determines what can be evaluated, and countries without it rely on more limited methods.
Independent evaluation
Evaluation commissioned by the body delivering the programme faces obvious incentive problems.
Which is why independent evaluation and pre-registration matter, and why funding arrangements that separate evaluator from deliverer are preferred.
What good evidence looks like
A credible comparison group, pre-specified outcomes, adequate sample, measurement of unintended effects and independent conduct.
Which is a short checklist that most cited evidence in political argument fails.
Where to find evidence
Evidence clearinghouses summarising what has been tested exist for several policy areas.
Which rate evidence quality and are freely accessible, and they are used far less than they could be.
Systematic reviews synthesising multiple studies are more reliable than any individual evaluation.
Which is what a careful reader should look for before accepting a claim that something worked.
Systematic review databases are searchable and free.