Governments frequently test a policy in a few sites before committing to it everywhere. The pilot succeeds, the expansion disappoints, and the pattern repeats often enough to be predictable.
Pilot sites are rarely typical
Somebody has to volunteer to run a trial. The places that volunteer usually have spare administrative capacity, engaged staff and leadership interested in the idea.
That is a selected sample by construction. The average site in a national rollout has none of those advantages and did not choose to participate.
Any result from a pilot therefore describes what happens under favourable conditions, which is useful information but not a forecast of the national average.
Attention itself changes the outcome
A pilot is watched. Designers visit, data is collected carefully, problems get escalated quickly and staff know their work is being evaluated.
None of that survives expansion. At scale the program becomes one of many duties, monitored by routine reporting rather than by the people who designed it.
Staffing quality does not replicate
Small programs can be staffed by people chosen for enthusiasm or expertise. A national program has to be staffed by whoever is available in every region.
Where the design depends on skilled judgement — assessing eligibility, tailoring a service, handling difficult cases — this difference is decisive.
Designs that survive scaling tend to be ones that work adequately with ordinary staff rather than ones that work brilliantly with exceptional staff.
Administrative plumbing appears only at scale
A pilot can process applications manually. A national program needs eligibility systems, appeals routes, fraud controls, interfaces with other agencies and a call centre.
Building that machinery takes years and money that the pilot budget never had to contemplate, and delays in it are often reported as the program failing.
Much of what looks like a policy not working is really the supporting infrastructure not yet existing.
What a well-designed trial can still tell you
Pilots remain worth running, but their honest purpose is to find out whether a mechanism can work at all and where it breaks, not to estimate national effects.
Trials that randomise sites, include unenthusiastic participants and run long enough to lose their novelty give far more transferable information than showcase projects.
The cheapest useful question a pilot can answer is which parts of the design are load-bearing and which are decoration.