Pilot-to-Scale Translation¶
Translation procedure — instantiates Scale-Bridging Translation
Adapts a live pilot's findings to full deployment by separating the pilot conditions that were essential from those that were accidental, then re-basing the result against ordinary target-scale conditions.
Pilot-to-Scale Translation takes a result that already worked in a small live setting and rebuilds it for full deployment. Unlike a lab result, the pilot happened in the messy real world — so the danger is not suppressed variance but scale itself: broader operational burden, heterogeneous local capacity, new governance and incentives, and a monitoring load the pilot never carried. Its defining move is sorting the pilot's conditions into accidental versus essential — which parts of the win came from a motivated site or a hand-picked team, and which from the intervention itself — and then re-basing the expected effect against the ordinary conditions the full rollout will actually meet.
Example¶
A benefits agency piloted a redesigned application-processing workflow in one district and cut processing time by roughly 40%. Before a national rollout, Pilot-to-Scale Translation does the sorting. Accidental: the pilot district had an unusually experienced caseworker team and a site lead who personally shepherded every case — the kind of extra attention that inflates a pilot. Essential: the redesigned form and a parallel-review step that removed a genuine bottleneck. The assumption log then records what must hold nationally — adequate staffing ratios, IT capacity, and no per-case hand-holding — flagging the staffing assumption as the one most likely to break. Because national scale adds heterogeneous district capacity, union work rules, and a compliance-monitoring burden the single site never faced, the team does not extrapolate the 40% directly; instead they run a staged probe in a handful of deliberately ordinary districts.
The probe re-bases the number: in average districts the gain is closer to 15% — still worth the rollout, but the business case is rebuilt on that figure, and targeted training is added to recover part of the gap the star team had quietly been covering.
How it works¶
What sets this procedure apart from the controlled-setting bridge is that it interrogates the pilot's representativeness, not its cleanliness:
- Sort conditions into accidental and essential. For every ingredient of the pilot's success, ask whether it will be present everywhere at scale or was a lucky feature of the pilot site.
- Log the load-bearing assumptions. Record which pilot conditions must persist nationally, marking those that must be redesigned, retested, or bounded before transfer.
- Probe ordinary sites before committing. Stage the deployment through deliberately average units, so the effect is re-based against typical capacity rather than the pilot's best-case one.
Tuning parameters¶
- Ordinariness of probe sites — how deliberately average the staged sites are. The blander the site, the more honest the re-based number and the less flattering the story.
- Essential/accidental threshold — how much doubt about a condition's transferability is enough to treat it as accidental and design around it.
- Capacity-assumption conservatism — whether staffing, IT, and governance are assumed at pilot quality or at the target scale's realistic floor.
- Staging depth — how many waves of probe precede full commitment; more waves buy evidence but delay the benefit.
When it helps, and when it misleads¶
Its strength is defusing the star-site trap — the pilot that succeeded partly because it was special, watched, and staffed with volunteers, and whose number cannot be multiplied out to a network of ordinary units.
Its central failure mode is a self-selected pilot: when the best-prepared district volunteered, the result is biased upward before translation even begins, and enthusiasm effects can flatter a small trial in ways that fade at scale.[1] The classic misuse is linear extrapolation — quoting the pilot's headline number as the rollout's expected result. The discipline is to probe deliberately ordinary sites, re-base the effect on what they show, and treat any condition you cannot guarantee everywhere as accidental until proven otherwise.
How it implements the components¶
scale_gap— names what scale adds: broader operational burden, heterogeneous capacity, new governance and incentives, and a monitoring load absent from the pilot.assumption_log— records which pilot conditions must persist at full scale and which must be redesigned or bounded first.target_scale_probe— the staged deployment through ordinary sites that re-bases the pilot effect before full commitment.
It maintains a working assumption log but not the standing, owned ledger — that is the Scale Assumption Register — and its staged probe is a re-basing check, not a representativeness sweep across defined strata, which is Stratified Target-Scale Rollout.
Related¶
- Instantiates: Scale-Bridging Translation — the small-real to big-real bridge for a live pilot.
- Consumes: Scale Assumption Register can hold and own the assumptions this translation surfaces.
- Sibling mechanisms: Lab-to-Field Translation · Stratified Target-Scale Rollout · Scale Assumption Register · Construct Mapping Table · Micro-to-Macro Model Translation · Macro-to-Micro Operational Translation · Individual-to-Population Policy Translation · Ecological Scale Translation · Multi-Level Model Check · Team-to-Organization Process Translation
References¶
[1] The Hawthorne effect — participants change their behavior because they are being observed or singled out — is one reason a closely-watched pilot can overperform relative to the same intervention rolled out routinely at scale. ↩