Staged Capacity Pilot¶
Procedure — instantiates Equilibrium-Aware Capacity Intervention Design
A reversible rollout procedure for capacity additions in self-optimizing networks.
No model of selfish behavior is fully trustworthy, so the safest way to learn how a capacity change actually shifts the equilibrium is to ship it small and reversibly and watch. Staged Capacity Pilot is the procedure for exactly that: release the addition to a limited slice of the network or population, hold pre-agreed thresholds on aggregate outcomes, and expand only as each stage clears — with a rollback path kept warm the whole time. Its defining move is controlled exposure with a live exit: at every step the change is small enough that a bad equilibrium is survivable and reversible, so real behavioral evidence is bought before full commitment. It changes no prices and imposes no standing access limit; it governs the quantity and reversibility of exposure — how much of the option exists, for whom, and how quickly it can be pulled.
Example¶
An international airport wants to add a new fast-track security lane. In isolation it speeds up the passengers who use it — but travelers self-optimize, and if everyone converges on the fast lane it could starve the standard lanes and migrate the crush to the shared document-check desk downstream. Rather than open it airport-wide, operations runs a staged pilot. Stage one: the lane is live in Terminal 2 only, during off-peak hours, with a defined rollback (close the lane, revert signage) rehearsed in advance. They hold thresholds on the invariant that matters — overall average wait across all lanes and the 95th-percentile wait — not just fast-lane speed. Stage one clears; stage two extends to peak hours; a threshold on the document-check queue wobbles, so they pause, adjust staffing there, and only then proceed to stage three (all terminals). Because each step was bounded and reversible, the one bad signal cost a pause, not an airport-wide failure.
How it works¶
- Define the exposure stages. Lay out the sequence of increasing rollout — by location, by segment, by time window, by fraction of population — as an explicit scenario ladder from smallest safe slice to full deployment.
- Attach go/hold/rollback thresholds to each stage. Before a stage opens, fix the aggregate outcomes that must hold to advance, the ones that trigger a pause, and the ones that trigger reversal — and rehearse the rollback so it's real, not notional.
- Expose, observe, decide. Run each stage long enough for the self-optimizing population to actually re-route, read the aggregate outcome against the thresholds, then advance, hold, or reverse.
- Keep the exit warm. Preserve the ability to roll back cheaply at every stage; the option to stop is what makes the whole procedure safe under behavioral uncertainty.
Tuning parameters¶
- Stage granularity — how many increments between zero and full rollout. Fine stages buy more evidence and contain harm better but take longer and cost more to run; coarse stages are fast but expose more at each jump.
- Exposure fraction per stage — how much of the network each step reaches. Small fractions are safer but may be too small for the equilibrium effect to show; large fractions reveal it but risk real harm.
- Dwell time — how long a stage runs before the go/hold decision. Longer dwell lets agents fully re-optimize (the effect you care about); shorter dwell moves faster but can advance on a not-yet-settled equilibrium.
- Rollback threshold strictness — how bad an aggregate reading must get to trigger reversal. Strict thresholds reverse early and safely but abort promising changes on noise; loose ones ride out noise but let harm run.
When it helps, and when it misleads¶
Its strength is that it replaces a modeled guess about agent response with observed behavior while keeping the downside bounded — the reversibility itself has real option value, because the freedom to stop after a cheap stage is worth more than a confident all-at-once launch under uncertainty.[n1] It is the natural partner to the pre-launch scenario test: the test predicts, the pilot verifies on the live population.
Its failure mode is that a small pilot can understate the paradox — equilibrium effects often appear only at scale, so a change that's harmless in one terminal or one region can tip badly once fully deployed, and a clean pilot can lull a team into skipping the thresholds on the final expansion. Pilots also leak: users outside the slice hear about the fast lane and behavior contaminates the control. The classic misuse is running the pilot as theater — expanding on schedule regardless of the readings, with a rollback that was never actually rehearsed and can't be executed. The guarding discipline is to keep the thresholds binding at every stage including the last, and to rehearse the reversal so the exit is genuinely available.
How it implements the components¶
marginal_capacity_scenario_set— the staged ladder is a scenario set realized in sequence: no-build, partial builds at increasing exposure, full build, each an evaluated case.staged_rollout_and_reversal_rule— it operationalizes the rollout-and-reversal rule going forward: threshold-gated advance with a live rollback at every stage.
It sets no price and alters no cost function (choice_cost_function, incentive_alignment_control — that is its procedural twin, the congestion pricing or toll rule, which changes what an option costs rather than how much of it exists), and it flags nothing itself (paradox_risk_indicator — the live readings it acts on come from the paradox risk dashboard).
Related¶
- Instantiates: Equilibrium-Aware Capacity Intervention Design — this procedure is the reversible-rollout arm of the toolkit.
- Consumes: Paradox Risk Dashboard supplies the per-stage aggregate readings the go/hold/rollback decisions hang on.
- Sibling mechanisms: Congestion Pricing or Toll Rule · Capacity Closure or Reversal Review · Braess Paradox Scenario Test · Route Access Metering Policy · Paradox Risk Dashboard
Editorial Notes¶
Form Classification¶
Form family: Experiment, Test & Rehearsal
Rationale: Staged Capacity Pilot operates as an active test, trial, simulation, drill, or rehearsal that generates evidence through a deliberate attempt or perturbation because it a reversible rollout procedure for capacity additions in self-optimizing networks.
Independent corroboration: The frozen evidence defines Staged Capacity Pilot as 'A reversible rollout procedure for capacity additions in self-optimizing networks', so its operative form is Experiment, Test & Rehearsal.
Nearest alternative: Protocol, Workflow & Routine — Staged Capacity Pilot includes features of a repeatable ordered procedure or handoff sequence that coordinates action, but its defining operation is an active test, trial, simulation, drill, or rehearsal that generates evidence through a deliberate attempt or perturbation.
Review outcome: Independent reviewer agreement; medium confidence.
Origin Attribution¶
Primary origin: Engineering & Design
Origin pattern: Cross-disciplinary synthesis
Present-day reach: Multi-domain
Rationale: A reversible live trial of incremental capacity is engineering commissioning combined with pilot deployment. GOV.UK beta guidance limits exposure while gathering evidence; systems feedback governs self-optimizing networks.
Related originating lineages:
- Computer Science & Software Engineering — computer_science contributes computer science and software-engineering practice to this mechanism's defining operation—A reversible rollout procedure for capacity additions in self-optimizing networks—without displacing the selected primary historical lineage.
- Innovation & Entrepreneurship — innovation_entrepreneurship contributes innovation management and experimental venture practice to this mechanism's defining operation—A reversible rollout procedure for capacity additions in self-optimizing networks—without displacing the selected primary historical lineage.
- Operations Research — Pilot evidence calibrates capacity.
- Systems Thinking & Cybernetics — Systems thinking, feedback control, and cybernetics supplies a parallel or contributing lineage for the mechanism's defining operation: a reversible rollout procedure for capacity additions in self-optimizing networks.
Review resolution: The blind reviewers disagree on primary lineage (systems_cybernetics versus engineering_design). Authoritative or primary research supports engineering_design as the best historical origin: A reversible live trial of incremental capacity is engineering commissioning combined with pilot deployment. GOV.UK beta guidance limits exposure while gathering evidence; systems feedback governs self-optimizing networks. The cited GOV.UK Service Manual, Limited Beta and Phased Rollout; NASA, Product Verification directly supports the mechanism's defining operation. All independently supported contributing domains are retained without an arbitrary cap. origin_mode=cross_disciplinary_synthesis records lineage, while domain_reach=multi_domain records later applicability separately from provenance.
Encyclopedia synthesis: The exact catalogued form synthesizes established practice rather than reproducing a single standard historical label.
Review outcome: Researched adjudication after independent review; high confidence.
Sources consulted:
Notes¶
[n1] A canary release deploys a change to a small fraction of traffic or users first, watches health metrics, and expands only if they hold — otherwise it rolls back. It is the software-delivery embodiment of staged, reversible exposure, and the reason a cheap early exit is worth more than a confident full launch under uncertainty. ↩