Skip to content

Pilot Project

Trial implementation — instantiates Hidden-Type Screening

Deploys a candidate solution, vendor, or approach at bounded scale under real conditions to reveal how it actually performs before committing to full rollout.

A Pilot Project screens by running the thing for real, but small. When the hidden attribute is how a system, vendor, or approach behaves in live conditions — something no demo, reference, or specification reliably reveals — the pilot creates the missing evidence by deploying it at bounded scale and watching. What makes it distinctive among the screening mechanisms is that it is a staged real-world trial of an artifact or approach rather than an inspection of a person: it accepts a limited commitment precisely in order to observe performance the candidate cannot fake in a slide deck, then uses what it learns to decide whether to expand, adjust, or abandon.

Example

A city wants a new pothole-reporting platform and has three vendors, each with a polished demo and glowing references. Rather than sign a citywide contract on promises, the public-works department pilots the leading candidate in a single district for one quarter. Now the hidden attributes surface: the app's uptime during a storm-driven spike in reports, how many duplicate tickets it generates, whether field crews actually use it or route around it, and how the vendor responds when something breaks at 2 a.m. The department instruments the pilot with pre-agreed metrics — resolution time, crew adoption, duplicate rate — and reviews them against thresholds set before launch.

The pilot ends with evidence a procurement document never could have produced: the platform handles routine load well but degrades under surge, and the vendor's support is slow. That finding routes the decision — expand with a contractual load guarantee, or reopen the search — on observed behavior rather than sales narrative.

How it works

  • Deploy small but real. Run the candidate in a genuine slice of the operating environment, not a sandbox, so the evidence reflects real conditions.
  • Instrument before you start. Fix the metrics and success thresholds in advance, so the pilot tests a hypothesis instead of generating a story to fit the decision.
  • Monitor through the trial. Track performance continuously and watch for the failure modes a demo hides — surge behavior, edge cases, human workarounds.
  • Route on the evidence. Use the results to expand, renegotiate, adjust, or walk away, treating the pilot as one stage of a graduated commitment.

Tuning parameters

  • Scope of the slice — how large and representative the pilot deployment is. Broader is more predictive of full rollout but costs more and raises the stakes of a bad candidate.
  • Duration — how long it runs. Longer catches slow-emerging failures and seasonal load, but delays the decision and lets sunk commitment build.
  • Success criteria strictness — how demanding the pre-set thresholds are. Loose criteria pass almost anything; overly strict ones reject a candidate that would have improved past a fixable early stumble.
  • Representativeness — whether the pilot conditions match rollout conditions. A pilot on the easiest district or a hand-picked team flatters the candidate and misleads the go/no-go.

When it helps, and when it misleads

A pilot is the right screen when the decisive attribute is real-world behavior under genuine conditions and the cost of a small trial is far below the cost of a wrong full commitment — it manufactures exactly the evidence claims and references cannot. Its central distortion is that a pilot is watched: the extra attention, the vendor's best team, and the volunteers who opted in all make pilot performance overstate rollout performance.[1] Pilots run on the friendliest slice, or without pre-set criteria, become theater — evidence assembled to ratify a rollout already decided. And the pilot itself builds sunk commitment that pressures a "yes" regardless of results. The disciplines are to pilot on representative conditions, pre-register the success thresholds, and preserve a real option to stop — the value of a pilot is the freedom to not proceed, which evaporates the moment expansion is treated as foregone.

How it implements the components

  • staged_commitment_path — it is a graduated commitment: a small, reversible stage whose outcome gates the larger one.
  • elicitation_channel — the bounded real deployment is the designed interaction that produces performance evidence otherwise unobservable.
  • monitoring_and_recalibration_loop — continuous in-trial measurement against pre-set metrics is how the pilot reads the candidate and updates the decision.

It does not screen an individual person's conduct over a trial of employment (that is Probationary Period), nor set the accept/reject threshold as a standing rule for a whole pool (that is Risk Scoring Model and Underwriting Assessment).

  • Instantiates: Hidden-Type Screening — the bounded-real-world-trial variant, revealing type through limited exposure.
  • Sibling mechanisms: Probationary Period · Work Sample or Audition · Diagnostic Test · Risk Scoring Model · Underwriting Assessment · Reference Check · Background Check · Credential Verification · Structured Application · Structured Interview · Self-Selection Menu · Challenge or Proof-of-Work

Notes

A pilot and a Probationary Period share the staged-commitment logic — reveal type through limited exposure, then decide — but differ in their subject: the pilot trials a thing or approach, the probation trials a person. Their tuning and fairness concerns diverge sharply as a result, which is why they are separate mechanisms rather than one.

References

[1] The Hawthorne effect — subjects change their behavior when they know they are being observed. A pilot is inherently observed, so its results tend to overstate steady-state rollout performance; representative conditions and honest thresholds are the correction.