Skip to content

Pilot Bundle Comparison

Evaluation method — instantiates Synergistic Combination Design

Tests the proposed combination against isolated elements or alternative bundles before committing to scale.

A Pilot Bundle Comparison is the experiment that decides whether a proposed combination actually reinforces before anyone scales it. Its defining move is running the combination against arms that isolate its parts — each element alone, the additive baseline, and alternative bundles — so the joint effect can be separated from what the elements would have produced independently. Everything else in the archetype assumes synergy exists; this mechanism is where the claim is put on trial. It generates interaction evidence by direct comparison rather than assertion, measures the combined effect against the isolated and additive baselines, and detects antagonism empirically — the case where the combination underperforms its own parts. It is a test, not a treatment: it produces the evidence other mechanisms consume, and it does not itself deliver, dose, or roll out the combination.

Example

An agricultural extension service believes that a new package for smallholder maize — improved seed, targeted fertilizer, and a low-cost irrigation method — will lift yields far more than any input alone. Rather than promote the package on that belief, it runs a factorial field trial. Plots are randomly assigned across arms that isolate the parts: seed only, fertilizer only, irrigation only, each pair, and the full three-way combination, plus a control. This structure is what lets the analysts read interaction evidence directly — they can see whether the full package beats the sum of the single-input gains, which is the whole synergy question. The combined measure compares the package's yield against both the isolated arms and their additive prediction. And the trial surfaces an antagonism finding the belief would have hidden: on the driest plots, fertilizer without irrigation actually depresses yield by scorching young plants, so one "complementary" pair is antagonistic outside the presence of the third element. The service scales only the combination the trial proves reinforcing, and drops the pairing the trial shows harmful — a decision it could not have made from plausibility alone.

How it works

  • Build arms that isolate the parts. Test the full combination against each element alone, against alternative bundles, and against the additive baseline — the comparison structure is the design, because synergy is only visible relative to isolated arms.
  • Read the interaction, not just the total. Compare the combination's result to what the separate elements predict added together; the gap above (or below) additive is the evidence of reinforcement (or antagonism).
  • Measure combined effect against baselines. Quantify the joint outcome against isolated and alternative-bundle arms, including the coordination and complexity costs the combination carries.
  • Flag antagonism explicitly. Treat a combination that underperforms its own parts as a first-class finding, not noise, and let it kill or reshape the bundle before scale.

Tuning parameters

  • Arm structure — how many isolating arms to run (single elements, pairs, full factorial). A full factorial reveals every interaction but needs many units and large samples; a reduced design is cheaper but can miss a specific pairwise antagonism.
  • Statistical power for interaction — how large a sample the test carries. Detecting an interaction effect needs far more power than detecting a main effect; underpowering guarantees an ambiguous verdict on the very thing being tested.
  • Pilot fidelity vs. scale realism — how closely the pilot mirrors real deployment. High-fidelity pilots generalize better but cost more and run slower; quick, artificial pilots answer fast but may not survive contact with scale.
  • Decision threshold — how much amplification over additive counts as "worth the coordination cost." A strict threshold guards against marginal bundles; a loose one green-lights combinations whose gain does not repay their complexity.

When it helps, and when it misleads

Its strength is that it is the archetype's truth-teller: it converts "these should reinforce" into a measured verdict, protects against bundle bloat and false-synergy claims, and is the one mechanism that can catch hidden antagonism before it scales. The factorial design is the classic instrument here — by crossing factors it isolates each interaction term, letting the analyst see whether the combined effect exceeds, matches, or falls short of the sum of the parts.[n1]

Its failure mode is a pilot underpowered or mis-structured for interaction: with too few units or no isolating arms, it reports a combined total but cannot attribute it, and a proud "the bundle worked" hides that the parts alone would have done as well. A subtler trap is a pilot so artificial it proves synergy that evaporates at scale. The classic misuse is running only a single-arm pilot of the full bundle and calling any good result proof of synergy — which measures the total, not the interaction. The guarding discipline is to insist on isolating arms and adequate power for the interaction term, and to treat an antagonism finding as a result to act on rather than an inconvenience to explain away.

How it implements the components

  • interaction_evidence — its core product: evidence of reinforcement generated by direct comparison of the combination against isolating arms, not by assertion.
  • combined_effect_measure — the quantified comparison of the joint outcome against isolated, additive, and alternative-bundle baselines, coordination costs included.
  • antagonism_check — the empirical detection of a combination that underperforms its own parts, treated as a first-class finding that can kill the bundle.

A pilot bundle comparison does not deliver the combination in production, set its coupling_or_sequence_rule, or carry a fallback_decoupling_path — that is combination_therapy_regimen, its nearest twin (it shares evidence and antagonism-screening) which consumes this evidence to safely deliver a live combination; nor does it phase a staged_combination_rollout, which belongs to policy_package_design. The pilot only tests the combination.

Editorial Notes

Form Classification

Form family: Experiment, Test & Rehearsal

Rationale: Pilot Bundle Comparison operates as an active test, trial, simulation, drill, or rehearsal that generates evidence through a deliberate attempt or perturbation because it tests the proposed combination against isolated elements or alternative bundles before committing to scale.

Independent corroboration: The frozen evidence defines Pilot Bundle Comparison as 'Tests the proposed combination against isolated elements or alternative bundles before committing to scale', so its operative form is Experiment, Test & Rehearsal.

Review outcome: Independent reviewer agreement; high confidence.

Origin Attribution

Primary origin: Statistics & Experimental Design

Origin pattern: Cross-disciplinary synthesis

Present-day reach: Multi-domain

Rationale: Pilot Bundle Comparison is rooted in experimental design and statistics: Factorial experimental design isolates component effects and the interaction value of a proposed bundle.

Related originating lineages:

Review resolution: Both blind reviewers agree that statistics and experimental design is the primary origin. Reconciliation resolves alternate_origin_disagreement, origin_mode_disagreement. Formative alternate lineages are retained as organizational_management; later breadth of use is recorded separately as domain_reach=multi_domain, while origin_mode=cross_disciplinary_synthesis describes the relationship among origin lineages.

Encyclopedia synthesis: The exact catalogued form synthesizes established practice rather than reproducing a single standard historical label.

Review outcome: Reconciled after independent review; high confidence.

Notes

Pilot Bundle Comparison is deliberately upstream of every delivery mechanism here: it is consumed by bundled_intervention_protocol, combination_therapy_regimen, product_bundle_design, and policy_package_design, each of which leans on its interaction evidence before scaling. Keeping the evidence step separate from delivery is what lets a team improve its test — more arms, more power — without re-litigating a rollout already underway.

[n1] A factorial design is an experiment that crosses two or more factors so that every combination of their levels is tested, allowing the analyst to estimate not only each factor's main effect but the interaction term — the extra (or missing) effect of applying the factors together — which is precisely the non-additive value this archetype is built around.