AAR Pipeline vs. Bare Chain-of-Thought — Re-run Report (2026-05-22)¶
Protocol: baseline (no §9 variation), per applications/aar_vs_cot_experiment_kit.md.
Corpus snapshot: 568 primes / 621 archetypes (605 generated + 16 hand-curated); MCP server reloaded 2026-05-22 07:02. Gate cleared (kit expects ~568 primes, ~600+ archetypes).
Scenario: the §2 coastal demersal fishery (open-access, ~120 operators / 5 villages, ~20% over MSY, recruitment-collapse threshold 4–7 yr; targets: biomass to MSY-range within 10 yr AND <20% aggregate income drop, with no bottom-quartile collapse below subsistence).
One run. Directional only — not evidence on its own.
1. Result at a glance¶
| Dimension | Pipeline (Cond A) | Bare CoT (Cond B) | Winner |
|---|---|---|---|
| 1. Structural coherence | 9 | 9 | tie |
| 2. Failure-mode anticipation | 9 | 9 | tie |
| 3. Implementability | 8 | 8 | tie |
| 4. Mechanism grounding | 9 | 9 | tie |
| 5. Stakeholder / equity fit | 9 | 9 | tie |
| 6. Threshold / time-horizon awareness | 9 | 8 | pipeline |
| TOTAL | 53/60 | 52/60 | pipeline (+1) |
Grader's pick: the pipeline output. The grader called the margin "small … but a real edge, not a coin flip," and singled out the pipeline's discipline of carrying explicit uncertainty bands on the rebuilding cap's glide path — refusing to treat the 4–7-year collapse estimate as a hard number — as "the single most sophisticated move in either document."
Dimensions where bare CoT beat the pipeline: none on score (it tied five, lost one). But the grader credited bare CoT with one genuine qualitative advantage the pipeline lacked: an explicit operational tie-breaker — "expand the bridge support before tightening further" if fisher income nears the danger line — which resolves the central income-vs-biomass tension under stress more crisply. The grader judged this "real but narrow, and easily ported into" the pipeline design.
2. Comparison against the baseline (first run, 2026-05-05)¶
| Pipeline | Bare CoT | Delta | CoT-won dimensions | |
|---|---|---|---|---|
| Baseline (2026-05-05) | 50/60 | 45/60 | +5 | implementability (8v7), mechanism grounding (9v8) |
| This run (2026-05-22) | 53/60 | 52/60 | +1 | none (CoT tied 5, lost threshold) |
Two things changed materially:
-
The gap collapsed (+5 → +1). The two outputs converged to near-identical recommendations this run — same skeleton (entry freeze → front-loaded sub-replacement cap → catch-share allocation with anti-concentration caps → off-the-top subsistence block + transition bridge → gear/size rules → co-management under a binding regulator cap with automatic triggers). A +1/60 margin from a single grader is well within noise; on its own it is indistinguishable from a tie.
-
The pipeline's edge moved. In the baseline, the pipeline won the three "catalog-favoring" dimensions (coherence, failure-mode, equity) and lost implementability and mechanism grounding. This run, those three catalog-favoring dimensions all tied at 9–9, and the pipeline's sole winning margin came on threshold/time-horizon awareness — one of the three dimensions the kit considers neutral to the catalog. The contribution traces to the
tipping_point_preventionarchetype (threshold-estimate-with-uncertainty, "intervention window as the scarce resource," false-precision failure mode), not to thecommons_governancecontent — which bare CoT reproduced essentially in full from training knowledge.
3. What the pipeline actually contributed¶
The pipeline anchored on two catalog archetypes: commons_governance (primary; a near-exact structural match whose components — access rule, condition-indexed use quota, monitoring signal, graduated sanctions, replenishment, legitimacy basis, dispute resolution, adaptation cadence — map directly onto fishery levers) and tipping_point_prevention (complement; supplying the urgency/irreversibility logic). Both passed their documented anti-signature checks cleanly, and commons_governance explicitly flags public_goods_provision as the wrong archetype for overuse of a rival stock — a useful negative check the pipeline recorded.
Honest read: bare CoT independently reconstructed almost the entire commons_governance design (allocation, anti-concentration, subsistence floor, co-management, monitoring, selectivity, triggers) without the catalog. The catalog's distinctive marginal contribution in this run was narrow but real — the explicit threshold-uncertainty discipline from tipping_point_prevention. That is the only place the grader saw daylight.
4. Limitations (carried from §8, plus run-specific)¶
- N=1; directional only. A +1/60 result from one problem and one grader does not support any general claim. The honest summary of this run is "the two methods produced near-equivalent recommendations; the grader narrowly preferred the pipeline on threshold handling."
- Operator contamination — flagged strongly this run. The same operator ran both conditions, and (by the kit's design) had the bare-CoT output in working context when finalizing the pipeline output. The pipeline's recommendation independently included landing/buyer choke-point monitoring + graduated sanctions — which in the baseline run was documented as bare-CoT's unique contribution that the pipeline missed. Its appearance in the pipeline this run may be partial borrowing rather than independent catalog derivation. This confound cuts against the pipeline's added-value claim (the convergence may be the pipeline absorbing CoT's strength), and it likely explains part of why the gap shrank. A clean re-run should isolate the operator (separate person/agent for each condition) or generate Condition A before ever seeing Condition B.
- Rubric overlap (3 of 6 dims favor the catalog). Note that this run those three dimensions tied, so the pipeline's win did not come from the biased dimensions — a point in the result's favor, but a fragile one at this sample size.
- Same model class for both conditions and the grader. Shared systematic biases cannot be ruled out.
- Single grader. No inter-rater check.
- Mature-domain selection. The fishery sits in a catalog-strong region (commons governance). The near-complete CoT reconstruction suggests this is a region where training knowledge is already dense, which may shrink the catalog's marginal value rather than showcase it. Says nothing about catalog-thin regions.
- Grader-format bias mitigated: both outputs were normalized to the same skeleton and length (~780–840 words) before grading, so style was not a tell.
5. Recommended next steps¶
The convergence + contamination make the §9.2 ablation the priority: a CoT agent given a generic structural-reasoning scaffold ("identify the structural patterns; map relationships; check failure modes") but no catalog. If that recovers most of the pipeline's edge, the protocol — not the catalog — is doing the work. Given that bare CoT already reconstructed commons_governance unaided here, this is the decisive test. Also worth running: multiple independent graders (§9.3) for an inter-rater check; a catalog-thin scenario to test whether the pipeline's edge widens where training knowledge is sparse (the maturity hypothesis); and a clean operator separation to remove the contamination flagged above.
Appendix — provenance¶
- Condition A (pipeline) full run incl. operative-prime set, context model, meta-model, dual-view reasoning, anti-signature evaluation:
outputs/condition_A_pipeline.md - Condition B (bare CoT, 0 tools) full output:
outputs/condition_B_bare_cot.md - Normalized outputs + private A/B mapping:
outputs/normalized_outputs_PRIVATE.md - Blind mapping: Recommendation A = bare CoT; Recommendation B = pipeline.