Skip to content

AAR Pipeline vs. Bare Chain-of-Thought — Re-run Report (2026-05-22)

Protocol: baseline (no §9 variation), per applications/aar_vs_cot_experiment_kit.md. Corpus snapshot: 568 primes / 621 archetypes (605 generated + 16 hand-curated); MCP server reloaded 2026-05-22 07:02. Gate cleared (kit expects ~568 primes, ~600+ archetypes). Scenario: the §2 coastal demersal fishery (open-access, ~120 operators / 5 villages, ~20% over MSY, recruitment-collapse threshold 4–7 yr; targets: biomass to MSY-range within 10 yr AND <20% aggregate income drop, with no bottom-quartile collapse below subsistence). One run. Directional only — not evidence on its own.


1. Result at a glance

Dimension Pipeline (Cond A) Bare CoT (Cond B) Winner
1. Structural coherence 9 9 tie
2. Failure-mode anticipation 9 9 tie
3. Implementability 8 8 tie
4. Mechanism grounding 9 9 tie
5. Stakeholder / equity fit 9 9 tie
6. Threshold / time-horizon awareness 9 8 pipeline
TOTAL 53/60 52/60 pipeline (+1)

Grader's pick: the pipeline output. The grader called the margin "small … but a real edge, not a coin flip," and singled out the pipeline's discipline of carrying explicit uncertainty bands on the rebuilding cap's glide path — refusing to treat the 4–7-year collapse estimate as a hard number — as "the single most sophisticated move in either document."

Dimensions where bare CoT beat the pipeline: none on score (it tied five, lost one). But the grader credited bare CoT with one genuine qualitative advantage the pipeline lacked: an explicit operational tie-breaker — "expand the bridge support before tightening further" if fisher income nears the danger line — which resolves the central income-vs-biomass tension under stress more crisply. The grader judged this "real but narrow, and easily ported into" the pipeline design.

2. Comparison against the baseline (first run, 2026-05-05)

Pipeline Bare CoT Delta CoT-won dimensions
Baseline (2026-05-05) 50/60 45/60 +5 implementability (8v7), mechanism grounding (9v8)
This run (2026-05-22) 53/60 52/60 +1 none (CoT tied 5, lost threshold)

Two things changed materially:

  1. The gap collapsed (+5 → +1). The two outputs converged to near-identical recommendations this run — same skeleton (entry freeze → front-loaded sub-replacement cap → catch-share allocation with anti-concentration caps → off-the-top subsistence block + transition bridge → gear/size rules → co-management under a binding regulator cap with automatic triggers). A +1/60 margin from a single grader is well within noise; on its own it is indistinguishable from a tie.

  2. The pipeline's edge moved. In the baseline, the pipeline won the three "catalog-favoring" dimensions (coherence, failure-mode, equity) and lost implementability and mechanism grounding. This run, those three catalog-favoring dimensions all tied at 9–9, and the pipeline's sole winning margin came on threshold/time-horizon awareness — one of the three dimensions the kit considers neutral to the catalog. The contribution traces to the tipping_point_prevention archetype (threshold-estimate-with-uncertainty, "intervention window as the scarce resource," false-precision failure mode), not to the commons_governance content — which bare CoT reproduced essentially in full from training knowledge.

3. What the pipeline actually contributed

The pipeline anchored on two catalog archetypes: commons_governance (primary; a near-exact structural match whose components — access rule, condition-indexed use quota, monitoring signal, graduated sanctions, replenishment, legitimacy basis, dispute resolution, adaptation cadence — map directly onto fishery levers) and tipping_point_prevention (complement; supplying the urgency/irreversibility logic). Both passed their documented anti-signature checks cleanly, and commons_governance explicitly flags public_goods_provision as the wrong archetype for overuse of a rival stock — a useful negative check the pipeline recorded.

Honest read: bare CoT independently reconstructed almost the entire commons_governance design (allocation, anti-concentration, subsistence floor, co-management, monitoring, selectivity, triggers) without the catalog. The catalog's distinctive marginal contribution in this run was narrow but real — the explicit threshold-uncertainty discipline from tipping_point_prevention. That is the only place the grader saw daylight.

4. Limitations (carried from §8, plus run-specific)

  • N=1; directional only. A +1/60 result from one problem and one grader does not support any general claim. The honest summary of this run is "the two methods produced near-equivalent recommendations; the grader narrowly preferred the pipeline on threshold handling."
  • Operator contamination — flagged strongly this run. The same operator ran both conditions, and (by the kit's design) had the bare-CoT output in working context when finalizing the pipeline output. The pipeline's recommendation independently included landing/buyer choke-point monitoring + graduated sanctions — which in the baseline run was documented as bare-CoT's unique contribution that the pipeline missed. Its appearance in the pipeline this run may be partial borrowing rather than independent catalog derivation. This confound cuts against the pipeline's added-value claim (the convergence may be the pipeline absorbing CoT's strength), and it likely explains part of why the gap shrank. A clean re-run should isolate the operator (separate person/agent for each condition) or generate Condition A before ever seeing Condition B.
  • Rubric overlap (3 of 6 dims favor the catalog). Note that this run those three dimensions tied, so the pipeline's win did not come from the biased dimensions — a point in the result's favor, but a fragile one at this sample size.
  • Same model class for both conditions and the grader. Shared systematic biases cannot be ruled out.
  • Single grader. No inter-rater check.
  • Mature-domain selection. The fishery sits in a catalog-strong region (commons governance). The near-complete CoT reconstruction suggests this is a region where training knowledge is already dense, which may shrink the catalog's marginal value rather than showcase it. Says nothing about catalog-thin regions.
  • Grader-format bias mitigated: both outputs were normalized to the same skeleton and length (~780–840 words) before grading, so style was not a tell.

The convergence + contamination make the §9.2 ablation the priority: a CoT agent given a generic structural-reasoning scaffold ("identify the structural patterns; map relationships; check failure modes") but no catalog. If that recovers most of the pipeline's edge, the protocol — not the catalog — is doing the work. Given that bare CoT already reconstructed commons_governance unaided here, this is the decisive test. Also worth running: multiple independent graders (§9.3) for an inter-rater check; a catalog-thin scenario to test whether the pipeline's edge widens where training knowledge is sparse (the maturity hypothesis); and a clean operator separation to remove the contamination flagged above.


Appendix — provenance

  • Condition A (pipeline) full run incl. operative-prime set, context model, meta-model, dual-view reasoning, anti-signature evaluation: outputs/condition_A_pipeline.md
  • Condition B (bare CoT, 0 tools) full output: outputs/condition_B_bare_cot.md
  • Normalized outputs + private A/B mapping: outputs/normalized_outputs_PRIVATE.md
  • Blind mapping: Recommendation A = bare CoT; Recommendation B = pipeline.