AAR v2 (refined pipeline) vs. Bare CoT — Hard Multi-Domain Run (2026-05-22)¶
What this run changed (two deliberate deviations from baseline at once): 1. Refined protocol ("AAR v2") — the Steps 5–9 upgrades agreed with Kurt: typed-graph JSON model with referential integrity; meta-model as a declared transform of the context model; recall-friendly retrieval then two-stage triage to ≤12 candidates; a per-archetype 4-axis scorecard with the anti-signature as a hard gate; explicit composition of multiple archetypes plus discard-to-CoT for residual tensions via a coverage check; and provenance tagging of every recommendation element. 2. A deliberately hard, catalog-rich, multi-domain scenario — a 6-region two-sided gig-services marketplace ("Tasca") in a trust-and-liquidity doom-loop. Chosen (per Kurt's instruction) NOT to test catalog completeness but to stress the new archetype-iteration machinery in a dense region: it pulls reinforcing feedback (the densest hub at 307 archetypes), network effects, trust, monitoring/bottleneck, diffusion/contagion, incentive/mechanism design, governance, threshold/critical-mass, and equity — with no single home archetype.
Corpus: 568 primes / 621 archetypes. Single blinded grader, same 6-dimension rubric, same model class. One run — directional only.
1. Result at a glance¶
| Dimension | AAR v2 pipeline | Bare CoT | Winner |
|---|---|---|---|
| 1. Structural coherence | 9 | 8 | pipeline |
| 2. Failure-mode anticipation | 9 | 7 | pipeline (+2) |
| 3. Implementability | 8 | 8 | tie |
| 4. Mechanism grounding | 9 | 8 | pipeline |
| 5. Stakeholder / equity fit | 9 | 8 | pipeline |
| 6. Threshold / time-horizon awareness | 9 | 8 | pipeline |
| TOTAL | 53/60 | 47/60 | pipeline (+6) |
Grader's pick: the pipeline ("Recommendation A"), explicitly and without splitting the difference. It judged the two "substantively the same intervention," differing in "depth of systems thinking and risk closure, not direction," and credited the pipeline with (a) pairing every named failure mode to a specific counter-measure, (b) an evidence-based subsidy sunset (does activity persist without support?) rather than a date-based one, and © explicit invariants that make the system's coupling legible. Bare CoT won zero dimensions (tied implementability, lost the other five).
2. Comparison against the prior runs¶
| Run | Scenario | Pipeline | Pipeline | CoT | Δ | CoT-won dims |
|---|---|---|---|---|---|---|
| Baseline (05-05) | fishery (catalog-strong, single-archetype) | v1 | 50 | 45 | +5 | implementability, mechanism |
| Re-run (05-22) | fishery | v1 | 53 | 52 | +1 | none (tied 5) |
| This run (05-22) | marketplace (catalog-rich, multi-archetype) | v2 | 53 | 47 | +6 | none (tied 1, lost 5) |
The gap is the widest yet (+6), and it widened most on failure-mode anticipation (+2) — the dimension the v2 scorecard/residual-risk discipline most directly targets. But see the attribution caveat below: this run changed two variables at once (harder scenario and refined protocol), so the +6 cannot be cleanly assigned to either.
3. What the refined pipeline actually did (and where the value showed up)¶
- Retrieval was rich, so triage did real work. Unlike the fishery (where source-prime retrieval returned essentially one archetype), here the candidate pool was ~16 and the two-stage triage cut to 8 finalists. This is the regime the refinement was built for.
- The anti-signature gate culled a keyword match.
participation_equity_and_inclusion_designmatched on "equity/inclusion" but tripped its own anti-signature ("the task is only to allocate resources/benefits equitably, with no collective activity or participation field to design") — it is structurally about voice in rituals/meetings. The gate discarded it; the equity floor was instead handled by incentive-design components plus CoT (opportunity reservation + subsistence-floor monitoring). The grader then scored the pipeline 9/10 on equity fit — i.e., discarding the wrong archetype and falling back to CoT produced better equity handling than forcing the superficial match would have. This is the "feel free to discard" behavior working as intended. - Composition + CoT-fallback was the operating mode, not single-archetype selection. The pipeline composed 7 archetypes (incentive_compatible_rule_design as spine; cycle_breaking + critical_mass_building for the two reinforcing loops; independent_verification_oversight + diffusion_containment for the trust/abuse subsystem; service_rate_matching for the bottleneck; network_effect_governance for platform rules) and explicitly routed three residual tensions to CoT (the equity floor, the legal control-vs-credential line, and the three-clock sequencing). Provenance tags recorded which was which.
- The grader-rewarded features trace to the archetype-handling refinements, not the representation. The evidence-based sunset came from
critical_mass_building's self-sustainability-test component; the paired failure/counter-measures came from the scorecard's residual-risk discipline; the explicit invariants are a Step-9 output. The typed-graph JSON (Step ⅚) is not visibly what earned the points. Consistent with the prior prediction: the value lives in retrieval/selection/composition, and the representation upgrade is lower-leverage.
4. Limitations (carry into any write-up; do not overclaim)¶
- N=1; single grader; same model class. A +6/60 from one problem and one grader is directional, not evidence.
- Two variables changed at once. Harder scenario AND refined protocol. This run cannot separate "the refinements helped" from "the pipeline does relatively better on hard multi-archetype problems." A clean test holds one fixed.
- Operator contamination persists and is the biggest threat. The same operator ran both conditions and saw the bare-CoT output before finalizing the pipeline output. The grader found the two plans "substantively the same intervention," which is exactly what contamination would produce. The pipeline's margin (coherence/failure-mode/threshold) is more defensible because it maps to specific protocol features, but the convergence may be operator-induced. A clean re-run needs operator isolation or Condition A generated before Condition B is seen.
- The bundle was tested as a bundle. Typed-graph + scorecard + gate + composition + CoT-fallback + provenance were all on at once; the gain cannot be attributed to any single refinement. This run, indirect evidence points to the archetype-handling pieces over the representation piece.
- Still no §9.2 ablation. The prompt-only structural-scaffold control (CoT given "find the patterns / map relationships / check failure modes / consider neighbors" but no catalog) remains the decisive missing comparison — especially now, since much of the pipeline's edge is procedural discipline (scorecard, coverage check, invariants) that a scaffold prompt might replicate without any catalog.
- Format normalized. Both outputs were re-rendered to the same skeleton/length (~820–880 words) and the pipeline's provenance tags/method vocabulary stripped before grading.
5. Recommended next steps¶
- Run the §9.2 ablation on this same marketplace scenario — CoT + structural-scaffold prompt, no catalog. If it lands between 47 and 53, that quantifies how much of the +6 is protocol discipline vs catalog content. This is now the single highest-value experiment.
- Isolate one variable: either (a) run AAR v2 on the fishery (refined protocol, easy scenario) to compare against the v1 fishery numbers, or (b) run v1 on the marketplace (old protocol, hard scenario). Either isolates protocol-vs-scenario.
- Operator isolation for at least one run, to test whether the cross-condition convergence is real or induced.
- Multiple graders to put an error bar on a margin that is now large enough to be interesting.
Appendix — provenance¶
- Condition A (AAR v2) full run incl. typed-graph JSON, meta-model transform, triage, scorecard with the anti-signature discard, composition + CoT-coverage check:
outputs/condition_A_pipeline_v2.md - Condition B (bare CoT, 0 tools):
outputs/condition_B_bare_cot_v2.md - Normalized outputs + private mapping:
outputs/normalized_outputs_v2_PRIVATE.md - Blind mapping: Recommendation A = pipeline; Recommendation B = bare CoT.
- Catalog-density sampling that selected the scenario:
feedback=307 archetypes (densest hub),monitoring=14,network_effect=13,trust=12,coordination=10, plus dense threshold/cascade/governance neighborhoods.
6. §9.2 ABLATION (added 2026-05-22) — and it revises the headline¶
The single most important control, run on the same marketplace scenario. Condition C = a CoT agent given the full AAR v2 structural process (operative-pattern ID → typed relational model → domain-stripped meta-model → enumerate/instantiate/score/anti-pattern-gate candidate solution templates → compose + discard-to-first-principles → neighbor check) but no catalog and no tools. All three conditions were then re-graded together by a fresh blinded grader on one scale (labels randomized).
| Condition | What it has | Total |
|---|---|---|
| Bare CoT | no process, no catalog | 46 / 60 |
| AAR v2 pipeline | process + full catalog | 53 / 60 |
| Scaffold CoT | process, NO catalog | 55 / 60 |
The grader ranked scaffold ≥ pipeline > bare CoT and adopted the scaffold. (Re-grade stability check: bare CoT 47→46 and pipeline 53→53 between the two grading sessions, so the scales are comparable.)
Interpretation — this flips the reading of the +6. The improvement over bare CoT is almost entirely the protocol discipline, not the catalog. The scaffold — which encodes the structural process but has no catalog access — recovered the entire gain and then some (55 vs the pipeline's 53; the +2 is within single-grader noise, so the honest statement is scaffold ≈ pipeline, both clearly above bare CoT). On this problem the catalog's marginal contribution over a good structural-reasoning prompt was undetectable. The grader specifically praised the scaffold's threshold realism (worst-but-savable triage, containing an already-dead region, margin-based subsidy withdrawal) and its paired failure-mode counters — process features, not catalog content.
What this does and does not establish.
- It does NOT say the catalog is worthless. It says: for a problem sitting in a catalog-rich, training-data-rich region (commons/marketplace design), a frontier model already holds the relevant solution patterns, so a prompt that makes it use them systematically captures the value, and a runtime lookup adds little. This is consistent with the fishery, where bare CoT reconstructed commons_governance unaided.
- It strongly reframes the project's likely value proposition: the catalog's contribution may be (a) as the source/teacher of the protocol discipline, and (b) as a resource for cross-domain transfer where the model's training knowledge is thin — exactly the regime the next experiment targets. Where CoT can reconstruct the pattern, the catalog is redundant; the open question is whether it earns its keep where CoT cannot.
Caveats (unchanged + new). N=1, single grader, same model class. Operator contamination persists across all three. And note the scaffold prompt is itself derived from the catalog-informed protocol — it is catalog-free in access but encodes the discipline the catalog taught us; the clean structure-free baseline is bare CoT (46), which is clearly worse. So the finding is "the disciplined process is the value, and it is separable from the catalog," not "unstructured reasoning suffices."
Consequence for next steps. The cross-domain transfer test is now the load-bearing experiment: it is the design most likely to isolate a contribution the catalog makes that the protocol-alone scaffold cannot — i.e., surfacing a structurally-apt pattern from a far/alien domain that CoT (even well-scaffolded) does not reach. Recommend running scaffold-CoT as a third condition there too, so we measure catalog-vs-protocol, not just pipeline-vs-bare-CoT.