Skip to content

Experiment Kit: Augmented Abstract Reasoning Pipeline vs. Bare Chain-of-Thought

Purpose. A self-contained, reproducible protocol for comparing the 9-step Augmented Abstract Reasoning (AAR) pipeline — executed against the Encyclopedia of Abstractions MCP server — against bare chain-of-thought (CoT) reasoning, judged by a blinded grader. This kit is designed so a fresh context window (with no memory of the original run) can reproduce the experiment, run variations, or scale it up.

Companion documents (read for background, not required to run): - applications/augmented_abstract_reasoning_pilot_comparison.md — the written report of the first run (the N=1 pilot, result 50/60 vs 45/60). - applications/augmented_abstract_reasoning.md — the full pipeline description and the fishery worked example (§5). - applications/substrate_independence_framework.md — the substrate-independence framing relevant to choosing scenarios.


0. Hypothesis (what the experiment tests)

The catalog complements CoT if, on problems where structural pattern recognition is load-bearing, a blinded grader prefers AAR-pipeline output over bare-CoT output produced by the same model class, on dimensions tied to structural reasoning (coherence, failure-mode anticipation, structural fit). One run is directional only; the claim requires many problems and multiple graders.


1. Prerequisites

  1. MCP server live. The Encyclopedia MCP server (mcp_server/server.py) must be connected and serving the current corpus. Verify with the corpus_stats tool; expect ~568 primes and ~600+ archetypes. If it returns stale counts, rebuild and restart:
    python3 scripts/mcp_preprocess.py
    python3 scripts/mcp_jsonl_export.py
    # then restart Claude Desktop (or `claude mcp restart encyclopedia-of-abstractions`)
    
  2. The 12 MCP tools available: search_prime, get_prime, search_archetype, get_archetype, find_archetypes_for_prime, find_related_primes, list_components, find_archetypes_using_component, list_mechanisms, find_archetypes_using_mechanism, corpus_stats, get_reasoning_pipeline_guide.
  3. Agent capability: ability to spawn independent sub-agents (for the bare-CoT condition and the blinded grader). Sub-agents must NOT share context with the pipeline operator.

2. The scenario (given verbatim to BOTH conditions)

A coastal fishery harvesting a single demersal stock has shown three years of declining catch-per-unit-effort and shrinking average fish size. The biologist's stock assessment estimates that current harvest exceeds maximum sustainable yield (MSY) by ~20% and that, on the current trajectory, the stock will cross a recruitment-collapse threshold within 4-7 years.

The fishery is open-access in practice (no enforced individual quotas); ~120 small-boat operators across 5 villages depend on it for income, with fish-processing buyers downstream. A regulator has authority to set total allowable catch but has historically deferred to the industry.

Decision required: What intervention design preserves the stock without destroying the livelihoods that depend on it?

Success criteria: - Stock biomass returns to MSY-supporting range within 10 years - AND <20% drop in aggregate fisher income (averaged across the cohort, not per-boat) - Failure = stock collapse OR a regulatory regime that pushes the bottom quartile of operators below subsistence.

(To run a variation, swap this scenario block. Keep the "decision required" + "success criteria" structure so both conditions and the grader have the same target.)


3. Condition A — AAR pipeline (operator executes against the MCP server)

The operator (the main context window) executes the 9-step pipeline live, using the MCP tools. Steps:

  1. Specify the problem statement from the scenario.
  2. Identify operative primes — search_prime for candidates, get_prime to inspect structural signatures. Err toward inclusion.
  3. Salience-rank the primes by load-bearing relevance.
  4. Prune to an operative subset (typically 4-9 primes), with explicit rationale.
  5. Build a context-specific model — a labeled relational graph (primes as node annotations, entities as nodes, relationships as labeled edges). (Discipline note: render this as an explicit typed structure, not just prose.)
  6. Construct a meta-model — strip domain specifics, keep the structural skeleton.
  7. Query archetypes — find_archetypes_for_prime (relation='source') for each operative prime; intersect/union to surface candidates; get_archetype for full records (components, mechanisms, anti-signatures, failure modes). search_archetype as fallback.
  8. Reason via both views — apply candidate archetypes' action logic to the context-specific model and the meta-model; reconcile.
  9. Evaluate fit & produce recommendation — walk each candidate's anti-signatures / trigger conditions / root tension; drop archetypes the scenario's anti-signatures trip; output a structured recommendation (chosen archetype, components/mechanisms instantiated, invariants, expected outcomes, residual risks tied to specific failure modes).

Tip: call get_reasoning_pipeline_guide(verbosity="full") at the start to load the protocol into working context.

Output: a final recommendation in plain prose. Save it. This is "Condition A output."


4. Condition B — Bare CoT (spawn an independent sub-agent)

Spawn a sub-agent with no tool access and no catalog, using exactly this prompt (substitute the scenario if running a variation):

Reason carefully through the following decision problem using chain-of-thought,
then produce a structured final recommendation. You have only your training
knowledge — no internet access, no external catalogs or frameworks beyond what
you already know.

[PASTE THE SCENARIO BLOCK FROM §2 — problem + decision required + success criteria]

Instructions:
1. Show your reasoning explicitly (chain-of-thought). Think about structural
   dynamics, stakeholder constraints, what could go wrong, tradeoffs, and
   neighboring approaches you considered and rejected.
2. Then produce a final structured recommendation. Use whatever section
   structure you think is most appropriate — there is no required format. It
   should be concrete enough that a regulator could implement it.
3. Stay under ~1000 words for the recommendation section.
4. Do NOT reference any specific external framework or catalog by name (no
   Ostrom, no IPCC, no named "design principles", no academic citations).
   Reason from first principles. This keeps the comparison on substance, since
   another method is being tested in parallel.

Write your reasoning, then the recommendation. No meta-commentary about the test.

Output: the recommendation section. This is "Condition B output."


5. Output normalization (remove method tells)

Before grading, re-render BOTH outputs into comparable plain prose and strip anything that identifies the method: - From Condition A: remove "archetype", "source primes", "anti-signatures", "the catalog", MCP tool names, and any encyclopedia-specific vocabulary. Keep the substance. - Both outputs should read as plain recommendations a consultant might write.

Assign labels A and B to the two normalized outputs. Record the mapping privately (which label = pipeline, which = CoT); the grader must not see it.

(Optional rigor upgrade: randomize the A/B assignment with a coin flip and have a third party hold the mapping.)


6. Grading (spawn a separate, blinded sub-agent)

Spawn a grader sub-agent that has NOT seen the pipeline run or the bare-CoT run. Give it both normalized outputs and this rubric:

You are an independent grader. Two recommendations were produced for the same
decision problem by two different reasoning methods. You do not know which
method produced which output. Score on substance only; do not try to identify
the methods.

[PASTE THE SCENARIO BLOCK FROM §2]

# Recommendation A
[paste normalized Condition A output]

# Recommendation B
[paste normalized Condition B output]

Score each recommendation on six dimensions, 1-10 each, with a one-sentence
justification per score:

1. Structural coherence — does it hang together as a system; would removing one
   piece predictably break another?
2. Failure-mode anticipation — does it name concrete, specific failure modes and
   propose a counter-measure for each?
3. Implementability — could the actors execute this within the stated
   constraints?
4. Mechanism grounding — are mechanisms concrete and actionable (specific rules,
   thresholds, paths) vs. vague aspirations?
5. Stakeholder/equity fit — does it engage with the heterogeneity of the actors
   and the success criterion's equity floor; is the equity logic robust under
   stress?
6. Threshold/time-horizon awareness — does it engage with non-linear collapse
   risk and the tension between the collapse window and the recovery target?

Output per recommendation:
  Recommendation [A/B]:
  1. Structural coherence: X/10 — [one sentence]
  ... (all six)
  TOTAL: XX/60

Then a ~250-word qualitative comparison: which is more likely to succeed, what
each does better, what each is missing, which residual risks each fails to
address. End with an explicit overall judgment: which would you adopt, and why.
Be honest. Don't artificially split the difference.

Output: per-dimension scores, totals, qualitative comparison, overall judgment.


7. Unblind & record

Reveal the A/B → method mapping. Record: - Per-dimension scores for pipeline vs CoT - Totals and the delta - The grader's overall pick - Any dimensions where CoT beat the pipeline (these are the most informative — they show where the catalog adds nothing or where bare CoT surfaced something the catalog missed)

Baseline result (first run, 2026-05-05): pipeline 50/60, bare CoT 45/60. Pipeline won coherence (9 vs 7), failure-mode anticipation (9 vs 6), equity-fit (9 vs 7); CoT won implementability (8 vs 7) and mechanism grounding (9 vs 8); threshold-awareness tied (8/8). Grader chose the pipeline output. Notably, the bare-CoT run independently surfaced a choke-point-enforcement insight (enforce at ~12-15 buyer/landing points rather than 120 boats) that the pipeline did not — a documented case of CoT contributing something the catalog missed.


8. Known limitations (carry these into any write-up)

  • N=1 per run; same model class for both conditions and the grader. Shared systematic biases cannot be ruled out.
  • Rubric overlap: three of six dimensions (coherence, failure-mode anticipation, structural/equity fit) are exactly what the catalog explicitly enforces, which biases toward the pipeline. On the other three dimensions alone the first run tied.
  • Operator conflict of interest: if the same person designs the kit, runs the pipeline, and analyzes results, that is a real bias. Blinded grading mitigates but does not eliminate it.
  • Grader-format bias: LLM graders may over-weight structured/list-formatted output; normalize formatting across both outputs.
  • Mature-domain selection: the fishery sits in a well-developed catalog region (commons governance). Results may not generalize to thin regions.

9. Variations worth running

  1. Different scenarios, stratified by catalog maturity. Pick problems whose operative primes are high-substrate-independence (catalog-strong) vs. low (catalog-thin). Hypothesis: pipeline's edge is larger in catalog-strong regions. (Use applications/substrate_independence_framework.md to choose.)
  2. Third condition — prompt-only structural scaffold. A CoT agent given a generic "identify the structural patterns; map relationships; check failure modes" prompt but NO catalog access. If this recovers most of the pipeline's gain, the protocol (not the catalog) is doing the work. This is the single most important ablation.
  3. Multiple independent graders (3+), report inter-rater agreement.
  4. Different base model for the conditions and/or grader, to test model-independence.
  5. Equal-effort control: give the bare-CoT agent the same wall-clock / token budget the pipeline consumed.
  6. Pre-registered rubric chosen by someone other than the pipeline author.

10. Reproduction checklist

  • MCP server live; corpus_stats sane
  • Scenario block fixed (original or variation)
  • Condition A: pipeline executed; output saved
  • Condition B: bare-CoT sub-agent run; output saved
  • Both outputs normalized; A/B mapping recorded privately
  • Grader sub-agent run on normalized A/B + rubric
  • Unblinded; scores + delta + grader pick recorded
  • Limitations noted; any CoT-wins flagged
  • (If formalizing) results appended to a results log with date + corpus snapshot