Skip to content

Semantic Retrieval Diagnostic — does the catalog's cross-domain content become retrievable? (2026-05-22)

Motivation. In the forward transfer test, honest scenario-derived lexical search never surfaced the procedural_fairness_due_process prime from the immunology scenario. Hypothesis (Kurt): the MCP search is naive/lexical; a semantic (vector) search should surface ideas by meaning, not verbiage. Sub-question (Claude): does the fix need the meta-model (domain-stripped query) to reach a framed prime across domains?

Setup. Local BGE-small-en-v1.5 (ONNX, the model Kurt staged in _models/), CLS-pooled + L2-normalized, BGE query instruction on queries. Embedded all 568 primes two ways — over core_idea (prose) and over structural_signature (domain-stripped). Two queries: Q_raw = the raw immunology scenario; Q_meta = a domain-stripped meta-model of the problem (autonomous agent, irreversible destructive action, ambiguous identity from noisy signals, catastrophic-vs-recoverable asymmetry, no oversight — no solution language, no legal terms). Cosine ranking; report rank of due_process (of 568) and the top hits.

Result: rank of procedural_fairness_due_process (of 568)

Config rank
Lexical (scenario terms) absent (never retrieved)
Semantic — Q_raw × prose 217
Semantic — Q_raw × signature 186
Semantic — Q_meta × prose 129
Semantic — Q_meta × signature 45

What surfaces at the top changes completely with the query. Under Q_raw the top hits are pure pharmacology/biomedicine (buffering, half_life, dose_response, therapeutic_window, receptor_saturation, pk_pd_modeling) — the raw scenario embeds into biomedical space, and due_process is buried at ~200. Under Q_meta the top hits become the right kind of abstraction: reversibility_and_irreversibility (#1), irreversibility (#2), authority_delegation_under_uncertainty (#3), consent (#7), moral_hazard, incentive_compatibility, sunk_cost_and_irreversible_commitment. The decision-theory / governance cousins jump to the top; due_process climbs to 45 but stays outside any practical top-k.

Interpretation — confirms both layers

Layer 1 (lexical brittleness): semantic search clearly fixes it. The structural/decision-theory cousins that lexical search missed or under-ranked (irreversibility, reversibility, consent, sunk-cost-commitment, authority-delegation) now rank at the top under a meta-model query. A vector index would reliably surface exactly the cluster that leads to irreversible_commitment_management / independent_verification_oversight. The pipeline's near-concept retrieval has been badly handicapped by keyword matching in every prior experiment; this is a real, fixable confound. Build the vector search.

The meta-model is confirmed as the key cross-domain retrieval artifact. Q_meta beats Q_raw decisively (due_process 45 vs 217; and the whole top-list flips from biomed noise to structural/governance primes), and embedding the structural_signature beats embedding the prose for the meta query (45 vs 129). So the right design is: build the domain-stripped meta-model, then retrieve it against structural-signature embeddings. This gives the representation work (Step ⅚) a concrete retrieval payoff — possibly its main value, separate from reasoning.

Layer 2 (cross-domain retrieval to the FRAMED prime): off-the-shelf embeddings do NOT solve it. Best case (Q_meta × signature) still leaves due_process at 45/568 — better than lexical's "absent," but below any realistic top-k cutoff (top-10/20). The framed prime stays partly stuck in its legal neighborhood even when queried with neutral, domain-stripped structure. This is exactly what the structural/framed theory predicts: a structural prime's signature is domain-neutral and embeds near the abstract query (irreversibility, reversibility, threshold rank near the top), but a framed prime's signature carries institutional/legal vocabulary that holds it at distance. The frame shows up as embedding distance. So the retrieval bottleneck is worst for precisely the framed primes that the forward test suggested carry the catalog's unique transfer value.

Consequences

  • Some of the earlier "the catalog adds nothing / the protocol does the work" findings carry a retrieval confound: the catalog's relevant content was under-retrieved by keyword search. A semantic + meta-model retriever partially lifts this — but the headline survives: the structural core is reachable by reasoning regardless, and even good semantic retrieval does not surface the framed prime itself (it surfaces its structural/governance cousins, which is how the manual forward run actually reached the procedural design).
  • Realistic catalog runtime value, restated: with semantic+meta-model retrieval the pipeline would reliably get the structural/governance neighborhood of a framed prime (and the archetypes those neighbors source), nudging toward the framed design — but not retrieve the framed prime by identity.
  • To actually reach framed primes across domains likely needs a stronger or fine-tuned embedder trained on catalog (instance ↔ prime) pairs — i.e., the synthetic-data approach from the project's origin. bge-small is a small general model; this result is a floor, not a ceiling.

Method caveats

Truncation 200 tokens (captures each entry's lead, incl. due_process's "notice → hearing → impartial decision → reasoned justification"); bge-small (384-dim) is a modest model; single scenario / single framed prime. The framed-vs-structural sweep (several of each) would turn this single data point into a pattern — and the structural primes already behaved as predicted here (irreversibility/reversibility/threshold ranked near the top under the meta query, due_process did not).

Script: outputs-side /tmp/diag2.py; queries /tmp/q_raw.txt, /tmp/q_meta.txt.