title: Experiment 12: Paired substrate-denial intervention¶
Experiment 12: Paired substrate-denial intervention¶
Part of Inverse Innovation with the Encyclopedia of Abstractions · Experiment 12: Paired substrate-denial intervention · Last revised August 2026
Status: complete prospective intervention with a triggered descriptive Max-effort follow-up
Run date: August 2026
Primary sample: 24 probability-sampled, previously untested generated archetypes × three fixed domains = 72 matched cells
Plain-language answer¶
Yes—but at a cost.
Earlier analysis found that 139 of 150 Experiment 9 proposals used either governance/process or computational/information as their main intervention. Experiment 12 asked whether that concentration reflected a hard limit or a default search habit. It reused Experiment 9's ordinary proposals as controls and instructed the same generator to solve the same 72 cells without making governance or computation the primary causal machinery.
The model followed that instruction in 67 of 72 cells. It produced 40 physical/material interventions and 27 measurement/instrumentation interventions that independent blinded classifiers judged compliant. This rules out the strongest interpretation of the earlier pattern: the tested pipeline is not incapable of leaving governance and computation.
The alternative proposals were usually worse, however. In blinded head-to-head comparison, the ordinary proposal won 68 cells and the constrained proposal won four. After authorized public-web screening, 49 constrained proposals and 60 ordinary proposals survived the same coarse prior-art screen. The restriction also changed what the system found: eight cells survived only in the constrained arm, while 19 survived only in the ordinary arm.
The best interpretation is therefore a soft but consequential default. Governance and computation are not the only substrates the system can use; they are generally the substrates through which it produced its strongest proposals in this sample. A ban can diversify the portfolio, but a ban alone is not an efficiency improvement.
Why this experiment was run¶
The post-hoc E10B and E9 substrate analyses found a two-level pattern:
- an archetype's structural-versus-framed character shifted allocation between governance and computational proposals; and
- across those differences, 92.7% of the balanced E9 proposals still remained in those two dominant channels.
That analysis was observational. It could not say whether the model was unable to create materially different interventions, whether the prompt merely made governance and computation convenient, or whether alternative substrates were available but lower quality. Experiment 12 intervened directly on the prompt while keeping the cell packet, model family, reasoning level, and matched ordinary output fixed.
Design¶
Primary comparison¶
The primary sample was the random-generated stratum from Experiment 9: 24 archetypes sampled from the previously untested generated-archetype corpus, each crossed with accounting/auditing, chemistry/materials, and computer science.
For every cell:
- Ordinary/High was the sealed Experiment 9 proposal generated by
gpt-5.6-solat High reasoning effort. - Constrained/High was a fresh isolated call to the same model and effort. It had to make the essential intervention physical/material, mechanical, chemical, biological, energetic, spatial, or measurement/instrumentation—or abstain.
Computation and governance could support the proposal, but removing software, algorithmic inference, incentives, permissions, reporting, or procedural enforcement could not remove its essential causal effect.
Blinded measurement¶
Two fresh gpt-5.6-terra High-reasoning passes classified every proposal under
opaque IDs. They assigned the primary substrate and one of four compliance
dispositions. A third blinded pass adjudicated every disagreement.
Two other fresh blinded passes compared each ordinary and constrained proposal in hash-randomized A/B order on structural fidelity, domain fidelity, causal coherence, operational specificity, testability, practicality, and contrivance risk. A third pass adjudicated winner disagreements. The judge did not see arm, archetype, domain, source score, or previous outcome. Abstention was a prespecified constrained loss.
Every compliant constrained proposal then received the unchanged Experiment 9 four-source web screen. The user explicitly authorized sending 67 primary constrained summaries and, after the Max trigger fired, 35 Max summaries and derived queries to public web services. All 102 screens completed without a failed transport call.
Frozen Max trigger¶
Before treatment generation, six archetypes—two from each source-score band—were selected by a deterministic hash rank and crossed with the three domains. Max effort would run only if:
- at least 24 constrained/High outputs were compliant; and
- the constrained arm won no more than 40% of decisive quality comparisons, or its usable yield was at least 15 percentage points below the ordinary control.
The trigger was a resource rule, not a confirmatory hypothesis test. It fired on both cost clauses.
Primary results¶
The model understood and usually obeyed the constraint¶
| Outcome | Cells | Share of 72 |
|---|---|---|
| Genuine allowed substrate | 60 | 83.3% |
| Mixed, but allowed substrate remained primary | 7 | 9.7% |
| Governance-primary proposal in disguise | 2 | 2.8% |
| Abstention | 3 | 4.2% |
| Compliant constrained output | 67 | 93.1% |
The Wilson 95% interval for compliant yield is 84.8%–97.0%. Among the 69 actual proposals, blinded adjudication assigned 40 to physical/material, 27 to measurement/instrumentation, and two to governance/process. The two governance cases failed the constraint. There were no forced-incoherent classifications.
Blinded proposal quality fell sharply¶
| Head-to-head winner | Cells | Share |
|---|---|---|
| Ordinary/High | 68 | 94.4% |
| Constrained/High | 4 | 5.6% |
| Tie | 0 | 0% |
The exact two-sided sign-test result is p = 4.62 × 10⁻¹⁶. Three ordinary
wins were deterministic abstention losses, but removing them would not change
the result's substance: among cells with two proposals, ordinary still won
65–4.
The four constrained winners were:
- a pressure-driven mechanical transfer totalizer for bulk-liquid cutoff audits;
- mechanically preloaded functional cassettes for electrochemical-cell experiments;
- a physical preimage atlas for colliding silicone-cure hardness states; and
- threshold-distributed hardener microcapsules for dampened thermoset cure.
Three of those four were chemistry/materials proposals. This is consistent with the idea that some domains naturally expose physical causal machinery, but four cases are too few for a general domain claim.
Web-screen survival also fell, more moderately¶
| Endpoint | Constrained/High | Ordinary/High | Difference, constrained − ordinary |
|---|---|---|---|
| Four-source light-screen survivor | 49/72 (68.1%) | 60/72 (83.3%) | −15.3 percentage points |
The Wilson interval for constrained survival is 56.6%–77.7%; the ordinary interval is 73.1%–90.2%. A 50,000-draw bootstrap clustered by archetype placed the paired difference between −30.6 and 0 percentage points. The paired binary table was:
| Ordinary survived | Ordinary failed | |
|---|---|---|
| Constrained survived | 41 | 8 |
| Constrained failed | 19 | 4 |
The exact paired McNemar/sign result for the 27 discordant cells is
p = 0.0522. That is not a license to declare the yields equivalent. The point
estimate and clustered interval favor ordinary generation, while the sample
does not locate the magnitude precisely.
The eight constrained-only survivors matter operationally. A diversified search can find candidates that ordinary generation misses even while lowering average yield. They included physical calibration or interlock proposals in chemistry/materials and computer science, plus one mechanically bounded audit sampler. They are researchable hypotheses under a coarse screen, not verified inventions.
Failure anatomy¶
Of the 72 constrained cells:
- 49 were compliant light-screen survivors;
- 18 were compliant but encountered established or substantially colliding prior art;
- two disguised governance as an allowed intervention; and
- three abstained.
This separates two bottlenecks. Only five cells failed to produce an eligible proposal. Most losses occurred after the model had obeyed the rule: the proposal was weaker than the ordinary alternative or collided with prior art.
Where the constraint hurt most¶
These subgroup results are descriptive. Domains were purposive, and the study was not powered for multiple subgroup tests.
By domain¶
| Domain | Compliant | Constrained usable | Ordinary usable | Constrained quality wins |
|---|---|---|---|---|
| Accounting/auditing | 21/24 | 19/24 | 23/24 | 1/24 |
| Chemistry/materials | 23/24 | 16/24 | 16/24 | 3/24 |
| Computer science | 23/24 | 14/24 | 21/24 | 0/24 |
Computer science is the clearest problem area. The generator could comply, but forcing a non-computational primary intervention often displaced the actual causal center of a computational problem. Chemistry/materials was the most natural home for the restriction: screen yield was equal across arms and it contained three of four constrained quality wins. Even there, ordinary proposals won 21 of 24 blinded comparisons, so equal light-screen yield should not be confused with equal proposal quality.
By archetype source-score band¶
| Band | Cells | Constrained usable | Ordinary usable | Constrained quality wins |
|---|---|---|---|---|
| Structural | 45 | 31 | 38 | 3 |
| Middle | 15 | 12 | 11 | 1 |
| Framed | 12 | 6 | 11 | 0 |
Framed archetypes were especially difficult. Concepts whose meaning depends on authority, consent, legitimacy, mentorship, incentives, or governance cannot always be converted into a physical device without losing what makes the archetype itself. Several constrained proposals turned those relations into keys, fixtures, gates, or coupled devices. They could be coherent artifacts while remaining strained representations of the source structure.
By allowed substrate¶
| Primary substrate | Proposals | Light-screen survivors | Constrained quality wins |
|---|---|---|---|
| Physical/material | 40 | 28 | 3 |
| Measurement/instrumentation | 27 | 21 | 1 |
Measurement was a common escape route. It yielded many researchable proposals, but only one defeated its ordinary counterpart. This suggests another local default: when software and governance are unavailable, the system often turns the abstraction into a sensor, witness, assay, or calibration device. That is sometimes exactly right, but it can also preserve observability while weakening the original intervention's causal force.
Max-effort follow-up¶
The 18-cell subset produced a four-arm comparison:
| Constraint | High | Max |
|---|---|---|
| Ordinary generation | sealed E9 proposal | fresh Max proposal |
| Substrate denial | primary E12 proposal | fresh Max proposal or abstention |
Constrained compliance was 16/18 at High and 17/18 at Max. Max therefore did not mainly help by teaching the model to obey the rule; compliance was already high.
Blinded quality¶
| Comparison | Result |
|---|---|
| Ordinary Max vs ordinary High | Max won 14–4 (p = 0.0309) |
| Constrained Max vs constrained High | Max won 12–6 (p = 0.2379) |
| Ordinary vs constrained at High | Ordinary won 16–2 (p = 0.00131) |
| Ordinary vs constrained at Max | Ordinary won 16–2 (p = 0.00131) |
The net Max advantage was +10 wins under ordinary generation and +6 under the constraint, an interaction of −4 wins across 18 cells (−0.222 per cell) in the opposite direction from preferential rescue. Max produced better-articulated proposals in both lanes, but it did not close the constraint gap.
Web-screen yield¶
| Arm | Survivors |
|---|---|
| Ordinary/High | 12/18 |
| Constrained/High | 11/18 |
| Ordinary/Max | 9/18 |
| Constrained/Max | 9/18 |
The Max outputs scored better in blinded proposal comparison but survived the light screen less often. This small, purposively balanced subset cannot show that greater effort causes more prior-art collision. It does show that internal proposal quality and researched survival are different objectives. More elaborate causal accounts may still describe established systems, and a novel-looking proposal can be operationally weak.
Priority problems revealed by the experiment¶
- Optimize quality within alternative substrates, not compliance alone. The instruction was understood. The central problem is preserving the archetype's causal structure without building a contrived physical analogue.
- Treat substrate feasibility as cell-specific. A hard ban is poorly matched to many computer-science and framed-archetype cells. A preliminary feasibility judgment could route a cell to physical, measurement, hybrid, or unrestricted generation rather than forcing the same rule everywhere.
- Do not let measurement become a universal fallback. A sensor or witness can reveal a problem without solving it. Future prompts should distinguish interventions that alter the causal process from instruments that merely observe it.
- Co-optimize against prior art. Max effort improved internal quality but not light-screen yield. Generation needs proposal-specific retrieval or a cheap evidence-bearing novelty check, not only more reasoning devoted to elaboration.
- Use diversification as a portfolio policy. The eight constrained-only survivors show a real option value. A minority quota or parallel alternative-substrate lane may add coverage without replacing the stronger ordinary lane.
- Concentrate expert review where the constraint was naturally aligned. Chemistry/materials and the constrained-only survivors are the best places to test whether alternative-substrate proposals contain real value rather than merely passing an automated screen.
What the result establishes—and what it does not¶
Experiment 12 supports three bounded claims about the tested prompt–model–corpus bundle:
- the governance/computational concentration is not a hard expressive ceiling;
- leaving those channels carries a large average internal-quality cost in this sample; and
- it carries a moderate, imprecisely estimated light-screen-yield cost while finding some different survivors.
It does not show that the base model has an intrinsic substrate preference. The effect could arise from the corpus, prompts, domain set, judge rubric, public-web evidence, model family, or their interaction. The control outputs were generated earlier rather than simultaneously. Calls were stochastic and had no fixed sampling seed. The web screen used four retained sources and can only detect obvious or nearby precedent; it cannot establish world novelty, commercial value, safety, or deployability. All semantic roles came from related model families, so blinding prevents label leakage but not shared priors.
The Max subset was small and selected to balance source-score bands, not to estimate a population effect. Reasoning effort is a product setting rather than a controlled unit of cognition. Its findings should remain descriptive.
Resource use¶
Experiment 12 completed 281 scientific model calls with no failed transport calls: 108 proposal-generation calls, 102 authorized web screens, 39 primary blinded measurement or adjudication calls, and 32 Max-follow-up blinded measurement or adjudication calls.
| Resource | Recorded amount |
|---|---|
| Input tokens | 44,256,895 |
| Cached input tokens | 29,580,032 |
| Uncached input tokens | 14,676,863 |
| Output tokens | 1,337,531 |
| Reasoning-output tokens | 660,600 |
| Summed call time | 43,522 seconds (12.1 hours) |
Summed call time includes concurrent calls and is not calendar duration. The web screens consumed most input tokens; Max constrained generation consumed more output and reasoning tokens than Max ordinary generation. Full phase accounting is in the runtime summary.
Reproducibility and artifacts¶
- Frozen method
- Original design seal
- README status-file audit correction
- Input manifest and frozen Max subset
- Primary machine-readable results and cell-level records
- Max machine-readable results
- Resource accounting
- Blinded primary inputs, sealed primary measurements, and revealed joins
- Blinded Max inputs and measurements
- Prompts, schemas, and run telemetry
The raw record also retains a pre-model adjudication runner key error and its resolution. It did not consume a model call or alter a scientific output.