Skip to content

Measuring Structural Transfer with Executable Kernels

A cross-domain analogy benchmark for language models, built on computed ground truth rather than asserted labels

A benchmark-design and evaluation-methodology paper. Draft — 2026-07-11. Companion to the Abstractopia foundational paper: where that paper describes an instrument for training and measuring cross-domain structural transfer in people, this one turns the same idea on language models. A runnable five-kernel proof-of-concept accompanies the paper (see Downloads; §8). This is a measurement contribution — it aims to make one consequential capability, genuine structural generalization versus surface pattern-matching, legible; it does not aim to advance model capabilities. Every empirical or bibliographic claim has been verified against its primary source.


Abstract

Language models solve many analogy problems zero-shot, which some read as emergent abstract reasoning [Webb, Holyoak & Lu, 2023]. Whether that reflects structural competence or surface shortcut is contested — counterfactual variants dissimilar from pretraining data sharply degrade model performance while humans hold steady [Lewis & Mitchell, 2024], a finding the original authors in turn dispute [Webb, Holyoak & Lu, 2025]. This live disagreement is exactly what a benchmark should adjudicate, and it cannot be adjudicated by a benchmark whose own labels are surface artifacts [Gururangan et al., 2018; McCoy, Pavlick & Linzen, 2019]. We propose a benchmark built on three commitments that answer those threats structurally: compute the ground truth instead of asserting it, measure invariances instead of accuracies, and treat difficulty as a measured continuous variable instead of a filtered binary. Each prime's structural signature is formalized as a small executable kernel; a vignette is a rendering of a kernel into a domain, so the answers to the benchmark's questions — does this share structure with that, who plays which role, what happens if you relieve this — are computed by running the kernel, not decided by an author. The atomic unit is a kernel × surface lattice whose two axes are the surface-swap (structure held, surface varied) and the structure-swap (surface held, sibling kernel substituted from the Encyclopedia's DAG). The headline metrics are a structural- invariance index and a surface-pull curve — model accuracy as a function of how hard surface features pull toward the wrong answer — which are interpretable at any overall accuracy and immune to the ceiling compression that cripples a far-minus-near gap. We report a runnable proof-of-concept: five primes formalized as generative kernels reproducing their signatures, one deliberately-chosen framed prime that would not become a kernel (locating the method's boundary), and a demonstration that the two metrics cleanly separate a structure-tracker from a surface-matcher on fiat-free computed ground truth. We are explicit that the approach measures transfer on the formalizable subset of the corpus and excludes the framed majority — a narrower construct than the corpus as a whole, and the honest price of rigor.


1. The question

There is a difference between recognizing a structure and recognizing a surface, and it is the difference that matters most for generalization. A model that has learned "a proxy placed under optimization pressure loses its link to the thing it measured" should recognize that pattern in a school teaching to a test, a fishery rewarded on tonnage, and a platform optimizing engagement — three situations with no shared vocabulary. A model that has instead learned that certain surfaces (metrics, quotas, dashboards) co-occur with certain answers will succeed on familiar surfaces and fail on unfamiliar ones, while producing identical-looking outputs on the cases where the two strategies happen to agree. Whether a model "has" an abstraction is, operationally, whether it tracks the structure or the surface — and that is answerable only by a test designed to pull the two apart.

Models plainly have some of this ability. Webb, Holyoak, and Lu found GPT-3 solved a broad range of analogy problems zero-shot, matching or exceeding humans in most of their settings [Webb, Holyoak & Lu, 2023]. But a capability shown on a task format is not a competence that transfers, and the gap between the two is where measurement has repeatedly failed.

2. Prior benchmarks, the counterfactual dispute, and the remaining gap

Four lines of work frame the problem and, together, locate a real gap.

Abstraction benchmarks exist, in idealized formats. Chollet reframed the target from skill on known tasks to efficiency of skill acquisition under novelty, and instantiated it in the Abstraction and Reasoning Corpus: grid puzzles where a solver induces a transformation from a few examples using core-knowledge priors [Chollet, 2019]. ConceptARC extended this to probe whether a system has grasped a concept rather than merely solved instances [Moskvichev, Odouard & Mitchell, 2023]. These are the right spirit but live in a visual-grid microworld; they do not test whether a model recognizes the same abstract structure across the natural-language surfaces of different real-world domains.

Natural-language analogy benchmarks exist, and one is close prior art. ARN builds 1.1k narrative triples explicitly operationalizing the cognitive distinction between surface and system similarity, with near/far partitions of analogies and disanalogies, and finds a clear model–human gap — the largest models still struggle with far analogies zero-shot [Sourati et al., 2024]. Our benchmark shares ARN's surface-vs- system organizing principle and its natural-language, cross-domain framing; we say so plainly and build on it. What ARN does not have, and what constitutes our gap, is below.

The validity threat is documented, and it is surface exploitation. A classifier seeing only the hypothesis of a natural-language-inference pair — never the premise — labels roughly two-thirds of SNLI correctly, because the collection protocol left surface artifacts [Gururangan et al., 2018]. HANS then showed strong models collapse on a controlled set where shallow syntactic heuristics are made to fail — "right for the wrong reasons" [McCoy, Pavlick & Linzen, 2019]. A benchmark that does not engineer against surface shortcuts measures the shortcut.

The dispute is live, and it is about exactly this. Lewis and Mitchell took the letter-string analogies Webb et al. used and built counterfactual variants — permuted or non-letter alphabets, dissimilar from pretraining — on which humans stay high while GPT models decline sharply, evidence of retrieval over reasoning [Lewis & Mitchell, 2024]. Webb, Holyoak, and Lu replied, reporting that models do generalize to counterfactual variants when allowed to write and execute code, and arguing against the mimicry reading [Webb, Holyoak & Lu, 2025]. The current draft of this paper cited only the 2023 opening move; the honest position is that the field has an unresolved empirical question, and the value of a benchmark is to adjudicate it rather than to add another disputed data point.

The gap that survives contact with all of this. No existing benchmark grades correctness against an external structural ontology with a known sibling topology, perturbs structure while holding surface fixed (ARN and the counterfactual work perturb surface; almost nobody perturbs structure), builds conflict items where surface recruits a sibling against the true structure, and — the load-bearing move — computes its ground truth rather than asserting it per item. Those are our differentiators, and each is a response to a specific failure above.

3. Three commitments

The rest of the design follows from three commitments, each aimed at a threat in §2.

First, compute ground truth, don't assert it. Every item that embeds an author's claim "this vignette instantiates P" carries item-level fiat, the tightest saturation loop. We remove it by formalizing signatures as runnable kernels and rendering vignettes from them (§4–§5), so match/role/intervention answers are computed.

Second, measure invariances, not accuracies. Raw accuracy is the number most easily inflated by scale and memorization and least diagnostic of structure. We headline paired-edit invariances and a curve, each of which uses an item as its own control (§7).

Third, parameterize difficulty, don't filter it. Selecting items that current models fail bakes in the first result and churns the item set across releases, destroying longitudinal comparability. We make difficulty a measured axis — surface pull and structural depth — and freeze a scoring core (§7, §9).

4. The substrate: executable kernels

The Encyclopedia of Abstractions is a curated corpus of ~1,300 primes, each a substrate-neutral structural pattern (roles and the relations among them) with a hierarchy of prerequisite and composition edges. In the original design the corpus was an external answer key. That over-claims: fixing the label space and the sibling topology (genuinely valuable, especially for distractor construction) does not fix the mapping from a particular vignette to a label. So the corpus's role shifts from answer key to generator schema. Each prime's signature is formalized as a small executable kernel — a typed relational template with dynamics — and a vignette is a rendering of a kernel into a domain. Feedback is a pair of coupled variables with a sign structure you can run; Goodhart is a proxy–construct coupling whose correlation decays as optimization pressure on the proxy rises; constraint is a binding min(); regression to the mean is selection-on-extremes over a noisy measurement of a stable latent. The interesting answers — does this match that, who plays which role, what happens if you relieve X — are then computed from the kernel, and item-level fiat disappears for every task where it matters.

Writing kernels is expensive, so the corpus tiers honestly, and the tiering is itself a finding. It is a mechanical, falsifiable version of the structural-versus-framed distinction the Encyclopedia has curated by hand: a prime is Tier-1 structural iff its signature runs.

  • Tier-1a — generative kernels you can run forward and intervene on (feedback, goodhart, constraint, regression to the mean, selection bias, and the causal / dynamical / game-theoretic family generally).
  • Tier-1b — predicate kernels with computable ground truth but no forward dynamics: you can test a deviation from a rule but there is no state to evolve, so intervention-prediction degenerates into re-checking the predicate. Sunk cost is the worked example (§8): it is not a system to run but a departure from a normative rule.
  • Tier-2 — framed primes where even the predicate needs domain semantics. This is the corpus majority — the ~160 framed primes an operator-driven survey of the corpus flagged as carrying heavy home-domain baggage — and it is the honest limit of the method.

The consequence must be stated up front rather than buried: this benchmark measures structural transfer on the Tier-1 (runnable) subset and excludes the framed majority — arguably the region where surface-matching does the most real-world damage. That is a deliberate trade of breadth for rigor, and it belongs in the abstract, not the limitations. A side benefit accrues regardless of whether the benchmark ships: writing kernels is a forcing function for signature rigor, and any prime whose signature resists execution had a signature doing less work than it appeared to — useful information for the Encyclopedia independent of this project.

5. The atomic unit: a kernel × surface lattice

Retire the single vignette as the unit of measurement. The atomic unit is a kernel × surface lattice: one kernel rendered into several surface domains, plus sibling kernels rendered into the same domains under a lexical-overlap constraint. The two minimal-pair edits are the lattice's axes, and both are quantified.

  • S-swap (structure held, surface varied): the same kernel rendered into a new far domain. The correct answer must stay the same. Surface distance is measured, not asserted — embedding distance between renderings and domain n-gram frequency against a reference corpus — which repairs the least-principled quantity in the original design (author-intuited "near/mid/far").
  • T-swap (surface held, structure varied): a sibling kernel from the DAG rendered into the same domain, constrained to maximize token overlap with the original. Defining a T-swap as substituting a DAG sibling solves the "flip to what?" problem mechanically: the answer always flips to a determinate prime, never to an undefined "not-P".
  • Role permutation (nearly free): same domain, same entities, same vocabulary, but the entities' role bindings in the kernel are permuted (the sensor becomes the actuator; the selected population becomes the selecting mechanism). This isolates role binding from entity association — the crux of structure-mapping in Gentner's sense [Gentner, 1983] — and costs almost nothing once kernels exist.

Rendering is done by a diverse ensemble of models with style randomization, and surface domains are rotated across primes so that, corpus-wide, no domain, register, or stylistic feature correlates with any label. Independence is enforced by construction and then verified by probes (§7).

6. The task battery: ontology-free first

The primary tasks never require the model to know the taxonomy's names, which severs the taxonomy-alignment confound — a model could have a coherent but differently- factored structural vocabulary and be wrongly scored against ours.

Matching (triads / odd-one-out). Three vignettes: two share a kernel across distant surfaces; the third shares surface with the query but runs a sibling kernel. This is ARN's format [Sourati et al., 2024], built on rather than around, and it has deep precedent in the analogy literature [Gick & Holyoak, 1983].

Role mapping. Given a source vignette and a far-domain target, produce or select the role correspondence ("which element here plays the role the thermostat plays there?"). Classification can be faked by co-occurrence; alignment is much harder to fake because it requires the relational skeleton with bindings intact. This is the workhorse task.

Intervention prediction. "If the constraint is relieved / the proxy is decoupled / the selection step removed — what happens?" Forced choice over outcomes, the correct answer computed by running the modified kernel. One guard is mandatory: a no-vignette baseline — pose the question with the vignette withheld. If domain lore alone answers it above chance, the item is invalid, because it can be solved without extracting any structure. This gate kills the domain-prior shortcut classification quietly permits.

Break-it edits (justification made checkable). Instead of scoring free-text justifications with an LLM judge — which inherits precisely the surface biases the benchmark targets — ask "which one of these candidate edits would make this scenario no longer an instance of the pattern it shares with the reference?" Correctness is computed: apply the edit to the kernel and check whether the signature survives. A model that names the load-bearing relation by severing it has demonstrated what a rubric can only approximate.

Compositional load-bearing. Couple two or three kernels, render, and ask which intervention moves the outcome. "Load-bearing" stops being theory-laden because it is literally executed — run each candidate relief and observe. Generation of such an item is a single forward pass; solution requires search over candidate structures, a genuine asymmetry that resists saturation for a different reason than conflict items do.

Classification against Encyclopedia labels survives as a secondary diagnostic track — valuable for error taxonomy via the sibling topology, and for the corpus — but no headline claim rests on it.

7. Metrics: an invariance index and a psychometric curve

Structural-invariance index (SSI). From the lattice, define the correct-flip rate under T-swap and the spurious-flip rate under S-swap, and report

SSI = P(correct flip | T-swap) − P(spurious flip | S-swap).

A perfect structure-tracker approaches +1 (flips when structure changes, holds when only surface changes); a pure surface-matcher goes negative (holds when the lexical twin changes structure, flips when a far-domain rendering keeps it). Each lattice cell is its own control, so memorizing an individual item's answer buys nothing.

Surface-pull curve. Take the artifact probes the original design left implicit — a vignette-only shallow classifier, an options-only model, an embedding-similarity retriever — and promote them from validity gates to the x-axis of the primary readout. Each probe assigns every item a surface pull: how strongly, and toward which answer, surface features alone point. Plot the full model's accuracy on the structure tasks as a function of surface pull toward the wrong answer. A genuine structure- tracker's curve is flat; a surface-matcher's tracks the probe downward. The headline statistics become the slope and the breakpoint — both continuous, both interpretable at any overall accuracy, both immune to the ceiling and floor compression that make a far-minus-near difference score unreliable and interpretable only in a mid-accuracy band one cannot guarantee. "Surface" is defined relative to the probe family, so an ensemble of probes is used and robustness reported across them.

Intervention accuracy at distance. The functional-transfer number: accuracy on intervention-prediction items at high measured surface distance.

The camouflage-vs-conflict hypothesis, as a curve-shape prediction. The original paper's §5.1 claimed plain surface "camouflage" is easy while structure–surface conflict is hard, and framed this as a binary comparison of hand-labeled item families. The curve subsumes it: a strong model flattens the low-pull region first and retains negative slope at high pull. That is sharper, falsifiable, and does not depend on hand-labeling which items are "conflict."

Required validity gates. Before an item ships, all mechanical: the shallow probe at chance on labels given the vignette only; the options-only model at chance (or the candidate-set topology leaks the answer); the no-vignette baseline at chance on interventions; a stylometric classifier at chance on renderer/label pairings (with a small author pool, per-prime authorial tics are annotation artifacts at the label level). The transfer gap survives only as a secondary, historically-interpretable measure, never the headline.

8. Proof of concept

Because committing a paper to the kernel approach is expensive and narrows the construct, we ran a one-day proof-of-concept before adopting it. It answers two questions: can the signatures be made executable, and where does the method break; and do the invariance metrics actually separate a structure-tracker from a surface-matcher on computed ground truth. The runnable code, results, and figure are downloadable (§Downloads; benchmark-kernel-poc).

Formalization. Five primes chosen for clean signatures all became runnable generative kernels whose forward dynamics reproduce their textbook behavior: cutting either coupling of feedback opens the loop and the signature fails; goodhart's proxy–construct correlation falls from 0.90 to 0.08 under optimization pressure and rewarding the construct directly breaks it; constraint's output moves when the binding resource is relieved (4→7) and not when a non-binding one is (4→4); regression_to_the_mean shows an effect of 1.48 with measurement noise and 0.00 without, and a control group selected the same way regresses identically (1.43 ≈ 1.48) — the tell against cause; selection_bias returns a biased estimate of 1.03 against a random-sample 0.02.

The boundary, deliberately probed. A sixth, framed prime — sunk_cost — was included as a stress test and would not become a generative kernel. Its ground truth is still computable ("does the decision track prior investment against prospective value?"), so fiat is still avoided, but there is nothing to run forward; its run() and intervene() raise by design. This is the Tier-1b boundary of §4, found by construction rather than asserted.

Metric separation. On a mini-lattice where T-swaps are genuine lexical twins (token overlap 0.27–0.67 with their base) and the structurally-correct S-swaps are lexically distant (overlap 0.00) — i.e. built to actively mislead a surface-matcher — two synthetic solvers took the same items. A structure-tracker (answering via the kernel signature) scored SSI = +1.0 with a flat surface-pull curve (slope +0.000). A surface-matcher (answering via token overlap) scored SSI = −1.0 with a curve collapsing from 1.00 to 0.00 as pull turned positive (slope −1.15). The instrument cleanly separates the extremes, and — the point — every answer was computed from kernel ground truth with no per-item human label.

The proof-of-concept deliberately does not test real models: its two solvers are bracketing cases that prove the instrument separates a perfect structure-tracker from a pure surface-matcher. Where actual language models fall on the SSI and slope axes is the empirical question the benchmark exists to answer, and is untouched here.

9. Saturation and contamination: publish the generator, freeze a core

Because ground truth is computed, the benchmark is a distribution, not a set. Publish the kernel library, the rendering pipeline, and a public development split. Keep a frozen private core scored identically forever — the longitudinal instrument, never re-filtered, so cross-generation claims survive. For each evaluation, mint fresh items from held-out kernel × surface combinations after the model's training cutoff, with canaries in anything public. This replaces the per-release empirical difficulty filter — which selects for a t=0 "models fail" result, churns the item set, and yields adversarial items that may not transfer across model generations — with difficulty that comes from the measured pull and structural-depth dials instead. The invariance metrics are themselves contamination-robust: a leaked item's memorized answer does not survive a structure-swap variant minted yesterday. Where empirical filtering is used at all, it is two-tier — a frozen core plus a rotating adversarial frontier, reported separately — and a failed item is adjudicated for why it failed (its computed break-it edit and the model's justification are the instrument) before admission, so "hard" is never confused with "broken."

10. Human validation in one protocol

Human validation collapses into a single task: naive solvers with no Encyclopedia training take the items, and their answers are compared to kernel output. (People trained on the ontology would share the author's carving of concept space and could only validate internal consistency, not construct validity.) Agreement above threshold certifies three things at once — rendering fidelity (the text faithfully realizes the kernel), item solvability (the structural evidence is dispositive, not a trick, handling the conflict-item hazard that surface cues are real probabilistic evidence), and the human baseline. It also inoculates against the objection at the center of the Webb–Lewis/Mitchell dispute — that hard variants might be hard for construct-irrelevant reasons like comprehension load: if naive humans recover the kernel's answers from a far-domain rendering, that rendering is certified comprehensible, and a model's failure cannot be waved off as vocabulary tax.

11. The generation/solution asymmetry, restated honestly

The original §5.1 argued that plain camouflage is not reliably hard for a capable model because generation (structure → surface) and recognition (surface → structure) draw on the same representation, so difficulty must live at structure–surface conflict. The premise needs care. West and colleagues document a generative AI paradox: models can generate content whose discriminative counterpart they cannot reliably handle, with hard negatives dropping a strong model from near-ceiling to well below it [West et al., 2024]. Generation and discrimination dissociate — which both strengthens and complicates the argument. It strengthens the claim that camouflage-generation is cheaper than adversarial recognition; but it undercuts the specific inference "a model that can generate P can therefore recognize P," because a model might render a clean camouflage item it then cannot invert. The defensible position drops the shared- representation premise: the real hardness is misleading-surface resistance (a bias the model exploits in both directions, so a surface that misleads will mislead it) and compositional search, and both are now measured as a curve shape (§7) rather than asserted. This is the honest residue of the "hash" intuition — not one-wayness, but resistance to a shared, surface-driven bias.

12. Threats to validity and limitations

Held to the standard the project teaches, here is where the design is vulnerable.

The kernel commitment narrows the construct to the formalizable subset. Tier-2 framed primes — the corpus majority, and arguably where surface-matching does the most real-world damage — escape measurement. This is the central limitation and is stated in the abstract, not hidden here.

Rendering pipelines have regularities a strong model could learn; renderer diversity and the stylometric gate are an arms race, not a solution. The surface-pull axis inherits its probes' blind spots, mitigated but not removed by a probe ensemble. Intervention items constrain vignettes toward semi-formal completeness, taxing length and naturalism. Contamination persists as a threat despite fresh minting and canaries, though the invariance metrics degrade it more gracefully than raw accuracy. The ontology's coverage and boundary cases inject label noise near sibling boundaries; items below an inter-annotator agreement threshold, set by multi-annotator labeling in the construction protocol (not as an afterthought), are cut. And a low SSI or a flat curve can be a true positive: if a model genuinely represents these structures it should post them, and that is a real result, not a benchmark failure — the design's value is as a diagnostic that ages gracefully, having established on a controlled, contamination-guarded, adversarially-validated test that a model tracks structure over surface, which current benchmarks cannot license.

13. Why it matters

Measurement drives progress, and the thing worth measuring is a failure mode: reliance on surface similarity in place of structural understanding, which shows up downstream as brittle generalization, over-application of familiar patterns to ill-fitting cases, and confident errors that look exactly like competence. An invariance-and-curve diagnostic exposes that failure mode directly, on natural-language content, against computed structural ground truth, with the shortcut engineered out. It is evaluation methodology — it measures whether models generalize structurally; it does not teach them to — and its orientation is safety-positive.

It also nearly completes a symmetry with the human instrument. Abstractopia asks whether people can be trained to see structure through surface and to know where it breaks; this benchmark asks whether models do, on the same primes, the same surface- distance axis, the same sibling distractors. The comparison is not quite "apples to apples" — models have gradient-level familiarity with every published exposition of these structures while humans take them cold — so it is best reported as the same instrument applied to different exposure histories, which is itself a research affordance: it locates where machine structural transfer diverges from human.

14. Conclusion

We have proposed a benchmark for cross-domain structural transfer in language models built on three commitments — compute ground truth, measure invariances, parameterize difficulty — realized through executable kernels, a kernel × surface lattice, ontology- free tasks, and two headline metrics (a structural-invariance index and a surface-pull curve) that a runnable proof-of-concept shows cleanly separate structural from surface strategies on fiat-free ground truth. Its central debt is to the lesson that a score is only as trustworthy as its resistance to shortcuts [Gururangan et al., 2018; McCoy, Pavlick & Linzen, 2019], its central differentiator is perturbing structure while holding surface, and its central honesty is that it measures the formalizable subset of structural transfer and says so. The single highest-leverage component is the kernels — and conveniently, that is the one that pays dividends back into the Encyclopedia whether or not the benchmark ever ships.


References

Barnett, S. M., & Ceci, S. J. (2002). When and where do we apply what we learn? A taxonomy for far transfer. Psychological Bulletin, 128(4), 612–637. https://doi.org/10.1037/0033-2909.128.4.612

Chi, M. T. H., Feltovich, P. J., & Glaser, R. (1981). Categorization and representation of physics problems by experts and novices. Cognitive Science, 5(2), 121–152. https://doi.org/10.1207/s15516709cog0502_2

Chollet, F. (2019). On the measure of intelligence. arXiv:1911.01547. https://arxiv.org/abs/1911.01547

Gentner, D. (1983). Structure-mapping: A theoretical framework for analogy. Cognitive Science, 7(2), 155–170. https://doi.org/10.1207/s15516709cog0702_3

Gick, M. L., & Holyoak, K. J. (1983). Schema induction and analogical transfer. Cognitive Psychology, 15(1), 1–38. https://doi.org/10.1016/0010-0285(83)90002-6

Gururangan, S., Swayamdipta, S., Levy, O., Schwartz, R., Bowman, S. R., & Smith, N. A. (2018). Annotation artifacts in natural language inference data. In Proceedings of NAACL-HLT 2018 (pp. 107–112). https://aclanthology.org/N18-2017/

Lewis, M., & Mitchell, M. (2024). Using counterfactual tasks to evaluate the generality of analogical reasoning in large language models. arXiv:2402.08955. https://arxiv.org/abs/2402.08955

McCoy, R. T., Pavlick, E., & Linzen, T. (2019). Right for the wrong reasons: Diagnosing syntactic heuristics in natural language inference. In Proceedings of ACL 2019 (pp. 3428–3448). https://aclanthology.org/P19-1334/

Moskvichev, A., Odouard, V. V., & Mitchell, M. (2023). The ConceptARC benchmark: Evaluating understanding and generalization in the ARC domain. Transactions on Machine Learning Research. https://arxiv.org/abs/2305.07141

Sourati, Z., Ilievski, F., Sommerauer, P., & Jiang, Y. (2024). ARN: Analogical reasoning on narratives. Transactions of the Association for Computational Linguistics, 12, 1063–1086. https://doi.org/10.1162/tacl_a_00688

Webb, T., Holyoak, K. J., & Lu, H. (2023). Emergent analogical reasoning in large language models. Nature Human Behaviour, 7, 1526–1541. https://doi.org/10.1038/s41562-023-01659-w

Webb, T., Holyoak, K. J., & Lu, H. (2025). Evidence from counterfactual tasks supports emergent analogical reasoning in large language models. PNAS Nexus, 4(5), pgaf135. https://doi.org/10.1093/pnasnexus/pgaf135

West, P., Lu, X., Dziri, N., Brahman, F., Li, L., Hwang, J. D., Jiang, L., Fisher, J., Ravichander, A., Chandu, K., Newman, B., Koh, P. W., Ettinger, A., & Choi, Y. (2024). The generative AI paradox: "What it can create, it may not understand." In Proceedings of ICLR 2024. https://arxiv.org/abs/2311.00059