Feigenbaum test¶
A proposed domain-specific variation of the Turing test in which a computer's performance is judged by whether it can reproduce the behavior or output of a recognized human subject-matter expert.
Core Idea¶
The Feigenbaum test is a proposed subject-matter-expert variant of the Turing test. Instead of asking whether conversation appears human in general, it asks whether a computer can reproduce the task performance or outputs of a recognized expert in a field such as chemistry or marketing.
A meaningful test needs a bounded domain, representative tasks, competent human comparators, controlled access to outputs, and explicit scoring. It evaluates behavior under that design, not consciousness, causal understanding, safety, or general intelligence. Edward Feigenbaum proposed the idea in 2003, and later futurist discussion broadened its visibility; the short source record supports the concept but not a universally standardized protocol.
Structural Signature¶
Sig role-phrases:
- bounded expert domain. Specifies the professional field and task family. Constitutive scope. If altered: General conversation is not the target.
- recognized human expert comparator. Provides a performance reference. Constitutive benchmark. If altered: Popularity or confidence is not expertise.
- machine system. Produces answers, decisions, or work products under comparable conditions. Constitutive candidate. If altered: Human assistance must be disclosed.
- blind or controlled evaluation. Limits evaluator access to source identity and aligns prompts or cases. Central validity condition. If altered: A product demonstration is not an imitation test.
- expert-equivalence judgment. Assesses whether outputs are indistinguishable or comparably competent under declared criteria. Identity-bearing outcome. If altered: Passing one task does not imply general intelligence.
What It Is Not¶
- Turing test. Is broad human imitation rather than expertise judged?
- Expert system. Is a system type being confused with an evaluation?
- Professional exam. Is expert output comparison and blinding present?
- AI benchmark. Is a recognized expert comparator included?
Scope of Application¶
Use Feigenbaum test for the named proposal and state field, tasks, expert selection, blinding, scoring, and inference limits.
- Artificial intelligence. Evaluates domain performance.
- Expert systems. Compares encoded expertise.
- Professional assessment. Defines expert benchmarks.
- Human-computer studies. Tests evaluator judgments.
- AI history. Tracks proposed intelligence tests.
Clarity¶
Expert-like output can reflect imitation, retrieval, or task competence without establishing the same internal understanding.
Manages Complexity¶
Results depend strongly on case sampling, expert disagreement, evaluator blinding, tool access, and scoring. One narrow pass should not be generalized beyond the tested field and conditions.
Abstract Reasoning¶
- Bound the subject-matter domain.
- Select qualified human comparators.
- Build representative controlled tasks.
- Blind and score outputs under one rubric.
- Limit conclusions to observed expert-level equivalence.
Knowledge Transfer¶
Comparator-based imitation tests transfer across professions, but the named expert-domain framing and historical proposal delimit the Feigenbaum test. The nearest stopping boundary is explicit: The ordinary Turing test is closest: it evaluates human-like conversational indistinguishability broadly, while the Feigenbaum variant narrows the comparator to field expertise. The inclusion test remains: An evaluation is a Feigenbaum test when a machine and recognized expert are compared on controlled tasks within a declared specialty for expert-level behavioral or output equivalence. The structure no longer applies when the case exits when no expert comparator, bounded subject domain, or controlled equivalence judgment is present.
Examples¶
Canonical¶
Chemists and a machine answer the same blinded structure-analysis cases; independent evaluators judge whether the machine's reports match expert quality under a preregistered rubric.
Mapped back: bounded expert domain → chemical analysis; recognized human expert comparator → qualified chemists; machine system → AI reports; blind or controlled evaluation → same cases and blinded outputs; expert-equivalence judgment → rubric-based comparison.
Applied / In Practice¶
A marketing model beats a historical click-rate baseline without comparison to expert work. It may be useful, but that performance benchmark is not a Feigenbaum test.
Mapped back: bounded expert domain → marketing; recognized human expert comparator → absent; machine system → prediction model; blind or controlled evaluation → historical benchmark; expert-equivalence judgment → not tested.
Structural Tensions¶
T1: observable equivalence vs. internal competence. Matching outputs does not reveal reasoning mechanism. Diagnostic: What exactly does a pass establish?
T2: expert standard vs. expert disagreement. Recognized specialists can disagree on hard cases. Diagnostic: How is the comparator distribution defined?
Structural–Framed Character¶
Description turns on bounded expert domain, recognized human expert comparator, machine system, blind or controlled evaluation, expert-equivalence judgment. Skeletal core. A candidate system is judged against a trusted reference performer on matched tasks under hidden identity. Domain-bound accent. Subject-matter experts, professional cases, machines, evaluators, and Turing-test history define the proposal. Transfer remains bounded because Why not prime. Comparator testing is portable; this is a named AI evaluation proposal. The negative boundary is concrete: Any benchmark, certification exam, chatbot conversation, expert system, automation demo, prediction contest, peer review, or ordinary Turing test is not automatically a Feigenbaum test. The test is evaluative-comparative: controlled expert and machine outputs support a bounded behavioral equivalence judgment. Its character: AI measured against specialist human performance rather than generic humanness.
Structural Core vs. Domain Accent¶
Skeletal core. A candidate system is judged against a trusted reference performer on matched tasks under hidden identity.
Domain-bound accent. Subject-matter experts, professional cases, machines, evaluators, and Turing-test history define the proposal.
Why not prime. Comparator testing is portable; this is a named AI evaluation proposal.
Instantiates / Related Primes¶
- Turing test. The proposal narrows imitation to expertise.
- Benchmark. A task set operationalizes the comparison.
- No strict parent is asserted.
Neighborhood in Abstraction Space¶
Feigenbaum test sits in a crowded region of the domain-specific corpus (37th percentile for distinctiveness): several abstractions share nearly its structure, so a description that fits it tends to fit its neighbors too.
Family — Group Dynamics & Collective Behavior (19 abstractions)
Nearest neighbors
- Clinical Equipoise — 0.89
- Face validity — 0.89
- Preventive action — 0.88
- Social comparison bias — 0.87
- Showdown Cooperative Learning — 0.87
Computed from structural-signature embeddings · 2026-10-08
Not to Be Confused With¶
- Turing test. Tell: Is broad human imitation rather than expertise judged?
- Expert system. Tell: Is a system type being confused with an evaluation?
- Professional exam. Tell: Is expert output comparison and blinding present?
- AI benchmark. Tell: Is a recognized expert comparator included?
References¶
- Frozen Wikipedia discovery revision: https://en.wikipedia.org/wiki/Feigenbaum_test (revision 1269540482).
- Preserved source candidate: https://archive.org/details/singularityisnea00kurz
The frozen Wikipedia revision is discovery provenance. The retained source set was reviewed for identity, formal or operational relation, and scope. The encyclopedia's structural synthesis is bounded to those claims; a thin authority surface is recorded as a nonblocking source-strengthening repair rather than concealed.