Evaluation¶
Core Idea¶
Evaluation applies a criterion-bearing frame to a bounded object, interprets the object's relevant observations or features against that frame, and produces a result that counts as a verdict, score, rank, grade, priority, or action-guiding judgment. The evaluator may be a person, a group, an institution, or a rule-governed procedure. The criteria may be explicit in a rubric, specification, objective, threshold, or reference class; they may also be reconstructable from repeated judgments. Without a standard that makes some features relevant to an evaluative purpose, however, there is description or reaction rather than evaluation.
The abstraction is a relation among four roles: object, criterion frame, relevant observations, and evaluative result. A teacher grades an essay, a panel assesses a proposal, a classifier scores an input, a court judges conduct under a standard, and an engineer reviews a design against safety and performance goals. The vocabulary is not being borrowed metaphorically across these cases. Each maps features of a bounded target through a criterion-bearing frame into a result that can guide acceptance, revision, ranking, or action.
Evaluation is therefore more than a statistics or institutional-review concept. Remove a particular language, profession, or documentary genre and the relation remains. What does not survive removal is the local meaning of the criteria and verdict: a pass, diagnosis, risk tier, aesthetic judgment, and predicted class are not interchangeable outcomes even though their role in the evaluative structure is the same.
Structural Signature¶
a bounded object — an evaluator or rule-governed procedure — a criterion-bearing frame — relevant observations or features — a comparison or interpretive mapping — an evaluative result — a traceable route from frame and observations to result
- Bounded object: a claim, artifact, action, person, proposal, system, state, or candidate is delimited as the target of judgment.
- Evaluator or procedure: some agent, collective, institution, or executable rule performs the mapping.
- Criterion frame: a threshold, standard, purpose, prototype, objective, rubric, comparison set, or weighted set of considerations determines what counts.
- Relevant observations: features, measurements, testimony, experience, reasons, or evidence are selected as inputs under that frame.
- Relational reading: the observations are interpreted or compared against the criterion rather than merely listed.
- Evaluative result: the operation yields a verdict, score, rank, grade, priority, classification, recommendation, or other judgment with action-guiding meaning.
- Auditability: even if the rationale is not recorded, the result claims a route from criteria and observations; where that route cannot be reconstructed, the output approaches unsupported preference or arbitrary labeling.
Remove the bounded target and the judgment has no object. Remove the criterion and there is no basis for relevance or valence. Remove the mapping and the result is detached from the inputs. Remove the evaluative result and only observation, measurement, or comparison remains.
What It Is Not¶
- Not description. Description records features without declaring what those features count for under a purpose-bearing frame.
- Not measurement. Measurement maps an attribute to a scale, yielding a value and uncertainty. Evaluation interprets one or more measured or qualitative features against criteria and gives the output judgment meaning.
- Not comparison alone. Comparison can report that A exceeds B or differs on dimension d. Evaluation makes that relation count toward a verdict, rank, score, or recommendation.
- Not verification. Verification is the narrower conformance evaluation in which a stated specification is fixed and a defined procedure yields evidence and an accept, reject, or qualified verdict.
- Not decision. A decision commits to one alternative and forecloses others. An evaluation can inform that commitment, rank the alternatives, or recommend action without itself authorizing the choice.
- Not review artifact. A review persists an attributable evaluation, verdict, and warrant as an addressable record. Evaluations can be transient, automated, private, or unrecorded.
- Not arbitrary preference expression. Saying “I like it” may report a state. It becomes an evaluation when a target, relevant considerations, and the basis on which the expression counts as a judgment can be recovered.
- Not every selective process. Bare differential retention can occur with no criterion-bearing result. A selection rule instantiates evaluation only when candidates are mapped against a reference, objective, or fitness condition into an evaluative output; otherwise Selection is the cleaner abstraction.
Broad Use¶
In education, rubrics map demonstrated work against learning criteria into feedback, grades, and mastery judgments. In science, reviewers and model-comparison procedures assess evidential support, methodological quality, explanatory fit, or predictive adequacy. In medicine, observed signs, measurements, histories, and risks are evaluated against diagnostic or treatment criteria. In engineering and design, artifacts are judged across performance, safety, reliability, usability, cost, and maintainability.
In law and governance, conduct, claims, and policies are evaluated under legal rules, institutional purposes, and competing public values. In organizations, candidate panels, investment committees, risk processes, procurement systems, and program reviews turn heterogeneous observations into scores, tiers, recommendations, or rankings. In computing, test harnesses, scoring functions, classifiers, moderation policies, search rankers, and optimization procedures implement rule-governed evaluative mappings.
Across these settings the criteria, observations, aggregation rules, and result types vary. The invariant is that a bounded target is read through a criterion-bearing frame and a result is produced whose meaning is evaluative rather than merely descriptive.
Clarity¶
Evaluation separates observing from judging. A thermometer can report 39°C; a clinical protocol evaluates that measurement together with symptoms and context as evidence of fever, urgency, or treatment need. A benchmark can report response time; a service review evaluates the value against a service-level objective. The numerical value does not contain the criterion that makes it acceptable or dangerous.
It also separates the act from its packaging and consequences. A review is a persisted artifact that contains an evaluation and warrant. A decision uses evaluative results to commit to a course. Verification is one disciplined evaluation form. These distinctions make it possible to ask whether a dispute concerns observations, criteria, aggregation, judgment, documentation, or authority rather than calling every stage “the assessment.”
Manages Complexity¶
Evaluation compresses a high-dimensional target and a potentially plural standard into an operable result. Rubrics, scorecards, test suites, diagnostic protocols, review panels, and multi-criteria models make the mapping explicit. Once the roles are exposed, disagreement can be localized: parties may share observations but reject the weights, share criteria but dispute evidence quality, accept individual scores but reject the aggregation rule, or accept the analysis while denying the evaluator's authority.
The compression is necessarily lossy. A grade does not preserve the essay; a risk tier does not preserve the causal model; a ranked list does not preserve every trade-off. Auditability therefore requires retaining enough of the criterion frame and input trace to reconstruct why the result was produced. When only the headline result survives, downstream users may mistake a purpose-relative judgment for a context-free fact.
Abstract Reasoning¶
Let (x) be the bounded object, (C) a criterion frame, (O(x)) the relevant observations selected under that frame, and (E) the evaluative procedure. The result is (r=E(x,O(x),C)). This representation makes three kinds of sensitivity explicit.
First, criterion sensitivity: (x) and (O(x)) can remain fixed while a change in purpose, threshold, comparison class, or value weights changes ®. Second, observation sensitivity: the criteria can remain fixed while new evidence or measurement changes the result or its confidence. Third, procedure sensitivity: the same observations and criteria can yield different results when aggregation, ordering, compensation, or veto rules change.
The reasoning sequence is therefore: bound the object; surface the criteria; identify which observations the frame renders relevant; inspect the comparison or interpretation rule; name the result type; and determine what uses the result licenses. If a supposed evaluation has no recoverable criterion, its authority rests on hidden framing. If it has criteria but no evaluative output, it is a protocol or comparison still awaiting judgment.
Knowledge Transfer¶
A grading rubric teaches the design reviewer to separate criteria from global impression. A software test suite teaches the policy evaluator that multiple checks require an explicit aggregation rule before they become an overall verdict. Peer review teaches the classifier designer that reproducible output can remain difficult to contest when no warrant or feature trace is exposed. Clinical assessment teaches organizational panels to distinguish evidence quality from decision threshold.
What transfers is not the local standard but the audit grammar: object, criterion frame, observations, mapping, result, and authorized use. This grammar lets techniques move while preserving their limits. A weighted scorecard can organize both supplier assessment and treatment selection, but the weights remain domain commitments; the shared form never makes those commitments interchangeable.
Examples¶
Formal / abstract¶
Three candidates have feature vectors (O(x_i)). A criterion frame supplies normalized dimensions, weights, a veto condition, and an acceptance threshold. The evaluator computes an aggregate score but first rejects any candidate violating the veto. The output is a rank plus acceptability verdict. Changing the weights can reverse the rank; changing the threshold can alter acceptance without altering the rank; removing the criterion frame leaves only feature vectors and pairwise relations. The example exposes evaluation as more than scoring: the result's meaning comes from the frame and decision-relevant output type.
Applied / industry¶
A committee evaluates proposals for public funding. It observes feasibility, expected benefit, cost, distributional effect, uncertainty, and execution risk. A rubric makes the criteria and weights explicit; reviewers interpret submitted evidence against each criterion; the panel produces ratings and a ranked recommendation. The evaluation is not identical to the measurements, comparisons, review documents, or final appropriation decision. It is the criterion-governed judgment that connects those stages.
Structural Tensions¶
T1 — Explicit criteria versus tacit expertise. Formal rubrics improve comparability and auditability, while expert judgment can notice qualities no rubric anticipated. Diagnostic: identify which dimensions may be used tacitly and require evaluators to surface them when they alter the result.
T2 — Compression versus fidelity. A single score travels easily but erases trade-offs and uncertainty. Diagnostic: preserve a profile or warrant when materially different feature configurations can produce the same aggregate.
T3 — Consistency versus criterion validity. A procedure can apply the wrong criterion perfectly. Diagnostic: audit both reliable application and whether the frame addresses the actual purpose; the latter is the Validation question.
T4 — Comparability versus context. Standardized criteria enable aggregation across objects but can suppress legitimate local differences. Diagnostic: declare which contextual adjustments are permitted and whether they change the frame or only the observations.
T5 — Neutral procedure versus embedded values. Automated scoring can make evaluation appear objective while criteria, labels, thresholds, and loss functions encode priorities. Diagnostic: trace every apparently technical parameter to the judgment it operationalizes.
T6 — Result production versus result authority. A procedure can generate a score without having standing to impose consequences. Diagnostic: separate whether the evaluation is coherent from whether the evaluator is authorized and whether a later decision should rely on it.
Structural–Framed Character¶
Evaluation is mixed-framed. Its object-criterion-observation-result relation is structural, travels intact, and can be implemented by nonhuman procedures. Yet the selection of relevant features, governing criteria, weights, thresholds, and result semantics is purpose-relative and often evaluatively loaded. The abstraction does not pretend those commitments disappear; it makes their position in the structure inspectable.
Substrate Independence¶
The abstraction survives removal of a particular profession, institution, language, or quantitative scale. Replace an essay with a bridge design, a legal claim, a clinical state, or a model output; replace a panel with a scoring rule; replace a grade with a risk tier or pass/fail verdict. The same roles remain.
It does not survive removal of the criterion-bearing frame. A physical process that merely changes state, or a selective process that merely retains some outcomes, is not thereby evaluating them. Calling it evaluation is justified only when the process instantiates a reference or objective and produces an output that plays an evaluative role. That boundary keeps broad transfer from dissolving into metaphor.
Relationships to Other Abstractions¶
Current abstraction Evaluation Prime
Parents (1) — more general patterns this builds on
-
Evaluation presupposes Comparison Prime
Evaluation presupposes Comparison because judging an object requires placing its relevant features and a criterion or reference in a shared frame.Comparison supplies the relation-reading operation inside every evaluation: the object's relevant feature is co-framed with a threshold, prototype, alternative, rule, purpose, or other criterion so that a relation can be read. Evaluation adds the purpose-relative criterion and an output that counts as a verdict, score, rank, grade, priority, or action-guiding judgment. A comparison can remain descriptive; an evaluation cannot omit the criterion-bearing use.
Children (43) — more specific cases that build on this
-
Arete Domain-specific is a kind of Evaluation
The proposed strict upward parent is
prime:evaluation.prime:evaluation is the nearest broader Prime; the source domain and invariant supply the autonomous residual. This is a proposal-only workspace relationship: the accepted Prime supplies a genuinely instantiated structural prerequisite or superclass, while Arete adds domain-specific constraints. The entry does not collapse into that parent because the domain-specific identity determined by the Greek text and period, bearer of excellence, role or function, practical, civic, intellectual or moral dimension, evaluator and community, relation to telos and eudaimonia, translation, and contrast with modern virtue are explicit It also declines a nearby thematic catalog node: the neighbor does not literally subsume the constitutive identity of Arete. This explicit assert-and-decline pattern keeps the proposed DAG narrow and prevents a merely thematic edge. The prospective workspace queue contains one strict upward edge toprime:evaluation. No live DAG mutation is authorized. -
Capability Maturity Model Domain-specific is a kind of Evaluation
The proposed strict upward parent is
prime:evaluation.prime:evaluation is the nearest broader Prime while the source-domain carrier and invariant supply the autonomous residual. This is a proposal-only workspace relationship: the accepted Prime supplies a genuinely instantiated structural prerequisite or superclass, while Capability Maturity Model adds domain-specific constraints. The entry does not collapse into that parent because the domain-specific identity fixed by the organization and assessment scope, five maturity levels and their ordering, key process areas or goals, institutionalization and evidence criteria, appraisal method, current and target level, improvement actions, measurement feedback, distinction between process maturity and output performance and relationship to CMMI are explicit It also declines a nearby thematic catalog node: the neighbor does not literally subsume the constitutive identity of Capability Maturity Model. This explicit assert-and-decline pattern keeps the proposed DAG narrow and prevents a merely thematic edge. The prospective workspace queue contains one strict upward edge toprime:evaluation. No live DAG mutation is authorized. -
Considered purchase Domain-specific is a kind of Evaluation
The proposed strict upward parent is
prime:evaluation.prime:evaluation is the nearest broader Prime; the source domain and invariant supply the autonomous residual. This is a proposal-only workspace relationship: the accepted Prime supplies a genuinely instantiated structural prerequisite or superclass, while Considered purchase adds domain-specific constraints. The entry does not collapse into that parent because the domain-specific identity determined by the buyer and stakeholders, product or service, stakes and reversibility, perceived risks, information search, alternatives and criteria, decision roles, timeline and postpurchase consequences are explicit It also declines a nearby thematic catalog node: the neighbor does not literally subsume the constitutive identity of Considered purchase. This explicit assert-and-decline pattern keeps the proposed DAG narrow and prevents a merely thematic edge. The prospective workspace queue contains one strict upward edge toprime:evaluation. No live DAG mutation is authorized.
- Cumulative Prospect Theory Domain-specific is a kind of Evaluation
CPT instantiates **Evaluation** by applying a criterion-bearing frame to prospects and producing a ranking.It relates to **Risk**, **Loss Aversion**, and **Expected Utility**, but is not a specialization of any of them: Risk names exposure, Loss Aversion one component, and Expected Utility a competing aggregation rule. Probability Weighting Function is the closest domain component.
- Doctrine of inherency Domain-specific is a kind of Evaluation
The proposed strict upward parent is `prime:evaluation`.prime:evaluation is the nearest broader Prime while the source-domain carrier and invariant supply the autonomous residual. This is a proposal-only workspace relationship: the accepted Prime supplies a genuinely instantiated structural prerequisite or superclass, while Doctrine of inherency adds domain-specific constraints. The entry does not collapse into that parent because the domain-specific identity fixed by the governing jurisdiction and current doctrine, patent claim and each limitation, single prior-art reference and date, express teaching, asserted unstated feature, inevitability rather than possibility, evidentiary support and burden, enablement and recognition issues and anticipation versus obviousness role are explicit It also declines a nearby thematic catalog node: the neighbor does not literally subsume the constitutive identity of Doctrine of inherency. This explicit assert-and-decline pattern keeps the proposed DAG narrow and prevents a merely thematic edge. The prospective workspace queue contains one strict upward edge to `prime:evaluation`. No live DAG mutation is authorized.
- Forecast skill Domain-specific is a kind of Evaluation
The proposed strict upward parent is `prime:evaluation`.prime:evaluation is the nearest broader Prime; the source domain and invariant supply the autonomous residual. This is a proposal-only workspace relationship: the accepted Prime supplies a genuinely instantiated structural prerequisite or superclass, while Forecast skill adds domain-specific constraints. The entry does not collapse into that parent because the domain-specific identity determined by the forecast and predictand, cases and matching, lead time and spatial scale, deterministic or probabilistic form, verification data, score orientation and decomposition, reference forecast, skill-score formula, uncertainty and missing cases are explicit It also declines a nearby thematic catalog node: the neighbor does not literally subsume the constitutive identity of Forecast skill. This explicit assert-and-decline pattern keeps the proposed DAG narrow and prevents a merely thematic edge. The prospective workspace queue contains one strict upward edge to `prime:evaluation`. No live DAG mutation is authorized.
- Form, fit and function Domain-specific is a kind of Evaluation
The proposed strict upward parent is `prime:evaluation`.prime:evaluation is the nearest broader Prime while the source-domain carrier and invariant supply the autonomous residual. This is a proposal-only workspace relationship: the accepted Prime supplies a genuinely instantiated structural prerequisite or superclass, while Form, fit and function adds domain-specific constraints. The entry does not collapse into that parent because the domain-specific identity fixed by the reference and candidate item, controlled specification and configuration, form attributes and tolerances, fit interfaces and installation envelope, function performance and operating conditions, verification evidence, interchangeability decision, excluded material process reliability or certification attributes and change-control consequences are explicit It also declines a nearby thematic catalog node: the neighbor does not literally subsume the constitutive identity of Form, fit and function. This explicit assert-and-decline pattern keeps the proposed DAG narrow and prevents a merely thematic edge. The prospective workspace queue contains one strict upward edge to `prime:evaluation`. No live DAG mutation is authorized.
- Full employment Domain-specific is a kind of Evaluation
The proposed strict upward parent is `prime:evaluation`.prime:evaluation is the nearest broader Prime while the source-domain carrier and invariant supply the autonomous residual. This is a proposal-only workspace relationship: the accepted Prime supplies a genuinely instantiated structural prerequisite or superclass, while Full employment adds domain-specific constraints. The entry does not collapse into that parent because the domain-specific identity fixed by the economy population and period, working-age labor force and participation, employment and unemployment definitions, frictional structural and cyclical components, potential-output or inflation criterion, estimated full-employment unemployment rate, underemployment and distribution, measurement uncertainty and policy-theory convention are explicit It also declines a nearby thematic catalog node: the neighbor does not literally subsume the constitutive identity of Full employment. This explicit assert-and-decline pattern keeps the proposed DAG narrow and prevents a merely thematic edge. The prospective workspace queue contains one strict upward edge to `prime:evaluation`. No live DAG mutation is authorized.
- High-stakes testing Domain-specific is a kind of Evaluation
The proposed strict upward parent is `prime:evaluation`.prime:evaluation is the nearest broader Prime while the source-domain carrier and invariant supply the autonomous residual. This is a proposal-only workspace relationship: the accepted Prime supplies a genuinely instantiated structural prerequisite or superclass, while High-stakes testing adds domain-specific constraints. The entry does not collapse into that parent because the domain-specific identity fixed by the assessed construct and test instrument, test taker and other accountable units, score scale and reliability, decision threshold or classification rule, attached benefits sanctions or access consequences, intended use and validity evidence, fairness accommodations and disparate impact, preparation and incentive effects, appeals retesting and error costs and distinction between assessment information and consequential policy are explicit It also declines a nearby thematic catalog node: the neighbor does not literally subsume the constitutive identity of High-stakes testing. This explicit assert-and-decline pattern keeps the proposed DAG narrow and prevents a merely thematic edge. The prospective workspace queue contains one strict upward edge to `prime:evaluation`. No live DAG mutation is authorized.
- Hostile work environment Domain-specific is a kind of Evaluation
The proposed strict upward parent is `prime:evaluation`.prime:evaluation is the nearest broader Prime while the source-domain carrier and invariant supply the autonomous residual. This is a proposal-only workspace relationship: the accepted Prime supplies a genuinely instantiated structural prerequisite or superclass, while Hostile work environment adds domain-specific constraints. The entry does not collapse into that parent because the domain-specific identity fixed by the jurisdiction and governing law, employment relationship, protected characteristic, unwelcome discriminatory conduct, subjective perception, objective severity or pervasiveness, frequency and context, workplace effect, actor and employer-liability rules and available defenses are explicit It also declines a nearby thematic catalog node: the neighbor does not literally subsume the constitutive identity of Hostile work environment. This explicit assert-and-decline pattern keeps the proposed DAG narrow and prevents a merely thematic edge. The prospective workspace queue contains one strict upward edge to `prime:evaluation`. No live DAG mutation is authorized.
- Industrial design rights in the European Union Domain-specific is a kind of Evaluation
The proposed strict upward parent is `prime:evaluation`.prime:evaluation is the nearest broader Prime while the source-domain carrier and invariant supply the autonomous residual. This is a proposal-only workspace relationship: the accepted Prime supplies a genuinely instantiated structural prerequisite or superclass, while Industrial design rights in the European Union adds domain-specific constraints. The entry does not collapse into that parent because the domain-specific identity fixed by the governing EU and national instruments, claimed product and visual features, disclosure and priority dates, novelty, informed-user overall impression and designer freedom, technical-function and public-policy exclusions, registered or unregistered status, term, territorial scope and infringement and validity tests are explicit It also declines a nearby thematic catalog node: the neighbor does not literally subsume the constitutive identity of Industrial design rights in the European Union. This explicit assert-and-decline pattern keeps the proposed DAG narrow and prevents a merely thematic edge. The prospective workspace queue contains one strict upward edge to `prime:evaluation`. No live DAG mutation is authorized.
- Jury Domain-specific is a kind of Evaluation
The proposed strict upward parent is `prime:evaluation`.prime:evaluation is the nearest broader Prime while the source-domain carrier and invariant supply the autonomous residual. This is a proposal-only workspace relationship: the accepted Prime supplies a genuinely instantiated structural prerequisite or superclass, while Jury adds domain-specific constraints. The entry does not collapse into that parent because the domain-specific identity fixed by the jurisdiction and court, jury type and legal authority, venire selection and eligibility, voir dire and challenges, sworn members and alternates, assigned questions of fact indictment penalty or verdict, admitted evidence and instructions, deliberation and voting rule, verdict and judicial review and protections for impartiality and secrecy are explicit It also declines a nearby thematic catalog node: the neighbor does not literally subsume the constitutive identity of Jury. This explicit assert-and-decline pattern keeps the proposed DAG narrow and prevents a merely thematic edge. The prospective workspace queue contains one strict upward edge to `prime:evaluation`. No live DAG mutation is authorized.
- Machine-Learning Learning Curve Domain-specific is a kind of Evaluation
**`prime:evaluation` — proposed strict subsumption parent.** The curve repeatedly evaluates model states or fits under a shared metric and produces an action-guiding diagnosis; the child adds ML exposure and split-specific obligations.**`prime:evaluation` — proposed strict subsumption parent.** The curve repeatedly evaluates model states or fits under a shared metric and produces an action-guiding diagnosis; the child adds ML exposure and split-specific obligations.
- MAGIC criteria Domain-specific is a kind of Evaluation
The proposed strict upward parent is `prime:evaluation`.prime:evaluation is the nearest broader Prime while the source-domain carrier and invariant supply the autonomous residual. This is a proposal-only workspace relationship: the accepted Prime supplies a genuinely instantiated structural prerequisite or superclass, while MAGIC criteria adds domain-specific constraints. The entry does not collapse into that parent because the domain-specific identity fixed by the claim and intended audience, measured magnitude, articulated pattern, population and context of generality, prior beliefs and interestingness, design and inferential credibility, tradeoffs among dimensions and supporting evidence are explicit It also declines a nearby thematic catalog node: the neighbor does not literally subsume the constitutive identity of MAGIC criteria. This explicit assert-and-decline pattern keeps the proposed DAG narrow and prevents a merely thematic edge. The prospective workspace queue contains one strict upward edge to `prime:evaluation`. No live DAG mutation is authorized.
- Mean integrated squared error Domain-specific is a kind of Evaluation
The proposed strict upward parent is `prime:evaluation`.prime:evaluation is the nearest broader Prime while the source-domain carrier and invariant supply the autonomous residual. This is a proposal-only workspace relationship: the accepted Prime supplies a genuinely instantiated structural prerequisite or superclass, while Mean integrated squared error adds domain-specific constraints. The entry does not collapse into that parent because the domain-specific identity fixed by the unknown function or density f, random sample and estimator f_n, integration domain and measure, pointwise error, squared L2 norm, expectation over samples, integrated variance and integrated squared bias decomposition, finite-sample MISE and AMISE, bandwidth or complexity choice and integrability assumptions are explicit It also declines a nearby thematic catalog node: the neighbor does not literally subsume the constitutive identity of Mean integrated squared error. This explicit assert-and-decline pattern keeps the proposed DAG narrow and prevents a merely thematic edge. The prospective workspace queue contains one strict upward edge to `prime:evaluation`. No live DAG mutation is authorized.
- Nimber Domain-specific is a kind of Evaluation
Nimber instantiates Evaluation because it maps each eligible game position to a contextually meaningful value that predicts outcome class and composes across independent subgames.The prospective workspace queue contains one strict upward edge to `prime:evaluation`. No live DAG mutation is authorized.
- Optimality criterion Domain-specific is a kind of Evaluation
The accepted reference-grade review places Optimality criterion under Evaluation because the child instantiates or depends on the parent's broader structure while retaining its own constitutive identity.An objective measure used to compare candidate statistical models for a hypothesis and designate the model with the best criterion value. The parent is defined more broadly: Apply a criterion-bearing frame to a bounded object, interpret its relevant features against that frame, and produce a verdict, score, rank, or action-guiding judgment.
- Outliers ratio Domain-specific is a kind of Evaluation
The proposed strict upward parent is `prime:evaluation`.prime:evaluation is the nearest broader Prime; the source domain and invariant supply the autonomous residual. This is a proposal-only workspace relationship: the accepted Prime supplies a genuinely instantiated structural prerequisite or superclass, while Outliers ratio adds domain-specific constraints. The entry does not collapse into that parent because the domain-specific identity determined by the video dataset and subjective experiment, objective metric and mapping, mean opinion score and dispersion, tolerance multiplier and interval, paired cases, outlier rule, denominator and uncertainty are explicit It also declines a nearby thematic catalog node: the neighbor does not literally subsume the constitutive identity of Outliers ratio. This explicit assert-and-decline pattern keeps the proposed DAG narrow and prevents a merely thematic edge. The prospective workspace queue contains one strict upward edge to `prime:evaluation`. No live DAG mutation is authorized.
- Performance Appraisal Domain-specific is a kind of Evaluation
Performance Appraisal strictly instantiates **Evaluation**.The employee’s work over the period is the bounded object; role expectations are the criterion frame; observed outputs and behavior are the relevant evidence; the appraisal method maps evidence to a rating or narrative judgment; and the result guides development or personnel action. The proposed DAG edge records this specialization. It is related to **Review Artifact**, because many systems persist an attributable judgment and warrant; **Feedback**, because the result is communicated and may influence later performance; **Summative Assessment**, when a period-ending judgment certifies standing; and **Goal Congruence**, when criteria align individual and organizational objectives. None is required as an additional parent: appraisal can be developmental, rating-free, or based on role standards other than cascading goals.
- Policy and charging rules function Domain-specific is a kind of Evaluation
The proposed strict upward parent is `prime:evaluation`.prime:evaluation is the nearest broader Prime while the source-domain carrier and invariant supply the autonomous residual. This is a proposal-only workspace relationship: the accepted Prime supplies a genuinely instantiated structural prerequisite or superclass, while Policy and charging rules function adds domain-specific constraints. The entry does not collapse into that parent because the domain-specific identity fixed by the standards release and network architecture, subscriber and service context, policy data and rule priorities, service-data-flow identification, quality-of-service and gating decision, online or offline charging rule, interfaces to enforcement and charging functions, session lifecycle and conflict and failure behavior are explicit It also declines a nearby thematic catalog node: the neighbor does not literally subsume the constitutive identity of Policy and charging rules function. This explicit assert-and-decline pattern keeps the proposed DAG narrow and prevents a merely thematic edge. The prospective workspace queue contains one strict upward edge to `prime:evaluation`. No live DAG mutation is authorized.
- Scoring Rule Domain-specific is a kind of Evaluation
**`evaluation`.** A scoring rule applies a criterion-bearing map to a bounded forecast and realized evidence, producing an interpretable score.This is the proposed strict subsumption parent.
- Social experiment Domain-specific is a kind of Evaluation
The proposed strict upward parent is `prime:evaluation`.prime:evaluation is the nearest broader Prime while the source-domain carrier and invariant supply the autonomous residual. This is a proposal-only workspace relationship: the accepted Prime supplies a genuinely instantiated structural prerequisite or superclass, while Social experiment adds domain-specific constraints. The entry does not collapse into that parent because the domain-specific identity fixed by the social question and hypothesis, participants or units, setting and time horizon, manipulated or naturally varied condition, comparison or counterfactual, outcomes and participant perspectives, assignment and confounding, observation and analysis plan, consent oversight and risk, external validity and interpretation limits are explicit It also declines a nearby thematic catalog node: the neighbor does not literally subsume the constitutive identity of Social experiment. This explicit assert-and-decline pattern keeps the proposed DAG narrow and prevents a merely thematic edge. The prospective workspace queue contains one strict upward edge to `prime:evaluation`. No live DAG mutation is authorized.
- Subjective validation Domain-specific is a kind of Evaluation
The proposed strict upward parent is `prime:evaluation`.prime:evaluation is the nearest broader Prime while the source-domain carrier and invariant supply the autonomous residual. This is a proposal-only workspace relationship: the accepted Prime supplies a genuinely instantiated structural prerequisite or superclass, while Subjective validation adds domain-specific constraints. The entry does not collapse into that parent because the domain-specific identity fixed by the person and prior beliefs, statement prediction interpretation or pair of events, ambiguity or base-rate weakness, personal relevance and emotional salience, selective matching and memory, perceived accuracy or causal relation, ignored mismatches and alternative explanations, reinforcement loop and experimental comparison or calibration are explicit It also declines a nearby thematic catalog node: the neighbor does not literally subsume the constitutive identity of Subjective validation. This explicit assert-and-decline pattern keeps the proposed DAG narrow and prevents a merely thematic edge. The prospective workspace queue contains one strict upward edge to `prime:evaluation`. No live DAG mutation is authorized.
- Test Drive Domain-specific is a kind of Evaluation
A Test Drive strictly specializes **Evaluation**.It applies a criterion-bearing frame to a bounded vehicle, interprets dynamic and embodied observations, and produces an action-guiding judgment. Evaluation is therefore the single proposed DAG parent. It also relates to **Validation**, because operation in a realistic context can determine whether the vehicle meets intended use; **Sampling (Representativeness)**, because a short route samples a larger ownership regime; **Joint vs. Separate Evaluation**, because side-by-side comparison changes which attributes become salient; and **Uncertainty-Driven Verification Premium**, because buyers may invest time and inspection expense when static evidence is insufficient. These are explanatory relations in prose, not additional direct parents.
- Test (law) Domain-specific is a kind of Evaluation
The proposed strict upward parent is `prime:evaluation`.prime:evaluation is the nearest broader Prime; the source domain and invariant supply the autonomous residual. This is a proposal-only workspace relationship: the accepted Prime supplies a genuinely instantiated structural prerequisite or superclass, while Test (law) adds domain-specific constraints. The entry does not collapse into that parent because the domain-specific identity determined by the jurisdiction and date, legal issue, authoritative source, elements or factors, logical relation, burden and standard of proof, evidence mapping, exceptions, remedy and precedential status are explicit It also declines a nearby thematic catalog node: the neighbor does not literally subsume the constitutive identity of Test (law). This explicit assert-and-decline pattern keeps the proposed DAG narrow and prevents a merely thematic edge. The prospective workspace queue contains one strict upward edge to `prime:evaluation`. No live DAG mutation is authorized.
- Transfer pricing Domain-specific is a kind of Evaluation
The proposed strict upward parent is `prime:evaluation`.prime:evaluation is the nearest broader Prime while the source-domain carrier and invariant supply the autonomous residual. This is a proposal-only workspace relationship: the accepted Prime supplies a genuinely instantiated structural prerequisite or superclass, while Transfer pricing adds domain-specific constraints. The entry does not collapse into that parent because the domain-specific identity fixed by the related entities and ownership or control relation, jurisdictions and tax periods, accurately delineated controlled transaction, functions assets and risks, contractual and actual conduct, transfer-pricing method and tested party, comparables and adjustments, arm’s-length price or range, documentation, corresponding and secondary adjustments and uncertainty or dispute mechanism are explicit It also declines a nearby thematic catalog node: the neighbor does not literally subsume the constitutive identity of Transfer pricing. This explicit assert-and-decline pattern keeps the proposed DAG narrow and prevents a merely thematic edge. The prospective workspace queue contains one strict upward edge to `prime:evaluation`. No live DAG mutation is authorized.
- Typical versus Maximum Performance Domain-specific is a kind of Evaluation
criterion and observation frame determine what performance verdict means.criterion and observation frame determine what performance verdict means.
- Ultimate Fact Domain-specific is a kind of Evaluation
**Evaluation** is the proposed immediate parent.Evidence, Deductive Reasoning, Verification, Burden of Proof, and Fact–Value Distinction are related primes. The prospective queue contains one strict edge to `prime:evaluation`. No live DAG mutation is authorized.
- User analysis Domain-specific is a kind of Evaluation
The proposed strict upward parent is `prime:evaluation`.prime:evaluation is the nearest broader Prime while the source-domain carrier and invariant supply the autonomous residual. This is a proposal-only workspace relationship: the accepted Prime supplies a genuinely instantiated structural prerequisite or superclass, while User analysis adds domain-specific constraints. The entry does not collapse into that parent because the domain-specific identity fixed by the product and decision scope, intended and excluded user populations, recruitment and sampling, goals and workflows, knowledge skills accessibility and constraints, physical social and technical context, evidence methods, segmentation and uncertainty, derived requirements or personas and traceability to design choices are explicit It also declines a nearby thematic catalog node: the neighbor does not literally subsume the constitutive identity of User analysis. This explicit assert-and-decline pattern keeps the proposed DAG narrow and prevents a merely thematic edge. The prospective workspace queue contains one strict upward edge to `prime:evaluation`. No live DAG mutation is authorized.
- Warranting theory Domain-specific is a kind of Evaluation
The proposed strict upward parent is `prime:evaluation`.prime:evaluation is the nearest broader Prime while the source-domain carrier and invariant supply the autonomous residual. This is a proposal-only workspace relationship: the accepted Prime supplies a genuinely instantiated structural prerequisite or superclass, while Warranting theory adds domain-specific constraints. The entry does not collapse into that parent because the domain-specific identity fixed by the target and observer, identity or trait claim, information cue and its source, target control or manipulability, warranting value, perceived credibility, cross-channel or offline corroboration, competing deception and impression-management explanations, context and observer differences and resulting judgment are explicit It also declines a nearby thematic catalog node: the neighbor does not literally subsume the constitutive identity of Warranting theory. This explicit assert-and-decline pattern keeps the proposed DAG narrow and prevents a merely thematic edge. The prospective workspace queue contains one strict upward edge to `prime:evaluation`. No live DAG mutation is authorized.
- Widely applicable information criterion Domain-specific is a kind of Evaluation
The proposed strict upward parent is `prime:evaluation`.prime:evaluation is the nearest broader Prime while the source-domain carrier and invariant supply the autonomous residual. This is a proposal-only workspace relationship: the accepted Prime supplies a genuinely instantiated structural prerequisite or superclass, while Widely applicable information criterion adds domain-specific constraints. The entry does not collapse into that parent because the domain-specific identity fixed by the observed data and pointwise likelihood units, posterior distribution and draws, log pointwise predictive density, variance penalty and effective parameter count, factor of minus two convention, standard error, model comparison and asymptotic relation and diagnostics for unstable contributions are explicit It also declines a nearby thematic catalog node: the neighbor does not literally subsume the constitutive identity of Widely applicable information criterion. This explicit assert-and-decline pattern keeps the proposed DAG narrow and prevents a merely thematic edge. The prospective workspace queue contains one strict upward edge to `prime:evaluation`. No live DAG mutation is authorized.
- Yau's conjecture on the first eigenvalue Domain-specific is a kind of Evaluation
The proposed strict upward parent is `prime:evaluation`.prime:evaluation is the nearest broader Prime while the source-domain carrier and invariant supply the autonomous residual. This is a proposal-only workspace relationship: the accepted Prime supplies a genuinely instantiated structural prerequisite or superclass, while Yau's conjecture on the first eigenvalue adds domain-specific constraints. The entry does not collapse into that parent because the domain-specific identity fixed by the unit ambient sphere and dimension, closed embedded minimal hypersurface, induced metric, Laplace–Beltrami operator and eigenvalue convention, first nonzero eigenvalue, asserted equality to n, coordinate-function eigenmodes, proved special cases and open general status are explicit It also declines a nearby thematic catalog node: the neighbor does not literally subsume the constitutive identity of Yau's conjecture on the first eigenvalue. This explicit assert-and-decline pattern keeps the proposed DAG narrow and prevents a merely thematic edge. The prospective workspace queue contains one strict upward edge to `prime:evaluation`. No live DAG mutation is authorized.
- Cognitive Appraisal Prime is a kind of Evaluation
Cognitive Appraisal is the organism-centered species of Evaluation whose object is a situation and whose criteria are goal significance and coping capacity.Cognitive Appraisal preserves Evaluation's bounded target, criterion-bearing frame, interpretation of relevant features, and action-guiding judgment. Its differentia are psychological: the target is a perceived situation, the frame is the organism's goals and coping resources, and the result organizes emotion and behavior. Evaluation need not involve an organism, emotion, threat, or coping, so the relation is strict subsumption rather than identity.
- Peer Review Prime is a kind of Evaluation
The accepted reference-grade review places Peer Review under Evaluation because the child instantiates or depends on the parent's broader structure while retaining its own constitutive identity.Evaluate bounded work or performance through people with relevant peer competence, routing criterion-grounded judgments into revision, authorization, learning, or accountability. The parent is defined more broadly: Apply a criterion-bearing frame to a bounded object, interpret its relevant features against that frame, and produce a verdict, score, rank, or action-guiding judgment.
- Truth value Prime is a kind of Evaluation
The accepted reference-grade review places Truth value under Evaluation because the child instantiates or depends on the parent's broader structure while retaining its own constitutive identity.A semantic value assigned to a proposition or formula to represent its status with respect to truth under a specified logic and interpretation. The parent is defined more broadly: Apply a criterion-bearing frame to a bounded object, interpret its relevant features against that frame, and produce a verdict, score, rank, or action-guiding judgment.
- Utility Prime is a kind of Evaluation
The accepted reference-grade review places Utility under Evaluation because the child instantiates or depends on the parent's broader structure while retaining its own constitutive identity.Represent the preference value, benefit or desirability of an outcome for a specified agent or evaluative system. The parent is defined more broadly: Apply a criterion-bearing frame to a bounded object, interpret its relevant features against that frame, and produce a verdict, score, rank, or action-guiding judgment.
- Verification Prime is a kind of Evaluation
Verification is the strict species of Evaluation whose criterion is a stated specification and whose defined checking procedure yields evidence and a conformance verdict.Verification preserves Evaluation's bounded object, criterion frame, relevant observations, and evaluative result, then supplies the exact differentia: the criterion is a specification taken as given, the mapping is a defined checking procedure, and the result is a conformance verdict supported by evidence. Evaluations may instead use open, plural, purposive, comparative, or aesthetic criteria and need not test specification conformance.
- Expectancy Disconfirmation Domain-specific is part of Evaluation
Expectancy Disconfirmation contains Evaluation because the signed expectation-performance gap becomes a criterion-bearing satisfaction judgment rather than remaining a descriptive residual.The model requires more than computing a discrepancy. A person evaluates the experienced performance relative to an expectation and converts the comparison into satisfaction, dissatisfaction, or confirmation. Evaluation supplies the criterion-bearing verdict; the child fixes its reference to a prior expectancy and makes discrepancy sign the central explanatory variable.
- Natural Capital Domain-specific presupposes Evaluation
Natural capital requires a criterion-bearing evaluation that treats ecological condition and future service capacity as an asset rather than unpriced background.Evaluation supplies the frame that makes a natural system legible as capital: a bounded stock is assessed against criteria for condition, productive capacity, persistence, and benefit generation. Remove that criterion-bearing judgment and the forest, wetland, soil, or fish stock remains a biophysical system but is no longer being treated as a value-bearing asset. Natural Capital adds the ecological stock-versus-service-flow and depletion accounting commitments.
- Peak-end rule Domain-specific presupposes Evaluation
Peak-End Rule presupposes Evaluation because it specifies how an extended experience is converted into a retrospective verdict that guides repeat-or-avoid choice.The rule does not govern storage or recall in general. It applies when an extended affective episode is judged after the fact against a good-bad, pleasant-painful, or repeat-avoid criterion. Evaluation supplies that criterion-bearing operation; Peak-End Rule supplies the distinctive weighting function used to form its verdict.
- Review Domain-specific is part of Evaluation
A Review strictly contains an Evaluation and adds attribution, warrant, persistence, addressability, and documentary affordances around it.Evaluation is the inner criterion-bearing operation that maps a bounded object and relevant observations to a verdict, score, rank, or judgment. Review crystallizes that evaluation into a persisted four-slot artifact by binding it to an identified evaluator and warrant. Removing Evaluation leaves a record with no evaluative act or verdict to document; removing the review packaging can leave a complete but transient or private Evaluation.
- Joint vs. Separate Evaluation Prime presupposes Evaluation
Joint-vs-Separate Evaluation presupposes Evaluation because it is a mode effect defined over procedures that map criteria and attributes into verdicts.The phenomenon cannot occur without an Evaluation: both joint and separate modes evaluate the same bounded options, but each mode changes the criterion frame by making a different subset of attributes evaluable and therefore can change the verdict. Evaluation supplies the common operation on which the mode toggle acts. Comparison remains foundational through Evaluation; the old direct edge was a flattened shortcut and was only typical because comparison is engaged explicitly in joint mode but suppressed in separate mode.
- News Values Domain-specific is a decomposition of Evaluation
Removing journalism-specific factors from News Values leaves Evaluation's object–criteria–observations–result mapping: candidate events are interpreted against a weighted frame to produce scores and coverage verdicts.Strip Galtung–Ruge and Harcup–O'Neill factors, newsroom institutions, news cycles, editorial culture, and coverage language. The preserved operation bounds a candidate, selects relevant features under a criterion-bearing frame, aggregates or interprets them, and produces an action-guiding score or verdict. News Values adds the journalism-specific criteria vector and its institutional meaning.
Hierarchy path (1) — routes to 1 parentless root
- Evaluation → Comparison → Self Checking
Neighborhood in Abstraction Space¶
Evaluation sits in a sparse region of abstraction space (80th percentile for distinctiveness): few abstractions share its structure, so a faithful description tends to retrieve it precisely rather than landing on a neighbor.
Family — Verification, Screening & Reference Standards (18 primes)
Nearest neighbors
- Boundedness — 0.72
- Evaluative Rating — 0.71
- Peer Review — 0.71
- Recognition Justice — 0.69
- Resolution Matching — 0.68
Computed from structural-signature embeddings · 2026-09-10
Not to Be Confused With¶
Comparison is Evaluation's strict prerequisite: the object's relevant features and the criterion must be placed in a shared frame so a relation can be read. Comparison alone need not render a verdict. Measurement supplies scaled observations but does not say what those values count for. Evidence relates traces to hypotheses and can support an evaluation, but aesthetic, interpretive, and rule-based evaluations need not use Evidence in that prime's narrower trace-to-unobservable-state sense.
Verification is a strict species: it fixes a specification and uses a defined checking procedure to produce evidence and a conformance verdict. Validation asks whether the verified artifact serves its intended real-world purpose. Cognitive Appraisal is the organism-centered species that evaluates situational significance and coping resources. Joint-vs-Separate Evaluation is a mode effect within evaluation procedures. Review is the domain-specific documentary artifact that persists an attributable evaluation together with a verdict and warrant.
Decision is downstream: it collapses alternatives into commitment. An evaluation may rank or recommend without authorizing action, and a decision may rely on several evaluations plus constraints that no single evaluation contains.
Solution Archetypes¶
No catalogued solution archetypes reference this prime yet.
References¶
- Scriven, M. (1991). Evaluation Thesaurus (4th ed.). Sage.
- House, E. R. (1993). Professional Evaluation: Social Impact and Political Consequences. Sage.