Skip to content

Educational Measurement

The evidence-and-inference practice that maps learner performances elicited by designed assessments onto scores or classifications whose reliability, validity, comparability, and fairness are evaluated for declared educational uses.

Version
v2 · 2026-09-06 · History
Domain-specific #
1740
Origin domain
education
Subdomain
educational assessment and psychometrics

Core Idea

Educational Measurement is the disciplined construction and evaluation of an inferential chain from what learners say, do, select, or produce under assessment conditions to scores, classifications, and claims about educationally relevant knowledge, skills, dispositions, or attainment. It does not discover proficiency by directly reading an internal quantity. It designs situations that elicit evidence, scores the resulting performances, models how those observations relate to a target construct, locates persons or performances on a scale, quantifies uncertainty, and evaluates whether the proposed interpretation and use are reliable, valid, comparable, and fair.

The National Research Council’s assessment triangle gives the compact logic: cognition + observation + interpretation. A model of learning identifies what competence means in a subject; tasks or situations elicit performances that could reveal that competence; an interpretation process turns the fallible observations into claims.[1] All three must be coordinated. A mathematically sophisticated scoring model cannot rescue tasks that elicit the wrong knowledge, and representative tasks do not yield warranted claims if scoring or interpretation is incoherent.

The current NCME handbook presents Educational Measurement as a mature measurement science whose recurring components include validity and validation, reliability, fairness, assessment design, administration and scoring, statistical modeling, scaling and equating, standard setting, reporting, instructional assessment, accountability, admissions, certification, and international comparison.[2] The joint AERA–APA–NCME Standards for Educational and Psychological Testing supplies professional obligations for test development, evaluation, administration, interpretation, and use, with validity, reliability/precision, and fairness as foundational concerns.[3]

The locked identity is therefore declared educational purpose and score claim + target learner construct + population and context + designed observation opportunities + response capture and scoring + measurement/interpretation model + score scale or classification + reliability/precision evidence + validation argument + comparability and fairness conditions + reporting and use limits -> Educational Measurement. A quiz, score, rubric, statistical model, or decision is only one component. The abstraction is the governed inferential system that makes the resulting educational claim more or less warranted.

Structural Signature

  • the educational purpose — diagnosis, feedback, progress monitoring, certification, selection, placement, accountability, comparison, or research, stated before choosing evidence;
  • the intended interpretation and use — the exact claim a score or classification is supposed to support and the decision, if any, it will inform;
  • the target construct or domain — the knowledge, skill, proficiency, disposition, or attainment to be characterized, including its content boundaries and learning progression;
  • the target population and context — learners, language backgrounds, instructional opportunities, administration conditions, institutions, and time period for which the interpretation is proposed;
  • the observation design — items, tasks, performances, portfolios, conversations, simulations, or other situations selected to elicit construct-relevant evidence;
  • the response process — how learners understand and engage with tasks, including strategies and accessibility barriers that may add construct-irrelevant demands;
  • the scoring rule — keys, rubrics, rater processes, automated models, partial-credit rules, or coding procedures that convert performances into observations or scores;
  • the measurement or interpretation model — classical test theory, item response theory, Rasch measurement, generalizability theory, diagnostic classification, qualitative teacher interpretation, or another declared bridge from evidence to claims;
  • the scale or reporting category — raw scores, scale scores, proficiency levels, subscores, percentiles, standards-based classifications, or narrative profiles, each with limited permissible interpretations;
  • the uncertainty and reliability account — sampling of tasks, occasions, raters, forms, and response variability, expressed through reliability, conditional standard errors, classification consistency, or another precision analysis;
  • the validity argument — evidence and assumptions supporting the proposed interpretation and use, not a permanent property stamped onto the test;
  • the comparability apparatus — standardization, linking, equating, common items, common populations, calibration, or moderation used when scores from different forms, administrations, or groups are compared;
  • the fairness and accessibility account — checks that construct-irrelevant barriers, bias, differential functioning, administration differences, and accommodations do not invalidate claims for relevant groups;
  • the decision rule — cut scores, admission thresholds, growth criteria, mastery rules, or instructional actions kept distinct from the underlying measurement;
  • the reporting interface — documentation that communicates scale meaning, uncertainty, appropriate comparisons, limitations, and intended uses to learners, educators, families, and institutions;
  • the monitoring loop — field testing, item analysis, rater monitoring, drift checks, consequence review, and periodic revalidation as populations, curricula, technology, and uses change.

Recognition test. Name the educational construct, the elicited performances, the scoring process, the interpretive model, the intended score claim and use, and the evidence for precision, validity, and fairness. If these roles cannot be mapped, the case may be instruction, grading, testing, data collection, or evaluation, but it is not yet Educational Measurement in the reference-grade sense.

What It Is Not

  • Not educational assessment as a bare event. Assessment gathers evidence through a quiz, project, observation, conversation, or examination. Educational Measurement is the inferential and technical system that defines how that evidence supports scores and claims.
  • Not a test. A test is one instrument or procedure. Measurement includes construct definition, assembly, scoring, modeling, validation, reliability, fairness, comparability, reporting, and use.
  • Not a score. A raw total or scale value is an output. Its meaning depends on the construct, form, model, reference population, precision, and proposed interpretation.
  • Not grading alone. A course grade can combine achievement, effort, participation, lateness, improvement, and teacher judgment. It is not automatically a coherent measurement of one construct or comparable across classrooms.
  • Not psychometrics as a whole. Psychometrics develops theories and methods for measuring psychological attributes across clinical, occupational, social, and educational settings. Educational Measurement selects and adapts those methods for educational constructs, populations, decisions, and institutions.
  • Not classical test theory, item response theory, or the Rasch model. Each is a family of interpretation models. The abstraction persists when a different justified model is used.
  • Not reliability. Reliability or precision concerns the consistency and error structure of scores. Consistently measuring the wrong construct can be highly reliable and invalid.
  • Not validity as a coefficient. Kane emphasizes that interpretations and uses, not tests or scores in isolation, are validated through an argument whose assumptions require evidence.[4]
  • Not test equating. Equating is one comparability operation for interchangeable forms under strong requirements. Linking, concordance, moderation, scaling, and equating are not synonyms.
  • Not standard setting. Standard setting maps a score scale to policy or performance categories through judgment and evidence. It uses measurement results but adds a decision boundary.
  • Not only standardized high-stakes testing. Classroom evidence, performance tasks, portfolios, adaptive assessments, and formative systems can participate when their inference and quality obligations are explicit.
  • Not formative or summative assessment specifically. Those terms classify how evidence is used in a learning cycle. Educational Measurement supplies evidence and warrant for both purposes.
  • Not program evaluation. Evaluation judges the merit or effectiveness of a curriculum, school, or policy. It may use educational measures, but its target and inferential question differ.
  • Not numerals assigned by convention without warrant. Coding responses numerically creates data; it does not establish that arithmetic, ordering, comparison, or latent-trait interpretation is meaningful.

Scope of Application

Educational Measurement covers classroom, institutional, regional, national, and international settings. In classroom diagnosis, a teacher may use carefully designed tasks and qualitative interpretation to identify a misconception and choose the next instructional move. The NRC explicitly allows the interpretation corner to be qualitative and teacher-mediated rather than requiring a formal statistical model.[1] In large-scale assessment, response patterns are commonly modeled statistically, scores are placed on scales, form-to-form comparability is maintained, and uncertainty is reported.

The same identity supports reading and mathematics achievement tests, language-proficiency assessments, science performance tasks, social-emotional questionnaires used for educational purposes, admissions and placement tests, certification examinations, learning-progress measures, accountability systems, course examinations, and evaluations of instructional interventions. The details differ, but every legitimate case connects a target construct, observation opportunity, scoring/interpretation model, and declared use.

The practice extends across the assessment lifecycle: construct and domain analysis; blueprinting; item and task development; accessibility review; cognitive laboratories; pilot and field testing; scoring design; rater training; model fitting; scale construction; linking or equating; standard setting; score reporting; validity synthesis; operational monitoring; and retirement or redesign. Mislevy, Almond, and Lukas’s evidence-centered design formalizes part of this lifecycle by coordinating claims, evidence, tasks, and statistical models.[5]

Its scope ends when the educational inferential obligation disappears. A generic opinion survey of students may be social research rather than Educational Measurement if it makes no claim about an educational construct. A school dashboard that simply counts attendance is administrative measurement unless attendance is used within a declared educational inference. A teacher’s unstructured impression may inform instruction, but it becomes educational measurement only to the degree that the target, observations, interpretive basis, and limitations are explicit enough to warrant the claim.

Clarity

Educational Measurement becomes clear when the claim is written before the number. Consider “Jordan scored 650.” The value is uninterpretable until one knows: on which scale; from which form and administration; measuring what construct; relative to which population or criterion; with what standard error; supporting which use; and under what validity and fairness evidence. A percentile is not a percent correct, a scale score is not necessarily an interval with a natural zero, and a proficiency category does not erase measurement uncertainty near its cut score.

Classical test theory offers a simple decomposition,

\[ X=T+E, \]

where \(X\) is an observed score, \(T\) is the model-defined true score, and \(E\) is error. Under the model, reliability is the ratio \(\rho_{XX'}=\operatorname{Var}(T)/\operatorname{Var}(X)\), and a commonly used standard error of measurement is

\[ \operatorname{SEM}=s_X\sqrt{1-\rho_{XX'}}. \]

These quantities describe precision under specified replications; they do not validate the construct claim. A reliability estimate can change with population heterogeneity and the error facets included.

An item-response model makes a different bridge. In the one-parameter logistic/Rasch form,

\[ P(X_{pi}=1\mid\theta_p,b_i)=\frac{e^{\theta_p-b_i}}{1+e^{\theta_p-b_i}}, \]

where \(\theta_p\) is a person location and \(b_i\) an item location on a common latent scale. The formula is not a measurement license by itself. Unidimensionality, local independence, model fit, item functioning, population coverage, and substantive interpretation must be evaluated. The clean diagnostic is: what assumptions turn this response pattern into this score claim?

Manages Complexity

Learner competence is multidimensional, context-sensitive, and incompletely observable. Educational Measurement compresses a sample of performances into a manageable score, profile, or classification while preserving an explicit record of what was sampled and what uncertainty remains. Without this compression, institutions could not compare performance across hundreds of tasks, thousands of learners, multiple forms, years, or jurisdictions. With careless compression, however, the score hides precisely the variation a decision needs.

Blueprints control content sampling. Rubrics stabilize interpretation of constructed responses. Measurement models pool information across items and learners. Scales provide a common reporting surface. Equating and linking manage form changes. Reliability and conditional standard errors indicate where information is strong or weak. Validity arguments organize heterogeneous evidence—from content alignment, response processes, internal structure, relations with other variables, and consequences—around the proposed interpretation. Fairness analyses ask whether irrelevant barriers alter the inference for subgroups.

This machinery also separates layers that organizations routinely collapse: performance observed on particular tasks; score produced under a scoring rule; scale location estimated under a model; proficiency category imposed through a cut score; decision made under policy; consequence experienced by a learner. Keeping these layers distinct makes disputes diagnosable. A bad classification may arise from noisy observations, a biased task, a poor model, an unstable cut score, or an inappropriate use; each demands a different repair.

Abstract Reasoning

  1. A score has no interpretation independent of its construct, scale, population, administration, model, and intended use.
  2. Reliability is necessary for many ambitious inferences but never sufficient for validity; precise scores can support the wrong claim.
  3. Validity belongs to proposed interpretations and uses. The same assessment can support one use and fail another.[4]
  4. More ambitious claims require more evidence. Ranking performance on sampled tasks is narrower than claiming durable, transferable mastery.
  5. Sampling more items can reduce random error but cannot repair construct underrepresentation if all items omit a critical domain.
  6. A higher total score need not imply a higher latent estimate under every multidimensional or weighted model; the score rule and model control the ordering.
  7. Item difficulty is not an intrinsic constant under every model. Population, curriculum, administration mode, language, and item exposure can alter functioning.
  8. Equating requires stronger conditions than showing correlation. Highly correlated tests can measure different constructs or differ systematically in difficulty.
  9. A cut score converts continuous uncertainty into a discrete outcome. Classification consistency near the boundary deserves separate analysis.
  10. A group mean difference does not by itself establish item bias; bias concerns construct-irrelevant differential functioning and interpretation within the intended framework.
  11. An accommodation is defensible when it removes an irrelevant barrier without changing the target construct; that judgment is construct-specific.
  12. Automated scoring can improve consistency while introducing construct mismatch, subgroup error, opacity, or drift. Agreement with raters is evidence, not complete validation.
  13. Teaching to curricular objectives can improve measured proficiency; coaching narrowly to item formats can inflate scores without the intended learning. Score change requires a validity argument.
  14. A test can remain physically unchanged while its interpretation becomes invalid because curriculum, population, technology, language, exposure, or use has changed.
  15. Overall reliability can conceal low precision in the score region where a decision cut lies; conditional uncertainty matters.
  16. Subscores require evidence that their distinct information exceeds their error and dependence on the total score.
  17. A score report that invites unsupported comparisons can create invalid use even when scoring itself is technically correct.
  18. Consequences do not alone prove invalidity, but consequential uses increase the evidence and monitoring burden.[4][3]

Knowledge Transfer

Within education, the roles transfer exactly. A kindergarten numeracy interview, a secondary-school writing performance, a computerized adaptive language test, and an international science assessment all require a target construct, observations, scoring, interpretation, precision, validity, fairness, and reporting. The carrier changes from spoken explanations to essays to selected responses to interactive simulations; the inferential chain remains.

The methods also transfer across educational purposes. An item-response model can support adaptive administration, scale construction, growth reporting, or form linking, but its fit and interpretation must be rechecked for each use. Rater designs from writing assessment transfer to portfolios and performance examinations. Differential item functioning analysis transfers from admissions to language and certification contexts. Evidence-centered design transfers from standardized tests to simulations and game-based assessment because claims, evidence, tasks, and models retain their roles.[5]

Some techniques travel into psychological, occupational, clinical, and machine-learning evaluation, where Psychometrics or generic Measurement may be the proper identity. The domain-specific node does not claim those cases merely because they use reliability or latent-variable models. Exact Educational Measurement requires an educational construct or use: learning, attainment, proficiency, placement in instruction, educational certification, or policy about learners and institutions. The portable residue belongs to Measurement, Construct Validity, Validation, and Measurement Uncertainty.

Examples

Common-scale reading assessment

A statewide reading program defines a construct map spanning decoding, vocabulary, literal comprehension, and integration across texts. A blueprint samples passages and item types across grades. Students’ responses are scored, field-test statistics identify weak items, and an item-response model estimates scale locations. Anchor items link the new form to the prior scale; conditional standard errors describe precision; content and response-process studies support the intended reading interpretation; differential item functioning and accessibility review investigate whether irrelevant language, disability, or cultural demands distort results.

The state reports scale scores and achievement levels. The scale score is the measurement output; the achievement level adds a standard-setting judgment; school accountability adds a policy use. Mapped back: construct map is cognition, item design is observation, IRT and scoring are interpretation, anchors provide comparability, uncertainty qualifies claims, and validity/fairness evidence bounds use. A high correlation with last year is not enough if the new form dropped essential reading content.

Constructed-response writing task

A university uses an essay to place incoming students into writing courses. The construct includes organization, evidence, reasoning, and language control. Prompts are reviewed for opportunity and accessibility; essays are scored by trained raters using an analytic rubric; double-scoring and adjudication estimate rater consistency; exemplar papers anchor rubric interpretation; drift monitoring detects changes in severity over time. A generalizability analysis separates person, prompt, rater, and residual facets, showing whether one essay and one rater support the placement claim.

Mapped back: the essay is not the abstraction. Educational Measurement is the system connecting a writing construct to prompts, performances, rubric judgments, error facets, placement categories, fairness review, and consequences. If prompt variance dominates, adding raters alone will not repair precision; more representative prompts are needed.

Classroom fractions diagnosis

A teacher wants to know whether students compare fractions by magnitude or by an erroneous “larger denominator means larger fraction” rule. She selects pairs designed to distinguish strategies, asks students to explain reasoning, and interprets response patterns against a learning model. The output may be a qualitative profile rather than a scale score. She uses it to group students and select the next lesson, then gathers new evidence.

Mapped back: the target misconception is the construct, discriminating tasks are observation, explanations are responses, the rule-based interpretation is the model, and instructional adjustment is the use. Because the claim is narrow and local, it may need less formal evidence than a statewide score, but it still needs coherent cognition–observation–interpretation alignment.[1]

Form-to-form equating boundary

An admissions test replaces exposed items with a new form. Raw totals are not directly comparable because the forms may differ in difficulty. A common-item design estimates the relationship between forms, checks anchor representativeness and stability, and maps new-form scores onto the reporting scale. If the new form changes the construct or administration enough that interchangeability is untenable, the operation is linking rather than equating. Holland and Dorans treat linking and equating as a specialized educational-measurement problem requiring explicit design and assumptions.[6]

Mapped back: common items and populations provide the bridge, statistical modeling estimates the transformation, and invariance checks protect meaning. Merely matching means and standard deviations cannot guarantee equivalent interpretations.

Structural Tensions

  • Reliability versus validity. Standardization and repeated similar items can produce precise scores while narrowing the construct or measuring a stable surrogate. Diagnostic: what evidence supports the intended claim beyond consistency?
  • Comparability versus curricular authenticity. Common forms and scoring rules enable comparison, while locally authentic tasks may reflect richer learning. Failure occurs when comparability strips away the domain or local authenticity makes scores incomparable. Diagnostic: which differences must be invariant, and which are educationally meaningful?
  • Standardization versus accessibility. Uniform conditions protect comparability, but identical presentation can impose irrelevant barriers. Accommodations improve access only if they preserve the construct. Diagnostic: does the change remove irrelevant difficulty or alter what is measured?
  • Broad construct representation versus feasible testing time. More tasks improve domain coverage but consume time and attention. A short test may be precise on a narrow slice and invalid for the broad label. Diagnostic: what content and cognitive processes remain unsampled?
  • Statistical fit versus substantive meaning. A model can fit response patterns while its latent dimension lacks educational coherence. Diagnostic: can content and learning theory explain the scale and item ordering?
  • Continuous evidence versus categorical decisions. Scores and uncertainty are continuous, while placement, proficiency, and certification require cutoffs. Diagnostic: how stable is classification near the cut and what losses attach to false positives and negatives?
  • Transparency versus security. Learners need score meaning and assessment criteria, but disclosing operational items can increase coaching and exposure. Diagnostic: what can be transparent without destroying future comparability?
  • Technical quality versus consequential use. A score valid for low-stakes feedback may not support graduation or accountability. Diagnostic: is evidence commensurate with the ambition and stakes of the use?
  • Fairness as equal treatment versus fairness as valid inference. Identical procedures can be unfair when they add irrelevant barriers; different accommodations can improve inferential comparability. Diagnostic: does each relevant group have an opportunity to demonstrate the target construct?
  • Instructional alignment versus score inflation. Alignment can focus teaching on valued learning, but narrow coaching to predictable formats can raise scores without broader mastery. Diagnostic: does improvement generalize to fresh construct-representative tasks?
  • Stable scale versus changing domain. Longitudinal reporting rewards continuity, while curriculum and social expectations evolve. Diagnostic: has scale continuity begun to preserve an obsolete construct?
  • Human judgment versus automated scoring. Human raters provide contextual interpretation but can drift; algorithms can be consistent but opaque or biased. Diagnostic: which error sources and subgroup failures are monitored for the chosen scoring system?

Structural–Framed Character

Educational Measurement is framed, with an aggregate of \(0.75\). Its mathematical components—sampling, scale construction, error models, linking, and statistical inference—are structural. Yet the full abstraction cannot be recognized without importing educational purposes and institutions: what counts as reading proficiency, which curriculum is authoritative, which population is relevant, what opportunity to learn is expected, which uses are acceptable, and what fairness requires.

The construct and use are not discovered in the way a physical dimension is discovered. They are specified through curricular, professional, and policy judgment, then constrained by empirical evidence. Standards, cut scores, accommodations, accountability rules, and reporting categories embed values and institutional authority. This does not make results arbitrary; it makes their frame part of the measurement claim. The abstraction is reliable only when that frame is explicit and revisable.

Structural Core vs. Domain Accent

The structural core is an evidence-to-inference chain: define an attribute, create observations, transform observations through a procedure and model, quantify error, and bound conclusions by validation. Measurement supplies the generic attribute–instrument–scale–uncertainty structure. Construct Validity supplies the proxy-gap audit. Validation supplies the evidence-and-argument process. Measurement Uncertainty supplies the discipline of reporting error rather than bare values.

The domain accent is educational: learner constructs and progressions; curriculum and opportunity to learn; items, tasks, rubrics, and assessment forms; formative, summative, placement, certification, and accountability uses; comparability across forms and cohorts; accessibility and subgroup fairness; and consequences for learners and institutions. These commitments create recurrent methods that generic Measurement does not entail.

The node remains domain-specific because its full role system does not transfer literally to unrelated substrates. A thermometer needs calibration and uncertainty but not curriculum alignment, opportunity to learn, rater rubrics, score-use validity, or educational fairness. Conversely, educational scores often lack a physical unit or directly observable true quantity. The general skeleton is prime-level; the inferential and institutional package is the residual domain abstraction.

  • Measurement is the strict genus and proposed DAG parent. Educational Measurement maps a learner attribute onto a scale through an assessment instrument and procedure, yielding a value or classification with uncertainty and a declared frame.
  • Construct Validity is a mandatory audit whenever scores claim to represent a latent educational construct. It asks whether tasks and scores capture the intended competence rather than a surrogate or irrelevant barrier.
  • Validation is the continuing evidence-and-argument process supporting interpretations and uses. It is not a one-time certification.
  • Measurement Uncertainty and Observational Noise appears in task, occasion, form, rater, and response variability and in estimation error.
  • Summative Assessment names final evaluative use; Formative Assessment names evidence used to adjust learning during instruction. Either may use educational measurement, but neither covers the full lifecycle or technical identity.

Relationships to Other Abstractions

Local relationship map for Educational MeasurementParents appear above the current abstraction, mutual partners to the right, and children below. Node labels state whether each abstraction is prime or domain-specific; colors identify relation types.EducationalMeasurementDOMAINPrime abstraction: Measurement — is a kind ofMeasurementPRIMEDomain-specific abstraction: Attribute Hierarchy Method — is a kind ofAttributeHierarchy MethodDOMAIN

Current abstraction Educational Measurement Domain-specific

Parents (1) — more general patterns this builds on

  • Educational Measurement is a kind of Measurement Prime

    Measurement is the strict genus and proposed DAG parent.

Children (1) — more specific cases that build on this

  • Attribute Hierarchy Method Domain-specific is a kind of Educational Measurement

    Educational Measurement is the proposed immediate parent.

Hierarchy path (1) — routes to 1 parentless root

Neighborhood in Abstraction Space

Educational Measurement sits in a sparse region of the domain-specific corpus (87th percentile for distinctiveness): few abstractions share its structure, so a faithful description tends to retrieve it precisely.

Family — Unclustered & Miscellaneous (1565 abstractions)

Nearest neighbors

Computed from structural-signature embeddings · 2026-09-08

Not to Be Confused With

The strongest frozen catalog neighbor is Summative Assessment at \(0.773946\). Summative Assessment concerns final evaluation after a learning period. Educational Measurement may produce evidence for summative decisions, but it also covers formative diagnosis, assessment design, scoring, modeling, equating, standard setting, reporting, and validation. Purpose and measurement system are different layers.

Measurement is the closest genus but not exact coverage. Its live identity maps an attribute to a scale through an instrument and procedure with uncertainty. Educational Measurement adds learner constructs, designed performance evidence, psychometric models, score-use validation, form comparability, fairness/accessibility, and educational institutional uses. Construct Validity and Validation supply indispensable quality checks but do not construct or operate the whole measurement system.

Formative Assessment and Mastery Learning are pedagogical neighbors. Formative Assessment uses evidence to alter ongoing instruction; Mastery Learning organizes progression around demonstrated mastery. Neither entails the technical score, scale, precision, equating, or fairness apparatus. Pedagogy organizes teaching rather than the inference from sampled performance. Effect Size summarizes magnitude; it is not a learner measurement system.

Do not use Educational Measurement as an alias for educational assessment, testing, psychometrics, grading, item-response theory, Rasch measurement, test equating, standard setting, score reporting, or program evaluation. Those are broader fields, instruments, models, subprocedures, or downstream judgments. The retained node is the coordinated evidence-and-inference practice that connects them.

References

[1] National Research Council. Knowing What Students Know: The Science and Design of Educational Assessment. Washington, DC: National Academies Press, 2001. Develops the cognition–observation–interpretation assessment triangle and the requirement that the three elements support one another. registry ↩a ↩b ↩c

[2] Cook, Linda L., and Mary J. Pitoniak, eds. Educational Measurement, 5th ed. National Council on Measurement in Education and Oxford University Press, 2025. Open-access field handbook covering validity, reliability, fairness, design, scoring, modeling, equating, standard setting, reporting, instructional use, accountability, admissions, certification, and international assessment. registry

[3] American Educational Research Association, American Psychological Association, and National Council on Measurement in Education. Standards for Educational and Psychological Testing. Washington, DC: AERA, 2014. Joint professional standards for test development, evaluation, administration, interpretation, fairness, and use. registry ↩a ↩b

[4] Kane, Michael T. “Validating the Interpretations and Uses of Test Scores.” Journal of Educational Measurement 50, no. 1 (2013): 1–73. Develops an argument-based account in which proposed interpretations and uses, rather than tests in isolation, are validated. registry ↩a ↩b ↩c

[5] Mislevy, Robert J., Russell G. Almond, and Janice F. Lukas. “A Brief Introduction to Evidence-Centered Design.” ETS Research Report RR-03-16, 2003. Coordinates claims about learners, evidence, tasks, and statistical models in assessment design. registry ↩a ↩b

[6] Holland, Paul W., and Neil J. Dorans. “Linking and Equating.” In Robert L. Brennan, ed., Educational Measurement, 4th ed., 2006, pp. 187–220. Distinguishes score-linking and equating purposes, designs, and requirements. registry