Skip to content

Psychometrics, Testing & Measurement Bias

← Back to Domain-Specific Families

Abstractions about constructing, comparing, and validating tests, scales, scores, and predictive measures. They cover item response and differential functioning, adaptive testing, content validity, concordance, propensity matching, outlier measures, ideological scales, and algorithmic bias.

24 abstractions in this family — domain-specific abstractions that sit near one another in structural-signature space (k-means over structural-signature embeddings). Each is shown with its short description.

  • Algorithmic bias — A systematic and repeatable tendency of an algorithmic sociotechnical system to produce unfairly differentiated outcomes across people or categories.
  • Bennett scale — The Developmental Model of Intercultural Sensitivity, a six-orientation framework describing movement from ethnocentric responses to increasingly ethnorelative engagement with cultural difference.
  • Computerized adaptive testing — Computer-administered assessment that updates an examinee ability estimate after each response and selects subsequent items to maximize information subject to content, exposure and stopping constraints.
  • Concordance correlation coefficient — An agreement coefficient combining Pearson correlation with penalties for differences in mean and scale between two measurements.
  • Content validity — The extent to which a measure's items adequately represent every relevant facet of its intended construct or content domain.
  • Differential effects — In observational causal comparison, the outcome contrast produced by applying one treatment rather than another, distinguished from differential assignment bias that can mimic that contrast.
  • Differential item functioning — A psychometric condition in which people from different groups with the same level of the measured trait have different probabilities of an item response.
  • Employment testing — Use standardized assessments as evidence in hiring, promotion, placement, or related employment decisions, requiring job linkage, reliability, validity, fair administration, accessibility, and jurisdiction-specific legal review.
  • Ethical positioning index — A proposed brand metric combining consumer perceptions of ethical conduct with the clarity and consistency of brand positioning.
  • Hot hand — A proposed short-run elevation in success probability following recent successes, especially in repeated sports or skill performance.
  • Item analysis — A psychometric evaluation and selection process that examines candidate questions for difficulty, discrimination, redundancy, model fit, fairness and construct coverage before assembling or revising a test.
  • Item response theory — A psychometric framework modeling the probability of an item response as a function of a latent trait and item parameters.
  • Item-total correlation — The correlation between one scored assessment item and a total or rest score, used to evaluate whether the item aligns with the construct measured by the scale.
  • Johanson analysis — A multidimensional media-analysis framework evaluating representation of women and girls through presence, agency, authority, gaze, sexuality and intersectional context.
  • Martin–Quinn score — A dynamic latent-variable estimate of each U.S. Supreme Court justice's ideological position inferred from voting alignments across terms.
  • Minnesota Paper Form Board Test — A paper-and-pencil spatial-visualization test asking examinees to identify which intact figure can be assembled from a displayed set of separated component shapes.
  • Outliers ratio — A legacy objective-video-quality metric reporting the fraction of model predictions lying outside a declared tolerance interval around subjective mean-opinion scores.
  • Pairwise comparison (psychology) — A psychometric elicitation method that presents two stimuli at a time and records a preference, similarity, or relative-attribute judgment for later scale estimation.
  • Predictive value of tests — The probability that a target condition is present or absent given a test result, combining test performance with condition prevalence in the population of use.
  • Propensity score matching — An observational causal-inference method that matches treated and untreated units with similar estimated probabilities of treatment given observed covariates.
  • Slope One — A family of item-based collaborative-filtering algorithms that predicts a user's rating from average pairwise rating differences between items and the user's ratings of neighboring items.
  • Thurstone scale — An equal-appearing-interval attitude scale built by having judges locate statement favorability and selecting items with separated medians and low ambiguity.
  • Thurstonian model — A latent-variable model that explains discrete choices, rankings or ordered responses by comparing noisy continuous psychological values, commonly modeled as jointly normal.
  • Wilson–Patterson Conservatism Scale — A survey instrument estimating liberal–conservative ideology from respondents’ approval or disapproval of a set of political and social issue labels.