Skip to content

Example, Counterexample, and Decision-Task Test

Test / assessment — instantiates Bidirectional Conceptual Translation

Puts representative, boundary, misleading, exception, absent-counterpart, and refusal cases to real users and compares how they classify, explain, rate confidence, and act across both frameworks.

Example, Counterexample, and Decision-Task Test is the performance check. It hands representative users a battery of cases — typical ones, near-boundary ones, misleading same-label "false friends," exception cases, absent-counterpart cases, and cases where the right answer is to ask for context or refuse to substitute — and watches what they actually do with the bridge: how they classify, how they explain, how confident they are, and which action they choose, in each framework. Its defining move is that superficial verbal agreement counts for nothing; the evidence is behavior on consequential cases, compared across frameworks, with disproportionate attention paid to boundaries and exceptions. Where an independent review checks whether meaning survived transmission on paper, this test checks whether people reason and act the same way once the bridge is in their hands — the thing text-level checks cannot prove.

Example

A team shipping an automated-driving feature keeps colliding over the word "safe." The engineers mean a hazard rate driven As Low As Reasonably Practicable; the legal team means a defensible duty of care with disclosed residual risk. Rather than argue definitions, they run a decision-task test. Both groups get the same case set: a routine scenario ("does the feature meet 'safe' here?"), a boundary case (a hazard reduced but not eliminated), a false-friend case (statistically low risk but an undisclosed known failure mode), an exception case (a vulnerable-road-user edge the metric doesn't cover), and a refusal case (insufficient data to claim "safe" at all). Each person classifies, explains, rates confidence, and states the action they'd take.

The patterns diverge exactly where predicted: on the false-friend case the engineers pass it while the lawyers flag non-disclosure; on the refusal case the engineers produce an estimate while the lawyers withhold the claim. Those divergences become concrete evidence — feeding the crosswalk (mark "safe" contested and context-conditioned) and the loss ledger — rather than a stalemate over whose "safe" is the correct one.

How it works

  • Build a covering case set. Typical, boundary, false-friend, one-to-many, absent-counterpart, and refuse-or-ask-for-context cases, with expected interpretation and action derived from the invariant set.
  • Recruit real users of each framework. People who will actually reason with the bridge, not raters of fluency.
  • Elicit behavior, not approval. Ask for classification, explanation, confidence, and chosen action in both frameworks — the four together, so a confident wrong action is visible.
  • Compare across frameworks and route results. Divergences and boundary failures feed the crosswalk and the ledger; they are not usability scores.

Tuning parameters

  • Case-mix ratio — how the set is weighted across typical, boundary, adversarial, absent-counterpart, and decision cases. Weighting toward rare high-harm exceptions catches costly failures; weighting toward frequency reflects everyday use.
  • User sampling — who is tested, across expertise, lived impact, and regional or institutional variation. Broader sampling detects more divergence but costs time and can fatigue participants.
  • Elicitation richness — classification only, or classification plus explanation, confidence, and action. Richer elicitation exposes confident-but-wrong reasoning; leaner elicitation scales.
  • Pass criterion — how much cross-framework agreement, on which cases, counts as passing. Segmenting by case type keeps a high overall pass rate from hiding a systematic failure on a protected exception.

When it helps, and when it misleads

Its strength is that it is the only check measuring use rather than wording, so it catches the bridge that reads well and still produces different decisions — the failure the whole archetype exists to prevent. Boundary and refusal cases surface exactly where a fluent translation breaks.

Its classic trap is testing only typical cases and reporting a high pass rate, so a systematic failure on a rare protected exception is averaged away — surface fluency confirmed, the dangerous edge never probed.[1] It can also drift into a satisfaction survey ("did this feel clear?") that rates comfort instead of correctness. The guarding discipline is to over-weight boundary, exception, and refusal cases relative to their frequency, elicit action and confidence rather than approval, and segment results so a minority-category failure stays visible.

How it implements the components

  • bidirectional_example_counterexample_and_boundary_test_set — it constructs the case set (typical, boundary, false-friend, absent-counterpart, refusal), with expected answers derived from the invariants.
  • native_participant_interpretation_and_use_validation — it recruits real users of each framework to classify, explain, rate confidence, and act, validating use rather than fluency.

It measures whether users act correctly but does NOT run the blind forward/back round-trip that checks meaning-in-transmission (that's the Back-Translation Review), define the invariants that set the expected answers (that's the Dual-Framework Map), or audit who gains authority and who bears the burden across affected parties (that's the Power-Effect Audit).

Editorial Notes

Form Classification

Form family: Experiment, Test & Rehearsal

Rationale: Example, Counterexample, and Decision-Task Test operates as a bounded trial, probe, simulation, or rehearsal that generates evidence from performance because it puts representative, boundary, misleading, exception, absent-counterpart, and refusal cases to real users and compares how they classify, explain, rate confidence, and act across both frameworks.

Independent corroboration: The frozen evidence defines Example, Counterexample, and Decision-Task Test as 'Puts representative, boundary, misleading, exception, absent-counterpart, and refusal cases to real users and compares how they classify, explain, rate confidence, and act across both frameworks', so its operative form is Experiment, Test & Rehearsal.

Review outcome: Independent reviewer agreement; high confidence.

Origin Attribution

Primary origin: Human-Computer Interaction

Origin pattern: Cross-disciplinary synthesis

Present-day reach: Multi-domain

Rationale: Human-computer interaction and user research formalized task-based testing with representative users, observed classifications, explanations, confidence, and action outcomes.

Related originating lineages:

  • Ethnography & Qualitative Methods — Cross-framework interpretive elicitation materially contributes participant explanation and contextual probing.
  • Linguistics & Semiotics — Contrastive examples, counterexamples, and classification judgments are rooted in semantic fieldwork and concept elicitation.
  • Statistics & Experimental Design — Cohort selection, task comparison, and response measurement materially draw on experimental design.

Review resolution: The mechanism's defining step is testing meanings through real user decisions, placing it in HCI. Linguistic boundary cases, designed comparisons, and qualitative explanation probes are retained as formative inputs.

Encyclopedia synthesis: The exact catalogued form synthesizes established practice rather than reproducing a single standard historical label.

Review outcome: Researched adjudication after independent review; high confidence.

Sources consulted:

References

[1] Buolamwini, Joy, and Timnit Gebru. "Gender Shades: Intersectional Accuracy Disparities in Commercial Gender Classification". Proceedings of the 1st Conference on Fairness, Accountability and Transparency, PMLR 81: 77–91 (2018). Shows that aggregate accuracy can conceal systematic failure for a less represented protected subgroup. registry