Round-Trip Assessment Redesign Test¶
Test / assessment — instantiates Identity-Safe Performance Context
Simulates participant and evaluator journeys to verify that cues, criteria, feedback, review, data flow, and monitoring work together.
A redesigned evaluation can pass every component review and still fail as a system: the cue audit removed the demographic screen, but the accommodation form re-introduces it; the appeal channel exists, but no one on the participant path is ever shown how to reach it; the data-separation rule is written, but a report export quietly re-joins identity to scores. Round-Trip Assessment Redesign Test is the pre-deployment dress rehearsal that catches these seams. It walks the whole loop — a simulated participant goes end-to-end through the real experience, and a simulated evaluator goes end-to-end through scoring and review — to verify the pieces actually interlock. Its defining property is being a one-time, integration-level dry run whose unit of analysis is the handoffs between components, not any single component. It exercises the pathways, the data flow, and the appeal path as a live assembly. It is not the ongoing monitor that watches real populations over time, and it is not the static audit that catalogs cues on paper.
Example¶
Before launching a redesigned civil-service hiring assessment, an agency runs a round-trip test. A team member plays a candidate from a stereotyped group and moves through the entire real journey: the invitation email, the registration flow, the (now deferred) demographic questions, the instructions, the work-sample task, the result notice, and an attempt to file a concern. A second member plays the evaluator, scoring the sample against the rubric and then processing the candidate's flag through the review channel. The rehearsal surfaces exactly the interaction failures a component review missed: the accommodation-request page still asks demographics before the task; the results email names no review route; and when the evaluator exports scores for the equity team, the export template re-attaches candidate names to identity fields.
The setup is a controlled simulation, run once before go-live and again after major changes. The intended outcome is a punch-list of integration defects — each a place where two individually-correct components fail at their seam — fixed before a real candidate ever hits them.
How it works¶
The test is defined by walking both journeys and instrumenting the joins:
- Two protagonists. A participant persona and an evaluator persona traverse the real system, because threats and leaks live in the transitions each one experiences, not in isolated screens.
- Walk the mapped pathways. For each hypothesized threat pathway from the audit, the participant persona checks whether the redesign actually closes it in the lived flow — confirming or refuting the map.
- Trace the data. The evaluator persona follows an identity field through collection, storage, scoring visibility, and export, verifying separation holds all the way to the last handoff.
- Exercise the exits. The participant persona actually tries to contest a result, confirming the appeal path is discoverable and walkable end-to-end, not just present in a policy.
- Punch-list, not pass/fail. Output is a list of seam defects with owners, re-run after fixes.
Tuning parameters¶
- Journey fidelity — paper walkthrough vs. live system with real accounts. Higher fidelity catches real leaks but costs setup.
- Persona coverage — one representative path vs. multiple intersecting identities and edge cases. Broader coverage finds more seams but multiplies runs.
- Adversarial stance — cooperative walk vs. actively trying to break separation and discover cues. A more adversarial run surfaces more but can chase implausible paths.
- Re-run trigger — launch-only vs. re-test after every material change. Frequent re-tests prevent regression but consume time.
- Scripting depth — free exploration vs. a fixed checklist tied to each mapped pathway. Scripts ensure coverage; free exploration finds the unscripted seam.
When it helps, and when it misleads¶
Its strength is catching emergent failures — the ones that exist only because two correct components meet badly — before a real participant absorbs them. It is essentially cognitive pretesting applied to an entire evaluation system: walking the artifact as a real user would to expose breakdowns a designer's-eye review cannot see.[n1]
Its failure mode is the rehearsal that certifies and then rots: a clean dry run at launch is treated as permanent proof, while later changes quietly re-open the seams it closed. A related misuse is confusing the dry run with real evidence — a simulated participant is not a stereotyped population under genuine stakes, so a passing round-trip test says the plumbing connects, not that the assessment is valid and equitable in the field. The guarding discipline is to re-run after material changes and to hand the ongoing question — does it actually behave validly across groups over time? — to the monitor, treating the round-trip test as an integration check, never as the equity verdict.
How it implements the components¶
identity_threat_pathway_map— it validates the map in the lived flow, walking each hypothesized pathway to confirm (or refute) that the redesign closes it, and updating the map with what the rehearsal reveals.identity_data_separation_boundary— it traces an identity field through the entire pipeline to the final export, verifying that collection stays decoupled from scoring at every handoff.identity_safe_contestation_path— it exercises the appeal route end-to-end, confirming a participant can actually discover and walk it, not merely that a policy names it.
It does NOT run the ongoing, population-level validity-and-equity monitoring (performance_validity_and_equity_monitor) — that is the Subgroup Outcome-Validity Dashboard; nor does it build the static cue catalog (evaluative_cue_inventory), which is the Identity-Cue Audit. The round-trip test is a one-time rehearsal of the assembled whole.
Related¶
- Instantiates: Identity-Safe Performance Context — supplies the pre-deployment integration-verification layer.
- Consumes: Identity-Cue Audit — it walks the threat pathways the audit mapped to check they are closed in the live assembly.
- Sibling mechanisms: Identity-Cue Audit · Subgroup Outcome-Validity Dashboard · Identity-Question Timing Protocol · Identity-Safe Review Channel
Editorial Notes¶
Form Classification¶
Form family: Experiment, Test & Rehearsal
Rationale: Round-Trip Assessment Redesign Test operates as an active test, trial, simulation, drill, or rehearsal that generates evidence through a deliberate attempt or perturbation because it simulates participant and evaluator journeys to verify that cues, criteria, feedback, review, data flow, and monitoring work together.
Independent corroboration: The frozen evidence defines Round-Trip Assessment Redesign Test as 'Simulates participant and evaluator journeys to verify that cues, criteria, feedback, review, data flow, and monitoring work together', so its operative form is Experiment, Test & Rehearsal.
Nearest alternative: Assessment, Review & Assurance — Round-Trip Assessment Redesign Test includes features of a bounded evaluation of existing evidence or work that produces a finding or disposition, but its defining operation is an active test, trial, simulation, drill, or rehearsal that generates evidence through a deliberate attempt or perturbation.
Review outcome: Independent reviewer agreement; medium confidence.
Origin Attribution¶
Primary origin: Education & Pedagogy
Origin pattern: Cross-disciplinary synthesis
Present-day reach: Multi-domain
Rationale: Testing the full participant-evaluator assessment journey is rooted in educational assessment design.
Related originating lineages:
- Human-Computer Interaction — Journey simulation materially tests cues, feedback, and data flow.
- Organizational & Management Science — Process redesign independently checks end-to-end governance and monitoring.
- Psychology — Experimental, clinical, and behavioral psychology supplies a parallel or contributing lineage for the mechanism's defining operation: simulates participant and evaluator journeys to verify that cues, criteria, feedback, review, data flow, and monitoring work together.
Review resolution: Both blind reviewers agree that education_pedagogy is the primary historical origin. Explicit reconciliation of alternate origin disagreement starts from reviewer_a’s mechanism-specific evidence: Testing the full participant-evaluator assessment journey is rooted in educational assessment design. Reviewer A proposed alternates=human_computer_interaction, organizational_management, origin_mode=cross_disciplinary_synthesis, domain_reach=multi_domain, and encyclopedia_synthesis=true; reviewer B proposed alternates=human_computer_interaction, psychology, origin_mode=cross_disciplinary_synthesis, domain_reach=multi_domain, and encyclopedia_synthesis=true. The final record retains every independently supported alternate from either review (human_computer_interaction, organizational_management, psychology) without an arbitrary cap, selects origin_mode=cross_disciplinary_synthesis to represent the combined lineage evidence, and keeps domain_reach=multi_domain and encyclopedia_synthesis=true from the more mechanism-specific assessment. Present-day transfer is recorded as reach and is not treated as proof of historical origin.
Encyclopedia synthesis: The exact catalogued form synthesizes established practice rather than reproducing a single standard historical label.
Review outcome: Reconciled after independent review; medium confidence.
Notes¶
[n1] Cognitive pretesting — walking real people through an instrument (a survey, a form, a task) and observing where their lived path breaks, misleads, or diverges from designer intent before fielding it. The round-trip test applies the same lived-walkthrough logic to a whole evaluation system. ↩