Skip to content

Usability or Field Test

Validation method — instantiates Affordance Shaping

Puts the shaped affordance in front of real users in a realistic setting and records what they actually do, so the design is judged by behaviour rather than by the designer's intention.

Every other mechanism in this archetype proposes; Usability or Field Test checks. It takes a shaped affordance and puts it in front of representative users doing real tasks, then records what they actually do — where they hesitate, misread, work around, or fail — rather than what they say or what the designer hoped. Its defining commitment is behaviour over intention and reality over the lab: it is the mechanism that can contradict the capability model and the cue design with evidence from actual use. That is why it earns its place next to a field setting, where the design has to survive real light, real hands, real interruptions, and real stakes.

Example

An agronomy team ships a crop-advisory app that tells smallholder farmers when to spray. In the office it demos flawlessly. The field test tells a different story. Researchers sit with farmers using their own phones, on their own plots, and simply watch and record. The "spray now" alert — obvious on an office monitor — washes out in direct sunlight; the offline mode, needed where there is no signal, hides its most important button one screen too deep; and several farmers, gloves on, tap the wrong row entirely.[n1]

None of that appeared in the lab. Because the test captured the actual behaviour stream — the taps, the pauses, the workarounds — the team could see precisely which shaped affordance failed and in which real condition. The design's intention was fine; its fit to the ecology was not, and only running it there revealed the gap.

How it works

  • Run it where it will live. The affordance is exercised in (or as close as possible to) the real setting — the field, not a clean room — so the conditions that break designs are present.
  • Capture behaviour, not opinion. What matters is the recorded stream of actions, hesitations, errors, and workarounds; stated preference is treated as a weak secondary signal.
  • Use representative users on real tasks. Participants resemble the intended population and do the actual job, so the observed failures are the ones that will really occur.
  • Hunt the intention-behaviour gap. The analysis targets exactly where what people did diverged from what the design assumed.

Tuning parameters

  • Lab vs. field — controlled setting versus real environment. The lab isolates causes; the field surfaces the conditions that actually break the design.
  • Sample size and representativeness — how many users, how well they match the population. A handful of representative users catches most gross failures; edge cases need more.
  • Task realism — tightly scripted tasks versus natural, self-directed use. Scripts compare cleanly; natural use reveals the unexpected.
  • Moderation — think-aloud probing versus unobtrusive observation. Probing explains why; silent watching keeps the behaviour undisturbed.
  • Instrumentation depth — notes only, or recorded traces, timings, and error logs. Richer capture finds subtler failures at higher analysis cost.

When it helps, and when it misleads

Its strength is that it is the one mechanism grounded in what people do, in conditions that matter — so it catches the failures that analysis predicted away and that design meetings never imagined. It is the reality check that keeps the rest of the archetype honest.

Its failure modes come from letting the check bend toward the answer the team wants. Small or unrepresentative samples get over-generalized; being watched changes behaviour; and, most insidiously, the test becomes usability theatre — run late, on friendly users, with tasks steered so the shipped design passes, its failures quietly dropped from the report. The discipline that guards against this is to pre-specify the tasks and success metrics before running, recruit users who genuinely represent the population, separate what users did from what they said, and report the failures as findings rather than embarrassments.

How it implements the components

Usability or Field Test fills the empirical, validation side of the archetype — the parts only real use can produce:

  • feedback_trace_capture — its core output: the recorded stream of what users actually did — actions, hesitations, errors, workarounds — attributable to specific shaped affordances.
  • ecological_fit_probe — by running in the real setting, it tests whether the affordance survives the conditions of its actual ecology rather than the tidy ones of design.

It validates; it does not design. The cues it tests come from Signifier Prototyping, the capability model it can refute from Task and Capability Analysis, and the default it observes from Safe Default or Preselected Path. Choosing statistically between variants at scale is Prototype A/B or Multivariate Test, and watching unprompted behaviour in the wild is Desire Path Observation.

Editorial Notes

Form Classification

Form family: Experiment, Test & Rehearsal

Rationale: Usability or Field Test operates as an active test, trial, simulation, drill, or rehearsal that generates evidence through a deliberate attempt or perturbation because it puts the shaped affordance in front of real users in a realistic setting and records what they actually do, so the design is judged by behaviour rather than by the designer's intention.

Independent corroboration: The frozen evidence defines Usability or Field Test as 'Puts the shaped affordance in front of real users in a realistic setting and records what they actually do, so the design is judged by behaviour rather than by the designer's intention', so its operative form is Experiment, Test & Rehearsal.

Nearest alternative: Assessment, Review & Assurance — Usability or Field Test includes features of a bounded evaluation of existing evidence or work that produces a finding or disposition, but its defining operation is an active test, trial, simulation, drill, or rehearsal that generates evidence through a deliberate attempt or perturbation.

Review outcome: Independent reviewer agreement; medium confidence.

Origin Attribution

Primary origin: Human-Computer Interaction

Origin pattern: Single lineage

Present-day reach: Universal

Rationale: Both independent reviews identify human computer interaction as the historical home of the operation—Puts the shaped affordance in front of real users in a realistic setting and records what they actually do, so the design is judged by behaviour rather than by the designer's intention.. The retained alternates document formative adjacent traditions; the reach field, not the origin field, carries later applicability.

Related originating lineages:

  • Computer Science & Software Engineering — Computer science and software-engineering practice supplies a parallel or contributing lineage for the mechanism's defining operation: puts the shaped affordance in front of real users in a realistic setting and records what they actually do, so the design is judged by behaviour rather than by the designer's intention.
  • Engineering & Design — Engineering design, reliability, and systems-safety practice supplies a parallel or contributing lineage for the mechanism's defining operation: puts the shaped affordance in front of real users in a realistic setting and records what they actually do, so the design is judged by behaviour rather than by the designer's intention.
  • Ethnography & Qualitative Methods — Ethnography and qualitative comparative inquiry supplies a parallel or contributing lineage for the mechanism's defining operation: puts the shaped affordance in front of real users in a realistic setting and records what they actually do, so the design is judged by behaviour rather than by the designer's intention.
  • Psychology — Psychology's perception, cognition, behavior, and risk-communication tradition contributes a separate formative lineage to the mechanism's usability or field test logic.

Review resolution: Both blind reviewers independently place the defining operation—Puts the shaped affordance in front of real users in a realistic setting and records what they actually do, so the design is judged by behaviour rather than by the designer's intention.—in human computer interaction. Their queued differences are secondary: alternate_origin_disagreement, origin_mode_disagreement, domain_reach_disagreement, encyclopedia_synthesis_disagreement. Reviewer A uniquely contributes ['psychology']; reviewer B uniquely contributes ['computer_science', 'engineering_design', 'ethnography_qualitative_methods']. I preserve the full evidence-supported union of 4 alternate domain(s), without a numeric cap. origin_mode=single_lineage reflects the more specific lineage judgment in reviewer B's evidence, while domain_reach=universal separately records present-day portability. The affirmative encyclopedia-synthesis finding is preserved, and confidence=high uses the more conservative reviewer level.

Encyclopedia synthesis: The exact catalogued form synthesizes established practice rather than reproducing a single standard historical label.

Review outcome: Reconciled after independent review; high confidence.

Notes

Do not conflate it with an A/B or multivariate test. A field test is usually small-n, qualitative, and in situ; it tells you why a design fails and surfaces the failures nobody predicted. An A/B test is large-n and quantitative; it tells you which variant wins but not why. They are complementary — using one to do the other's job is the common boundary error.

[n1] Ecological validity is the degree to which findings obtained in a study generalize to real-world conditions (a notion associated with Egon Brunswik). A field test deliberately trades some experimental control for ecological validity, which is exactly why it can catch failures a lab study designs away.