Skip to content

Tensions in Practice: Checkable procedures in tension with richer intent

Delivering instructions that people can use

Suppose a service aims to help people use instructions. It can turn one part of that aim into a checkable requirement: deliver the complete instruction set. A delivery procedure can satisfy that requirement while readers still fail to understand what to do. Keeping the remaining intent visible allows a separate check of use, but that check takes work and judgment; it does not turn a rich purpose into a complete mechanical certificate.

Make execution checkable

State a requirement clearly enough to build a procedure and verify its result.

Preserve the wider purpose

Keep the user’s ability to use the instructions in view when assessing success.

Why these aims pull against each other

A thin requirement is easy to check because it omits parts of the wider intent. Treating success against that requirement as proof of the entire purpose hides the omitted remainder.

Compare the arrangements

Certify delivery alone

Translate the aim into a complete-delivery requirement, run the delivery procedure, and check whether the required instruction set arrived.

What it protects
The delivery contract is explicit and can be checked reproducibly.
What it costs
Understanding and successful use remain outside this certificate; announcing the whole aim as achieved would overstate what was checked.
When it fits
Sufficient only for a delivery claim, or where other justified arrangements address the remaining intent.

Illustration note: The delivery example instantiates the source’s spec-reduction risk. The diagram intentionally labels the certificate narrowly instead of endorsing an invalid claim of complete success.

Keep the remainder visible

Keep the same delivery procedure and certificate. Separately ask intended users to try the instructions, then assess the delivery result alongside this evidence and its gaps.

What it protects
A passed delivery check no longer silently substitutes for the whole purpose; failures of use can become visible.
What it costs
Tryouts and interpretation require effort, can miss users or contexts, and may leave disagreements about what adequate use means.
When it fits
The additional check must address the actual intended use. A convenient satisfaction score alone could repeat the original reduction.

Illustration note: The tryout is an editorial response to the source’s instruction to retain the unformalized remainder. It is evidence for judgment, not proof that every reader will succeed.

What this illustration does—and does not—establish

Operationalization: Checkable specification versus inexpressible intent (measurement) identifies the risk that making a purpose checkable turns it into a thin proxy. Its core distinguishes the specification, executable procedure, and correctness contract. The instruction-delivery example and reader tryout are editorial arrangements under those limits.

  • The scenario is editorial; it is not a report of actual reader testing of this website or its instructions.
  • Successful execution must still satisfy the stated specification. A good wider intention does not excuse failure of the agreed delivery contract.
  • No incentive loop or deliberate gaming is needed for specification reduction; this is distinct from a measure degrading after becoming a target.
  • A richer checklist or another proxy can still omit important intent. The second arrangement exposes a remainder rather than claiming to eliminate it.

Source entries

Operationalization

Prime · Source of the tension

The source’s Checkable specification versus inexpressible intent (measurement) passage supplies this contextual tension. The arrangements below are bounded editorial illustrations, not an additional empirical finding.

T6 — Checkable specification versus inexpressible intent (measurement).

T6 — Checkable specification versus inexpressible intent (measurement). Operationalization presumes the specification can be stated precisely enough that "does the procedure satisfy it?" is answerable — but some intents resist crisp specification (vague legislative purpose, a contested latent construct), and forcing a checkable spec can distort the intent into what is measurable. Here the boundary is with operationalization's own risk of goodhart substitution. The failure mode is *spec reduction*: replacing an irreducibly rich intent with a thin checkable proxy, then discharging the proxy faithfully while the real intent is unserved. Diagnostic: check whether the specification captures the intent or merely a measurable shadow of it — where intent exceeds what can be crisply specified, the residual gap includes the un-formalised remainder, and procedural correctness against the proxy does not certify the intent.

Read the source section

What the operation commits to

The essential commitment is a *level-of-description shift* paired with a *correctness contract*: the procedure is not merely "a way of doing it" but a way that is meant to discharge the specification, and the question "does this procedure satisfy this specification?" is a meaningful one with substrate-specific answers.

Read the source section

A gap can exist before anyone targets it

- Not goodharts_law. Goodhart concerns a measure degrading once targeted; operationalization's adjacent risk is *spec reduction* (replacing rich intent with a thin checkable proxy), which is a failure mode of lowering, not the prime — though the two meet where the proxy is then optimised.

Read the source section