Skip to content

Tensions in Practice: Exact representation in tension with equivalent-form reuse

Structured records · declared member-order equivalence

The records “a:1, b:2” and “b:2, a:1” contain the same fields and values but different character sequences. Hashing their original bytes can keep those representations apart. Sorting fields first makes them share a comparison form, provided member order has no meaning in this application. The hash does not decide that equivalence; the normalization rule does.

Preserve representational distinctions

Detect a change in the original sequence, including order or formatting when those matter.

Reuse genuinely equivalent records

Keep irrelevant spelling or ordering differences from creating separate comparison keys.

Why these aims pull against each other

A normalization rule gains matches by deliberately erasing distinctions before fingerprinting. Whether those distinctions are irrelevant is an application contract, not a property of the hash.

Compare the arrangements

Fingerprint the raw form

Hash the original representation without rearranging its members. Tiny records with distinct fields. F1/F2/F3 stand for distinct fingerprints in this stipulated collision-free example; they are labels, not a proposed hash function.

Original representations remain separate
Hash inputToken
a:1, b:2a:1, b:2F1
b:2, a:1b:2, a:1F2
a:1, b:3a:1, b:3F3
What it protects
Different representations remain distinguishable in this collision-free toy.
What it costs
Equivalent-for-purpose records can occupy separate comparison keys and miss reuse opportunities.
When it fits
Fits byte-sensitive provenance or a record contract where order or formatting is meaningful.

Illustration note: The table stipulates distinct tokens for these different inputs. Different bytes do not universally imply different hashes.

Fingerprint a canonical form

Sort these unique fields into one specified order before hashing on both sides of comparison. The declared record meaning ignores member order, not values. Sort fields before hashing: the first two records share F1; changing b from 2 to 3 still produces a different comparison form.

Member order is ignored; values remain distinct
Hash inputToken
a:1, b:2a:1, b:2F1
b:2, a:1a:1, b:2F1Order ignored
a:1, b:3a:1, b:3F3
What it protects
The first two records reuse one key while the changed value remains distinct.
What it costs
Normalization adds work and can hide meaningful differences if the equivalence rule is wrong or changes.
When it fits
Fits a record type whose member order is irrelevant and a versioned, consistently applied rule.

Illustration note: Only member order changes here. No Unicode, numeric-precision or general JSON-standard guarantee is imported.

What this illustration does—and does not—establish

Hashing: Fingerprint versus Meaning supplies the distinction between bits and chosen meaning; Canonicalization Pipeline supplies the risks of erasing distinctions and inconsistent placement.

  • Neither a raw nor canonical digest reconstructs the original; retain originals separately when they matter.
  • A token match is not a mathematical proof of object identity; collision handling remains necessary for the application.
  • Duplicate field names and types are excluded from this tiny record contract.

Source entries

Hashing

Prime · Source of the tension

Hashing: Fingerprint versus Meaning supplies the conflict examined here.

Fingerprint versus Meaning

Diagnostic: ask whether equality should be over *bits* or over *meaning*; if meaning, canonicalize the input before hashing rather than expecting the hash to understand it.

Read the source section

Canonicalization Pipeline

Mechanism · Related concept

Supplies explicit normalization rules, their scope and original-retention obligations.

Tuning parameters

- Rule-set aggressiveness — how many distinctions the rules erase. More rules fold more variants together but risk collapsing inputs that were meaningfully different. - Placement — ingest-side only, query-side only, or both. Normalizing on only one side reintroduces mismatches between what was stored and what is looked up. - Original retention — whether the pre-canonical form is discarded or kept alongside the canonical one. Keeping it preserves the ability to recover an over-normalized distinction.

Read the source section