Tensions in Practice: Exact representation in tension with equivalent-form reuse¶
Structured records · declared member-order equivalence
The records “a:1, b:2” and “b:2, a:1” contain the same fields and values but different character sequences. Hashing their original bytes can keep those representations apart. Sorting fields first makes them share a comparison form, provided member order has no meaning in this application. The hash does not decide that equivalence; the normalization rule does.
Preserve representational distinctions
Detect a change in the original sequence, including order or formatting when those matter.
Reuse genuinely equivalent records
Keep irrelevant spelling or ordering differences from creating separate comparison keys.
Why these aims pull against each other
A normalization rule gains matches by deliberately erasing distinctions before fingerprinting. Whether those distinctions are irrelevant is an application contract, not a property of the hash.
Choose an arrangement to see what changes and what remains difficult.
Finite illustrative comparisons. Labels carry the meaning; color does not establish a preference or measured effect.
What this choice protects
What it costs
When it fits
Compare the arrangements
Fingerprint the raw form
Hash the original representation without rearranging its members. Tiny records with distinct fields. F1/F2/F3 stand for distinct fingerprints in this stipulated collision-free example; they are labels, not a proposed hash function.
| Hash input | Token | |
|---|---|---|
| a:1, b:2 | a:1, b:2 | F1 |
| b:2, a:1 | b:2, a:1 | F2 |
| a:1, b:3 | a:1, b:3 | F3 |
- What it protects
- Different representations remain distinguishable in this collision-free toy.
- What it costs
- Equivalent-for-purpose records can occupy separate comparison keys and miss reuse opportunities.
- When it fits
- Fits byte-sensitive provenance or a record contract where order or formatting is meaningful.
Illustration note: The table stipulates distinct tokens for these different inputs. Different bytes do not universally imply different hashes.
Fingerprint a canonical form
Sort these unique fields into one specified order before hashing on both sides of comparison. The declared record meaning ignores member order, not values. Sort fields before hashing: the first two records share F1; changing b from 2 to 3 still produces a different comparison form.
| Hash input | Token | |
|---|---|---|
| a:1, b:2 | a:1, b:2 | F1 |
| b:2, a:1 | a:1, b:2 | F1Order ignored |
| a:1, b:3 | a:1, b:3 | F3 |
- What it protects
- The first two records reuse one key while the changed value remains distinct.
- What it costs
- Normalization adds work and can hide meaningful differences if the equivalence rule is wrong or changes.
- When it fits
- Fits a record type whose member order is irrelevant and a versioned, consistently applied rule.
Illustration note: Only member order changes here. No Unicode, numeric-precision or general JSON-standard guarantee is imported.
What this illustration does—and does not—establish
Hashing: Fingerprint versus Meaning supplies the distinction between bits and chosen meaning; Canonicalization Pipeline supplies the risks of erasing distinctions and inconsistent placement.
- Neither a raw nor canonical digest reconstructs the original; retain originals separately when they matter.
- A token match is not a mathematical proof of object identity; collision handling remains necessary for the application.
- Duplicate field names and types are excluded from this tiny record contract.
Source entries
Hashing
Hashing: Fingerprint versus Meaning supplies the conflict examined here.
Fingerprint versus Meaning
Diagnostic: ask whether equality should be over *bits* or over *meaning*; if meaning, canonicalize the input before hashing rather than expecting the hash to understand it.
Canonicalization Pipeline
Supplies explicit normalization rules, their scope and original-retention obligations.
Tuning parameters
- Rule-set aggressiveness — how many distinctions the rules erase. More rules fold more variants together but risk collapsing inputs that were meaningfully different. - Placement — ingest-side only, query-side only, or both. Normalizing on only one side reintroduces mismatches between what was stored and what is looked up. - Original retention — whether the pre-canonical form is discarded or kept alongside the canonical one. Keeping it preserves the ability to recover an over-normalized distinction.