Calibration Workshop¶
Facilitated ritual — instantiates Prototype-Centered Category Modeling
Convenes the people who judge the category to align on shared reference cases and on how context reweights typicality, so their independent calls converge.
When several people apply the same fuzzy category, they will disagree — not because someone is wrong, but because each carries a slightly different prototype in their head. A Calibration Workshop is the facilitated session that pulls those private prototypes into one shared one. Its distinguishing move is that it works on the judges, not the cases: they score the same examples independently, then surface and argue the gaps until they converge on a common set of reference cases and a shared understanding of how context shifts the call. The output is not a decision about any one item — it is a group of people who will now make the same decision the next time they are apart.
Example¶
A company's hiring committees keep disagreeing about what a "strong hire" looks like — one interviewer's clear yes is another's maybe. Before the next round they run a calibration workshop. Everyone independently rates the same six anonymised past candidate write-ups on a "strong hire" scale, then the spread is put on the wall. The loud disagreements become the agenda: one interviewer has been weighting polish, another substance. They talk it through, agree on two write-ups as the canonical strong hires and one as the canonical near-miss, and write down a weighting rule — for junior roles trajectory outweighs current polish; for senior roles the reverse. Those agreed examples and that context rule become the reference the whole committee carries into live interviews, so two interviewers who never speak still grade the next candidate the same way.
How it works¶
The distinguishing mechanic is independent scoring, then reconciled disagreement. Judges rate a shared set blind to each other, the spread is made visible, and the divergences — not the agreements — drive the discussion, because that is where the private prototypes differ. Reconciliation produces two durable artifacts: a small library of agreed reference cases (the examples everyone now points to) and an explicit rule for how context reweights the call. It is a recurring ritual, not a one-off — re-run periodically, it keeps a group of judges from drifting apart again.
Tuning parameters¶
- Case-set difficulty — calibrating on easy cases wastes the session; the payoff is in deliberately including the borderline examples where judges split.
- Blind vs. open scoring — scoring independently before discussion stops the loudest voice anchoring everyone; open scoring is faster but converges on personality, not judgement.
- Convergence target — how tight an agreement to demand before stopping (full consensus vs. "close enough"); chasing perfect agreement on the rim can burn the room out.
- Participant mix — who is in the room decides whose prototype becomes the shared one; omit a constituency and you calibrate to a partial view.
- Cadence — how often to re-convene; frequent workshops hold alignment tight but cost everyone's time.
When it helps, and when it misleads¶
Its strength is attacking inconsistency between judges directly — the failure mode where a category is nominally shared but privately means something different to each person applying it. By reconciling on real cases it raises inter-rater agreement in a way no written definition alone achieves.[1] The reference library and weighting rule it leaves behind are what let independent judges stay aligned between sessions.
It misleads when the room converges on the wrong prototype with new confidence — calibration produces agreement, which is not the same as accuracy, and a biased or unrepresentative panel will simply standardise its bias. It can also enforce a false consensus that silences a minority reading which was actually catching something real. The disciplines are to seed the panel across the perspectives that matter, to check the agreed prototype against an external benchmark rather than only against itself, and to treat persistent principled disagreement as a signal the category may need to split, not as noise to hammer flat.
How it implements the components¶
The workshop fills the alignment-side components a facilitated ritual can produce:
calibration_case_library— its core output: the set of agreed reference cases every judge now anchors to.context_weighting_rule— the session makes explicit how context (seniority, channel, stakes) reweights typicality, so the same item is judged consistently across situations.prototype_anchor_set— reconciliation ratifies which examples the group treats as the canonical members and near-misses of the category.
It aligns people but does not measure the disagreement statistically or trace its causes — that's a Classification Disagreement Audit — and it builds the reference library rather than monitoring it for staleness, which is Drift Sample Review.
Related¶
- Instantiates: Prototype-Centered Category Modeling — the ritual that turns many private prototypes into one shared one.
- Sibling mechanisms: Classification Disagreement Audit · Drift Sample Review · Boundary Case Review Panel · Golden Case Benchmark · Typicality Rating Exercise · Positive / Negative Example Deck
References¶
[1] Inter-rater reliability — the degree to which independent judges assign the same case the same label, often summarised with a statistic such as Cohen's kappa. A calibration workshop is the standard intervention for raising it when a category is judged by many hands. ↩