Skip to content

Taxonomy Merge Workshop

Facilitated ritual — instantiates Equivalence Class Consolidation

A facilitated session where domain experts decide whether whole categories should be grouped, split, or treated as near-equivalent, and who will own the result.

A Taxonomy Merge Workshop consolidates at the level of categories, through structured human deliberation. When the question is not "do these two records match?" but "should these two whole categories be one — or should this overbroad one be split?", the judgment is contested, evidence-laden, and political, and no algorithm can settle it. This mechanism is the facilitated gathering that does: it puts a bounded set of categories on the table, brings the people with domain authority to argue grouping versus splitting to a decision, assigns each resulting class an owner, and records the outcome with a route to revisit. What makes it this mechanism and not the lighter Synonym Merge Review is scale and stakes — it restructures a classification and its ownership, not just the labels on a concept.

Example

A working group of systematists convenes to resolve the status of two long-separately-named beetle "species" that new genetic and morphological evidence suggests are one. This is a merge decision over categories. The group scopes the session to a specific set of related taxa rather than the whole family, weighs the evidence in the room, and concludes the two names denote a single species. By the Principle of Priority, the older published name becomes the valid one and the newer is demoted to a junior synonym.[1] A curator is named as the authority responsible for the accepted taxon in the reference database, and the decision — with its evidence and any dissent — is recorded so that a future finding can reopen it. The same session might go the other way on a different taxon, splitting one broad "species" into two when the group judges the differences to be real. The tension between merging and splitting is the whole substance of the room.

How it works

Its distinguishing traits are synchronous expert deliberation, bounded scope, and assigned ownership:

  • It fixes which categories are in play for the session, bounding the scope so the discussion converges instead of sprawling into an endless reorganization.
  • It brings domain authority into deliberation — the decision is reached by argued judgment over evidence, not computed — and resolves each case to group, split, or hold.
  • It assigns an owner to each accepted class and records the decision, dissent included, so the outcome is both stewarded and reopenable.

Tuning parameters

  • Participant set — who is in the room and whose judgment carries weight. Broad expertise catches more real distinctions; a narrow room decides faster but risks blind spots.
  • Scope per session — how many categories are put on the table at once. Tight scope produces durable decisions; sprawling scope exhausts the group and yields shallow ones.
  • Decision rule — consensus, majority vote, or designated-authority ruling. Consensus legitimizes but can deadlock; an authority ruling is decisive but contestable.
  • Lumper/splitter default — the group's bias when evidence is ambiguous: merge unless proven distinct, or keep separate unless proven equivalent. This default silently shapes most borderline calls.
  • Bindingness — whether the outcome is provisional or authoritative, and how a later challenge reopens it.

When it helps, and when it misleads

Its strength is settling category boundaries that are genuinely contested and evidence-dependent — exactly the cases automation cannot resolve — while surfacing the tacit domain knowledge that lives only in experts' heads, and producing a decision that is owned and therefore legitimate.

Its failure modes are social. The lumper/splitter bias swings outcomes as much as evidence does, so a group's disposition can over-merge away real distinctions or over-split into fragmentation. Expert politics and the loudest voice can dominate, and a workshop can ratify a grouping for organizational convenience rather than because the categories are truly equivalent. The classic misuse is convening the session to rubber-stamp a reorganization already decided elsewhere. The discipline is to put the equivalence criterion and the evidence explicitly on the table, name owners and record dissent, and keep a revisit path so a decision made under today's knowledge can be corrected under tomorrow's.

How it implements the components

  • comparison_scope — the workshop's first act is to bound which categories are under review for the session, keeping the deliberation tractable.
  • class_owner — it assigns an accountable steward or authority to each accepted class going forward.
  • merge_split_review_path — the session is the deliberative route for grouping, splitting, and revisiting categories, and it records each decision for future challenge.

It does not infer equivalence from record data — that is the Identity Resolution Model — nor curate individual term synonyms case by case (the lighter Synonym Merge Review), nor map the resulting categories across external schemes (Crosswalk Table).

  • Instantiates: Equivalence Class Consolidation — the workshop is the expert deliberation that decides how whole categories group, split, or align.
  • Sibling mechanisms: Synonym Merge Review · Crosswalk Table · Alias Resolution Table · Canonicalization Pipeline · Deduplication Workflow · Equivalence Test Suite · Identity Resolution Model · Master Record Consolidation · Policy Equivalence Rule · Unit Normalization Table

References

[1] Lumpers and splitters names the enduring disposition split among taxonomists toward merging or dividing categories; in zoological nomenclature the Principle of Priority resolves a confirmed merge by making the oldest available name valid and the rest junior synonyms — a real, codified route for the grouping decisions such a workshop makes.