Interlaboratory Comparison¶
A cross-site comparison exercise — instantiates Traceable Measurement System Design
Sends the same or comparable targets to independent labs, sites, or methods and compares their qualified results — separating real site-to-site bias from true differences in the things measured.
A method that works in one lab may not travel. Interlaboratory Comparison tests whether it does: it distributes a common or comparable target to independent sites, methods, or teams, has each measure it under its own real conditions, and then compares the qualified results — each with its stated uncertainty — to see how far apart independent implementations of "the same measurement" actually land. Its defining move is preserving independence while circulating a shared target, so the spread it reveals can be split into site bias (an implementation that reads high or low) versus target heterogeneity (the items really differing). Where a within-lab study asks whether one system repeats, and a reference comparison asks whether one method matches a standard, this mechanism asks the reproducibility question that only many independent operators can answer: does the measurement mean the same thing everywhere it is made?
Example¶
A proficiency scheme for grain-testing labs mails every participant a portion from the same well-mixed lot of maize spiked with a known mycotoxin level. Each lab measures it with its ordinary method, staff, and instruments, and returns a qualified result. The organizer plots the results against the assigned value and computes a normalized deviation for each lab.
Most cluster near the assigned value; two sit conspicuously high, and their deviations — scaled by their own reported uncertainties — exceed the acceptance criterion. Because every lab measured a piece of the same material, that high reading cannot be blamed on a different sample: it is real site bias, most likely a calibration or extraction-recovery problem. The scheme flags those two labs for corrective action and re-testing, and the round doubles as evidence that the rest of the network produces comparable, transferable numbers — the kind of reproducibility evidence a single lab can never generate about itself.
How it works¶
The distinguishing work is comparing independent executions of a shared target without flattening them:
- Circulate a stable, comparable target. The item must be homogeneous and stable enough that differences between labs aren't just the sample changing in transit.
- Specify enough, but not too much. Provide common information without over-standardizing away real local practice — the point is to see how the measurement behaves as actually implemented, not in a scripted ideal.
- Preserve independent execution. Sites measure and report without conferring; results copied to force agreement destroy the signal.
- Separate bias from heterogeneity. Analyze between-site dispersion, normalize deviations by uncertainty, investigate outliers, and document harmonization or the limits of comparability.
Tuning parameters¶
- Target stability and comparability — how identical the circulated items are. Truly identical items isolate site effects cleanly but may not represent the diversity of real samples.
- Degree of standardization — how much common protocol to impose. Over-standardize and you hide the very method spread you convened the comparison to measure.
- Participant breadth — how many and which sites. A broad panel is more representative of the real network but harder to coordinate.
- Scoring statistic — raw deviation, z-score, or an uncertainty-normalized error. Normalizing by each lab's uncertainty is what makes a "big" deviation meaningful rather than merely large.
- Reference versus consensus value — whether "truth" is an assigned reference or the participants' consensus. A consensus can be biased if the whole field shares a systematic error.
When it helps, and when it misleads¶
Its strength is real-world transfer evidence: it exposes site effects that no lab can see in its own data, and it supports harmonization across a network — the logic of proficiency testing[n1] as an external check on reproducibility.
It misleads when the comparison's own weaknesses are ignored. An unstable or heterogeneous target confounds site bias with sample variation; selectively excluding awkward participants flatters the network; hidden method differences masquerade as noise. The sharpest trap is treating the consensus mean as truth — if a systematic error is common to most participants, the consensus is confidently wrong, and an accurate outlier gets penalized. The classic misuse is ranking or punishing labs on raw spread without accounting for their uncertainties, which turns a learning exercise into an unfair league table. The discipline that keeps it honest is to verify target stability, normalize deviations by uncertainty, prefer a traceable reference value to a bare consensus, and use results for method improvement rather than punitive exposure.
How it implements the components¶
Interlaboratory Comparison fills the cross-implementation side of the chain:
repeatability_and_reproducibility_profile— it supplies the reproducibility half of the profile: agreement under changed sites, instruments, and teams, which only a multi-site exercise can measure.measurement_stewardship_and_versioning— its rounds feed cross-site governance: harmonizing methods and versions, onboarding new sites and methods on comparable terms, and tracking corrective actions to closure.
It does not separate within-lab repeatability from operator variance in a controlled study (that's the within-system focus of Gauge Repeatability and Reproducibility Study), anchor a method to a certified reference (calibration_and_traceability_chain, validity_and_selectivity_evidence — Reference Material Comparison), or control human-rater expectancy (operational_definition — Blinded Rater Assessment).
Related¶
- Instantiates: Traceable Measurement System Design — supplies the evidence that a measurement is reproducible and comparable across independent implementations.
- Consumes: Measurement Protocol supplies the common method specification circulated to participating sites.
- Sibling mechanisms: Gauge Repeatability and Reproducibility Study · Reference Material Comparison · Calibration Traceability Record · Instrument Drift Control Chart · Blinded Rater Assessment · Limit of Detection Estimation · Measurement Protocol · Measurement System Validation Study · Uncertainty Budget Table
Editorial Notes¶
Form Classification¶
Form family: Experiment, Test & Rehearsal
Rationale: Interlaboratory Comparison operates as a bounded trial, probe, simulation, or rehearsal that generates evidence from performance because it sends the same or comparable targets to independent labs, sites, or methods and compares their qualified results — separating real site-to-site bias from true differences in the things measured
Independent corroboration: The frozen evidence defines Interlaboratory Comparison as 'Sends the same or comparable targets to independent labs, sites, or methods and compares their qualified results — separating real site-to-site bias from true differences in the things measured', so its operative form is Experiment, Test & Rehearsal.
Nearest alternative: Assessment, Review & Assurance — Independent labs deliberately repeat measurement on a shared target, making comparison a replication trial rather than ordinary review.
Review outcome: Independent reviewer agreement; medium confidence.
Origin Attribution¶
Primary origin: Statistics & Experimental Design
Origin pattern: Convergent development
Present-day reach: Specialized
Rationale: Comparing qualified results across laboratories is a core experimental reproducibility and proficiency-testing practice.
Related originating lineages:
- Chemistry & Materials Science — Analytical chemistry and metrology materially institutionalized interlaboratory comparison around reference materials.
- Engineering & Design — Metrology, calibration traceability, and laboratory quality systems materially define reference targets and acceptance criteria.
- Medicine & Healthcare — External quality assessment in clinical laboratories provides a mature recurring implementation.
Review resolution: Both independent reviews place the primary lineage in statistics_experimental_design. The queued differences (alternate_origin_disagreement) concern secondary metadata rather than primary provenance. The final retains engineering_design, chemistry_materials, medicine_healthcare only where a reviewer supplied a formative-lineage rationale; this does not convert downstream applicability into origin. origin_mode=convergent because the reviewers document independently established or materially co-developing traditions. domain_reach=specialized records application breadth separately from provenance.
Review outcome: Reconciled after independent review; high confidence.
Notes¶
[n1] Proficiency testing — the use of interlaboratory comparisons to determine the performance of participating labs against pre-established criteria (the basis of external quality-assessment schemes). It is the standardized, recurring form of this mechanism, and the normalized-error and z-score statistics it uses are what make a lab's deviation interpretable rather than merely large. ↩