Reporting-Pyramid Undercount¶
Core Idea¶
Reporting-Pyramid Undercount is the structural pattern in which a true population of countable events must pass through several ordered detection, recognition, classification, escalation, and recording stages before it appears in an official dataset. Each stage admits only a fraction of the events that reached it. The final capture fraction is therefore the product of the layerwise fractions, and the recorded count can be a small remainder even when no single layer looks catastrophically lossy.
The loss is normally non-random. Severe, visible, well-documented, easy-to-classify, low-cost, or institutionally safe events cross more gates than mild, ambiguous, stigmatized, inaccessible, or costly ones. The recorded set is consequently not a uniformly shrunken image of the true set. Its composition is systematically displaced toward the kinds of events that survive the pipeline.
The prime has a second load-bearing commitment: the gates can move. Awareness campaigns, new reporting tools, changes in definitions, expanded testing, exhausted staff, political pressure, and altered incentives all change layerwise capture. A rise in the recorded count can therefore be caused by more true events, by a more permeable reporting pipeline, or by both. A fall can indicate improvement or reporting decay. The observed series is jointly generated by the event process and the observation process.
Structural Signature¶
Sig role-phrases:
- the true event population — the events that meet the target definition whether or not any reporting system sees them
- the ordered inclusion gates — sequential stages an event must cross before entering the dataset
- the layerwise capture fractions — the conditional probabilities of surviving each gate
- the multiplicative attenuation — the product of the layer fractions that determines aggregate capture
- the non-random loss — systematic variation in capture by event severity, visibility, population, place, access, or incentive
- the recorded remainder — the biased subset that appears in the official count
- the moving-filter confound — changes in capture that imitate changes in the underlying incidence
- the external calibration apparatus — audits, population surveys, linked records, capture-recapture studies, or sentinel systems used to estimate what the routine pipeline misses
The strict signature requires both an ordered pipeline and differential inclusion. A single independent sampling gate is ordinary Selection Bias. A sequence of purely random gates produces undercount but not the systematic compositional distortion at the center of this prime.
What It Is Not¶
It is not a generic claim that every official number is wrong. A reporting system can have high or stable capture, and its count can be fit for a bounded purpose. The prime identifies a mechanism that must be demonstrated or plausibly specified layer by layer.
It is not reporting delay. Delayed events are expected to enter the dataset later and can often be handled by nowcasting or censoring corrections. Reporting-pyramid losses are events that never cross one or more gates and therefore never arrive.
It is not simple missing data. A dataset can have blank fields without losing events, and events can be absent without passing through a sequential reporting chain. The pyramid adds ordered conditional gates and a product structure.
It is not merely measurement error. Measurements may be noisy or miscalibrated for every recorded event while the event count remains complete. Reporting-Pyramid Undercount changes which events become visible and usually changes their composition.
It is not capture-recapture. Capture-recapture is one estimator that can help recover a hidden population from overlapping imperfect lists. It is an instrument for some undercount problems, not the filtering pattern itself.
Broad Use¶
Epidemiology and public health. A true infection may have to produce symptoms, cause care-seeking, prompt testing, satisfy a case definition, and reach a surveillance registry. The familiar surveillance pyramid is one concrete domain instance.
Pharmacovigilance. An adverse drug reaction must be experienced, recognized as potentially drug-related, communicated, reported through an approved channel, accepted as a valid case, and entered into a database. Severe and unusual reactions have different capture probabilities from common or ambiguous ones.
Occupational safety. An injury may need worker disclosure, supervisor recognition, recordability classification, and submission to a regulator. Incentives around safety performance, compensation, employment security, and administrative burden can change several gates.
Crime, abuse, and discrimination statistics. An experienced event may need to be recognized, disclosed, accepted by an authority, classified under a legal category, and entered into a record. Social risk and institutional trust make the missing population compositionally unlike the observed one.
Cybersecurity and operations. A failure or intrusion must be instrumented, detected, triaged, escalated, declared, and logged. Changes in telemetry or incident policy can create apparent increases without any change in the underlying event rate.
Public administration. A person entitled to a benefit, remedy, or service may need to learn of eligibility, reach an office, complete documentation, pass an administrative definition, and remain in the process until entry. Administrative counts then measure successful traversal as much as underlying need.
Clarity¶
The prime separates four quantities that a raw count silently fuses:
- the number of true events;
- the fraction crossing each gate;
- the kinds of events preferentially lost;
- changes in the gates themselves.
This is more precise than saying “the data undercount reality.” It asks where the loss occurs, whether the losses compound, which attributes predict survival, and whether the observation system changed during the period being compared. Once those questions are explicit, a count is no longer mistaken for a transparent window onto incidence.
The distinction between level and composition is especially important. Multiplying a recorded count by an estimated inverse capture rate may recover a plausible total while leaving the distribution of severity, geography, demographics, or type badly distorted. Scaling a biased sample does not unbias its composition.
Manages Complexity¶
A reporting ecosystem can contain dozens of institutions, forms, incentives, thresholds, and handoffs. The pyramid compresses that machinery into a short ordered chain of conditional gates. For \(k\) stages, the analyst estimates or bounds \(k\) capture fractions rather than inventing an undifferentiated “dark figure.”
The decomposition also gives a routing rule for intervention. If most loss occurs before recognition, a better submission portal will not recover it. If reports are made but rejected during classification, awareness campaigns target the wrong layer. The prime directs effort toward the binding coefficient while predicting which apparent count changes the intervention itself will create.
Abstract Reasoning¶
Let \(N_t\) be the true number of target events during period \(t\). Let \(p_{it}(x)\) be the conditional probability that an event with attributes \(x\), having survived gates \(1,\ldots,i-1\), crosses gate \(i\). The expected recorded count is
If capture is treated as homogeneous, this reduces to \(E[R_t]=N_t\prod_i p_{it}\). The reciprocal \(1/\prod_i p_{it}\) is an aggregate multiplier, but it is valid for recovering the total only under assumptions about the population and the estimation of the component fractions.
Two consequences follow. First, moderate losses compound: five gates each capturing 70 percent leave only about 17 percent of true events in the final count. Second, \(R_t\) is not identified with \(N_t\) when the \(p_{it}\) change. The same recorded rise can result from increased incidence or increased permeability at any gate. Auxiliary measures are required to separate the two.
The model also exposes irreversibility. A downstream reform cannot recover events already lost upstream unless it creates a new route around the earlier gate. Better classification cannot count events never detected; mandatory submission cannot count events never recognized.
Knowledge Transfer¶
The abstraction transfers exactly across systems with an in-principle countable event class and a sequential inclusion path. A drug reaction, workplace injury, assault, software incident, and infection have different substantive content, yet each can be mapped to the same roles: true events, recognition, disclosure, institutional acceptance, classification, and entry.
Transfer fails when the “pyramid” is only a picture of decreasing quantities with no conditional traversal by the same event units. A food chain, organizational hierarchy, or funnel of unrelated populations may have a pyramid shape without instantiating Reporting-Pyramid Undercount. What travels is the event-level passage through non-random gates, not the geometry.
Examples¶
Formal/abstract¶
Suppose 10,000 events occur. Eighty percent are detectable, half of those are recognized, 60 percent are submitted, and 75 percent of submissions are accepted and entered. The final count is \(10{,}000 \times .8 \times .5 \times .6 \times .75 = 1{,}800\). No layer alone loses more than half, but 82 percent of the events disappear through compounding.
If a new portal raises submission from .6 to .9 while incidence stays fixed, the recorded count rises to 2,700—a 50 percent apparent increase caused entirely by one observation-layer change.
Mapped back: the 10,000 are the true event population; the four conditional rates are the layerwise capture fractions; 1,800 is the recorded remainder; and the portal creates the moving-filter confound.
Applied/in practice¶
A hospital safety office launches anonymous mobile incident reporting. Recorded medication errors rise sharply the next quarter, especially for near misses. It would be mistaken to infer that care suddenly became less safe. The new channel lowered the disclosure cost and raised capture for events that staff previously judged too minor or risky to report.
Mapped back: medication errors are the true event population; recognition, willingness to disclose, classification, and database entry are the ordered inclusion gates; fear and perceived severity create non-random loss; and the mobile tool changes a gate while leaving the event process potentially unchanged.
Structural Tensions¶
T1: Incidence versus ascertainment. A reported trend is the product of events and capture. Increasing detection can look like worsening incidence; collapsing reporting can look like improvement. Diagnostic: Which independent indicators can distinguish a change in true events from a change in one or more gates?
T2: Total correction versus compositional correction. An aggregate multiplier may recover a plausible total while preserving severity, demographic, geographic, or access bias. Diagnostic: Is the inference only about how many events occurred, or does it also claim who or what kinds were affected?
T3: Parsimony versus localization. A single multiplier is easy to communicate, but an intervention needs layer-specific coefficients. Diagnostic: Does the proposed repair act on the layer that accounts for the loss, or merely on the most visible downstream stage?
T4: Stable comparison versus changing infrastructure. Better reporting is desirable, yet it breaks comparability with the older series. Diagnostic: Was the apparent change contemporaneous with a new definition, tool, incentive, capacity, or disclosure environment?
T5: Upstream loss versus downstream control. The largest losses may occur before the institution has any contact with the event. Diagnostic: Can a downstream mandate reach events lost before detection or disclosure, or is a new upstream sampling route needed?
T6: Estimation versus governance. Measuring the hidden population may require linking sensitive records or intensifying surveillance, creating privacy, trust, and incentive costs that themselves alter reporting. Diagnostic: Will the calibration apparatus change the gates it is intended to estimate?
Structural–Framed Character¶
Reporting-Pyramid Undercount is structurally portable but institutionally realized. Its mathematical core—sequential conditional inclusion with multiplicative attenuation—does not depend on a particular profession or value judgment. Its common vocabulary of reporting, cases, eligibility, and official counts does presume an observer or institution maintaining a record.
The term “undercount” evaluates the recorded count against a declared target population. That target definition is framed: institutions decide what should count. Once the target is fixed, however, whether events cross the gates and how the probabilities multiply are structural facts. The prime therefore combines a framed boundary with a portable mechanism.
Substrate Independence¶
The specific layers can be removed without breaking the abstraction. Strip symptom expression and testing from disease surveillance; what remains is a true event set crossing ordered, non-random gates. Strip supervisor recording from workplace safety; the same roles remain. The mechanism survives changes in people, organizations, software, and legal regimes.
Its limits are equally clear. It does not generalize to any loss of information, any funnel, or any difference between reality and data. It requires event units, a target count, sequential conditional gates, and a final recorded subset. Those constraints give it more substance than the generic metaphor of an iceberg.
Relationships to Other Abstractions¶
Current abstraction Reporting-Pyramid Undercount Prime
Parents (1) — more general patterns this builds on
-
Reporting-Pyramid Undercount is a kind of Selection Bias Prime
Reporting-Pyramid Undercount is Selection Bias specialized to an ordered sequence of non-random inclusion gates whose capture fractions multiply.It retains Selection Bias's defining mismatch between the target population and the systematically selected observed sample, then fixes the selection mechanism to a sequence of reporting or ascertainment stages. Each stage admits a non-random fraction of the events surviving the previous stage, so both the total and composition of the final count can diverge from the true event population.
Children (1) — more specific cases that build on this
-
Outbreak Underascertainment Domain-specific is a kind of Reporting-Pyramid Undercount
Outbreak Underascertainment is Reporting-Pyramid Undercount specialized to infectious-disease surveillance and its symptom, care-seeking, testing, confirmation, and notification layers.It retains the ordered non-random gates, multiplicative capture fraction, biased recorded subset, and moving-filter confound, then fixes the event population and calibration apparatus to outbreaks, seroprevalence, multiplier studies, and disease-surveillance practice.
Hierarchy paths (6) — routes to 6 parentless roots
- Reporting-Pyramid Undercount → Selection Bias → Bias
- Reporting-Pyramid Undercount → Selection Bias → Statistical Inference → Inductive Reasoning
- Reporting-Pyramid Undercount → Selection Bias → Statistical Inference → Uncertainty
- Reporting-Pyramid Undercount → Selection Bias → Vantage-Induced Omission → Viewpoint
- Reporting-Pyramid Undercount → Selection Bias → Statistical Inference → Probability → Measure → Set and Membership
- Reporting-Pyramid Undercount → Selection Bias → Statistical Inference → Probability → Measure → Aggregation → Micro Macro Linkage
Neighborhood in Abstraction Space¶
Reporting-Pyramid Undercount has no computed distinctiveness yet.
Family — Unclustered & Miscellaneous (429 primes)
Nearest neighbors
Computed from structural-signature embeddings · 2026-07-26
Distinction from Neighbors¶
Selection Bias is the strict parent. It covers any systematic difference between a target population and the units selected into observation. Reporting-Pyramid Undercount adds ordered sequential gates, multiplicative capture, and a layer-localized intervention model.
Missing Data concerns absent observations or fields and need not describe how event units traverse a reporting system. Attrition usually begins from an enrolled or once-observed population and tracks loss over time; reporting-pyramid events can disappear before the first record. Reporting Delay concerns eventual arrival and time censoring rather than permanent exclusion. Measurement Error changes values assigned to recorded units. Misclassification places units in wrong categories; it can operate at one pyramid gate but is not the whole pipeline. Capture-Recapture is an estimation method using overlapping imperfect lists, not the undercount mechanism.
Streetlight Effect reallocates search toward accessible regions. It can help create an undercount when institutions look only where detection is easy, but a reporting pyramid can exist under complete search effort because events still fail later gates.
Solution Archetypes¶
No catalogued solution archetypes reference this prime yet.
Notes¶
This entry was drafted during mixed-DAG puzzle pass Recursion 100 because the domain-specific Outbreak Underascertainment entry repeatedly named the cross-domain reporting-pyramid mechanism as its missing parent. It is intentionally compact and requires Claude style re-authoring, density expansion, FACT anchors, and independent source verification before publication.
The node should remain below Selection Bias rather than receive parallel flattened heads to Missing Data, Pipeline, or Measurement Error. Those concepts are neighbors or possible mechanisms. The exact genus is systematic selection into the recorded set; the ordered product of reporting gates is the differentiator.
References¶
- Gibbons, C. L., et al. “Measuring Underreporting and Under-Ascertainment in Infectious Disease Datasets: A Comparison of Methods.” BMC Public Health 14 (2014): 147. Citation lead; independently verify bibliographic details.
- Havers, F. P., et al. “Seroprevalence of Antibodies to SARS-CoV-2 in 10 Sites in the United States, March 23–May 12, 2020.” JAMA Internal Medicine 180, no. 12 (2020): 1576–1586. Citation lead; independently verify.
- Scallan, E., et al. “Foodborne Illness Acquired in the United States—Major Pathogens.” Emerging Infectious Diseases 17, no. 1 (2011): 7–15. Citation lead; independently verify.
- Hazell, L., and S. A. W. Shakir. “Under-Reporting of Adverse Drug Reactions: A Systematic Review.” Drug Safety 29, no. 5 (2006): 385–396. Citation lead; independently verify.