Skip to content

Sequencing Coverage

How fully and evenly usable sequence reads represent a declared genomic target, expressed through per-position depth and aggregate depth or breadth summaries.

Version
v2 · 2026-10-03 · History
Domain-specific #
13603
Domain group
Natural Sciences
Origin domain
Biology & Ecology
Subdomains
Genomics, Sequencing → Biology & Ecology
Aliases
Sequence coverage, Read-based coverage

Core Idea

Sequencing coverage describes how a declared genomic target is represented by usable sequence reads. At a position, its depth records how much qualifying read evidence overlaps that position. Across a region, coverage is summarized by a mean or median, a distribution of depths, or the fraction of bases at or above a declared threshold—often called breadth at that threshold. These are linked views of a positional profile, not interchangeable claims that every base has the reported average depth.[1][2]

The read and base eligibility rules matter. Mapping quality, base quality and other filtering choices change which observations count. The elementary estimate N × L / G compares read count N times average read length L with target length G; it is a useful nominal mean under simplifying assumptions, not a promise that all output maps evenly or remains after filters. The source article's broad word “coverage” also includes physical span and genomic breadth; this entry uses the narrower read-based sequencing identity and names those differences explicitly.[1][2][3]

Structural Signature

Sig role-phrases: declared target → counted aligned evidence → local depth function → optional aggregation and threshold → task-specific interpretation.

  • Declared target: A genome, exome or named set of intervals supplies the positions and denominator. Without it, a read count alone does not establish the coverage of anything.[2]
  • Counted aligned evidence: Reads or read bases assigned to target positions are included under stated alignment and quality rules. Different filters can yield different effective depth from the same raw run.[2]
  • Local depth function: Each target position has its own representation by qualifying reads. This positional structure reveals gaps and thin areas that an output total cannot.[1]
  • Optional aggregation and threshold: A local-depth statement needs no regional summary. When the question concerns a region, a mean, median, histogram, or percentage of bases beyond a stated depth compresses the positional function. A threshold, if used, is selected for the application, not an inherent universal cutoff.[1][2]
  • Task-specific interpretation: The profile is compared with the intended sequencing question. More depth can support rare-event investigation, but coverage by itself neither proves an accurate variant call nor guarantees every region is adequate.[1]

What It Is Not

It is not sequencing throughput. Producing a large number of reads says how much data was generated, not how those reads map to a target. Nor is mean depth a guarantee that every position has that depth: Illumina's coverage histograms explicitly show uneven distributions, and a low-depth tail can coexist with a high mean.[1]

It is also not automatically molecular uniqueness. A depth count depends on filtering and duplicate conventions; the seed's phrase “number of independent reads” would overclaim if repeated evidence from the same molecule were counted. Physical coverage may count regions spanned between paired reads even when the intervening bases were not sequenced; it should not be silently substituted for per-base read depth. Breadth asks what fraction of the target meets a defined depth, which is an aggregate of a different form, not a synonym for the depth at one base.[2][3]

Scope of Application

Whole-genome studies, exomes and targeted intervals can all carry the same literal positional measure: identify a target, align usable reads, then report local depths and appropriate summaries. GATK documents per-locus, per-interval and per-gene summaries, as well as mean, median and fractions above chosen thresholds. The identity remains the same when the target size or assay changes; the requirements for using the resulting profile may differ.[2]

The metric does not by itself settle whether a genotype, assembly or low-frequency variant is correct. Those conclusions additionally depend on base errors, alignment ambiguity, assay design and analysis decisions. Illumina therefore discusses desired coverage by method and application rather than supplying one universal number. A coverage claim is incomplete when its target, eligibility criteria or summarized statistic is left ambiguous.[1]

Clarity

Coverage reports often compress a whole region into “30×.” That can mean a nominal output-to-genome ratio, a mapped mean, or a filtered mean. The profile view asks which one is meant and whether poorly covered positions are hidden by the average. It separates depth at a base, mean depth across bases and breadth above a threshold. These distinctions resolve apparent disagreements between reports that used different denominators or filters.[1][2]

It also makes the title's ambiguity visible: the candidate Wikipedia page groups read depth, physical span coverage and percentage of bases covered. They are related but answer different questions. A precise report names the read-based quantity and its aggregation instead of treating “coverage” as a single self-explanatory number.[3]

Manages Complexity

Millions of read-to-position overlaps can be reduced to a small, interpretable set: the target intervals, eligibility rules, positional depth distribution and one or more task-relevant summaries. A histogram or quartiles preserve information about unevenness that a single mean loses. A breadth-at-threshold statistic answers whether enough target positions clear a chosen floor without listing every locus.[1][2]

This compression has a known cost. A region-level summary can still conceal a particular clinically or scientifically important position. The analyst returns to local depth when the decision concerns a specific locus. Thus coverage provides a ladder of views—local, interval and aggregate—rather than replacing the underlying evidence with one definitive figure.[2]

Abstract Reasoning

From a high nominal N × L / G, one may infer that much sequence was produced relative to target length, not that every base has usable support. Check mapped and filtered depths, then inspect their distribution and the share of target positions meeting the intended floor. If the mean is high but breadth is low, additional data alone may not fix uneven assignment; the target design or mapping behavior needs examination. This is a diagnostic inference from profile shape, not a universal laboratory prescription.[1][2]

If two assays have the same mean but different low-depth tails, they offer different evidence coverage for a task requiring a minimum at each site. Conversely, a lower mean with more uniform representation may cover more target positions above a modest threshold. The comparison is meaningful only when target and read-counting rules are aligned.[1][2]

Knowledge Transfer

Literal use transfers within genomics from genome-wide sequencing to exomes and named panels: the carriers remain genomic positions and aligned reads, although the target, filters and interpretation change. The positional profile can be summarized at a base, gene or interval without changing what “read representation” means.[2]

Other fields also speak of test coverage or geographical coverage. Those may share a broad idea of representation over a target, but they are analogies or instances of a broader, separately argued abstraction. They do not make this named sequencing measure a prime abstraction. No cross-domain parent is claimed on the strength of the word alone.

Examples

Historical whole-genome setting. NHGRI described a Human Genome Project working draft with an average fold coverage. That aggregate describes repeated representation across a genome; it does not assert identical depth at every base. Mapped back: declared target = designated genomic regions; counted aligned evidence = qualifying sequence assigned to those regions; local depth function = representation at each position; optional aggregation and threshold = reported mean fold coverage, without an implied threshold; task-specific interpretation = characterization of a working draft, not certification of every locus.[4][1]

Targeted interval assessment. GATK's documented coverage analysis can summarize a specified panel of intervals by median depth and percentage of bases at or above a selected threshold. This is an illustration of the tool's documented capability, not a claim that a particular sample met a threshold. Mapped back: declared target = the panel intervals; counted aligned evidence = reads and bases after stated mapping/base-quality filters; local depth function = counts at each target locus; optional aggregation and threshold = median and fraction clearing the chosen floor; task-specific interpretation = whether the intervals have the representation relevant to the analysis.[2]

Structural Tensions

Compact mean versus positional evenness. A mean is easy to report, but can hide low-depth target bases; a full distribution is more informative but less compact. Leaning entirely on the mean may conceal gaps, while listing all positions can obscure the overall picture. Diagnostic: What fraction of bases reaches the task-relevant floor?[1][2]

Inclusive output versus quality-filtered evidence. Counting every generated read raises a nominal total, while filtering suspect mappings or bases can lower usable depth. The first may exaggerate evidence; the second must state its rules so comparisons remain fair. Diagnostic: What exactly was eligible to be counted?[2]

More output versus remaining weak regions. More reads can improve coverage for some questions, but an uneven profile may persist. Extra sequencing has a cost and is not itself a guarantee of a trustworthy call. Diagnostic: Are weak loci short of read output, or is representation systematically uneven?[1]

Structural–Framed Character

The entry is predominantly structural: once a target and eligibility rules are fixed, read-to-position overlap yields a depth profile and numerical summaries. Those values are not conferred by an institution's judgment. They are nevertheless framed by choice of reference, filtering, target intervals and depth threshold. Two analysts can report different valid coverage summaries of the same raw run if those choices differ.[2]

Its evaluative weight is conditional: a high value is not automatically good, and a low value is not automatically bad, without a task. Its human-practice dependence lies in study design, alignment and reporting conventions, not in the arithmetic of counts. Its institutional origin is genomic sequencing practice, while the quantity itself is not an institutional status. Its vocabulary travels among genomics assays literally, but “coverage” in unrelated fields imports a broad metaphor rather than recognizing this read-based metric there. Its character: a structurally computable genomic evidence profile whose interpretation depends on declared measurement and research frames.[1][2]

Structural Core vs. Domain Accent

The portable skeleton is representation of a declared target by observed evidence, summarized over its parts. That resemblance to other coverage metrics is a possible future-prime question, not a verified live parent or proof that the named sequencing identity is itself cross-domain. The domain accent is decisive here: aligned sequence reads, genomic positions, quality filters and per-base depth. Physical span coverage and raw throughput can resemble portions of the skeleton without instantiating this exact read-based measure.[1][2][3]

Why not prime: the supported literal applications remain genomics assays. A general representation-over-target abstraction would need its own definition, boundary and multi-domain evidence; broadening this entry's name would erase the distinctions its evidence supports.

No strict typed parent relation is asserted in the current DAG. The entry measures read representation over genomic positions. Shotgun Sequencing is a production method rather than a necessary genus of this metric.

Neighborhood in Abstraction Space

Sequencing Coverage sits in a sparse region of the domain-specific corpus (68th percentile for distinctiveness): few abstractions share its structure, so a faithful description tends to retrieve it precisely.

Family — Molecular Biology & Genetic Engineering Methods (13 abstractions)

Nearest neighbors

Computed from structural-signature embeddings · 2026-10-08

Not to Be Confused With

  • Sequencing throughput: output volume before target mapping and positional aggregation.[1]
  • Physical coverage: a paired-fragment span can include unsequenced interior bases.[3]
  • Breadth at a threshold: a fraction of target positions meeting a selected depth, not the depth at one position.[2]
  • Variant-call accuracy: coverage can contribute evidence but cannot alone establish correctness.[1]

References

[1] Illumina, “Sequencing Coverage for NGS Experiments”, coverage estimation, histograms, uniformity and application-specific recommendations. registry ↩a ↩b ↩c ↩d ↩e ↩f ↩g ↩h ↩i ↩j ↩k ↩l ↩m ↩n ↩o ↩p ↩q ↩r

[2] GATK Team, “DepthOfCoverage (BETA)”, Overview and summary outputs. Cited as analysis-tool documentation, not a universal clinical standard. registry ↩a ↩b ↩c ↩d ↩e ↩f ↩g ↩h ↩i ↩j ↩k ↩l ↩m ↩n ↩o ↩p ↩q ↩r ↩s ↩t

[3] “Coverage (genetics)”, lead and Physical Coverage sections. Used only for the frozen candidate title's terminology and scope boundary. registry ↩a ↩b ↩c ↩d ↩e

[4] National Human Genome Research Institute, “Human Genome Project and SNP Consortium Announce Collaboration to Identify New Genetic Markers for Disease”, working-draft and depth-of-coverage definitions. registry ↩