Skip to content

Data Science & Analytics

← Back to Domain-Specific Abstractions by Origin Domain

37 domain-specific abstractions whose origin domain is Data Science & Analytics. They span 40 subdomains — sort by that column to group them, or click any subdomain to filter to it.

AbstractionSubdomainDescription
Access URLReach a resource through a published, dereferenceable handle that names a route rather than the bytes — so location, storage, and hosting stay hidden behind a routing layer and identity can outlive the link.
Adjusted mutual informationBy adopting a hypergeometric model of randomness, it can be shown that the expected mutual information between two random clusterings is.
Alluvial diagramA flow visualization whose blocks represent groups at successive stages and whose width-scaled streams show how members or quantities move among those groups over time or categories.
Annotation DriftThe gradual, undocumented shift over time in how annotators apply an unchanged rubric, so identical instances get different labels — a silent recalibration of a measurement instrument, invisible to within-slice agreement and detectable only by re-annotating frozen gold.
Bitemporal modelingRecord each fact along both valid time—when it holds in the modeled world—and transaction or system time—when the database records it—so corrections preserve what was believed earlier while enabling as-of knowledge and as-of reality queries.
Consistent Overhead Byte StuffingA reversible byte code that removes a reserved delimiter from packet bodies while bounding worst-case expansion to roughly one byte per 254 input bytes.
Covariate-Shift Blind SpotThe deployment failure in which a model's input distribution P(X) drifts outside its training support while P(Y|X) holds — and goes undetected because monitoring watches lagging outcome metrics instead of the immediately-available input signal.
Data and information visualizationThe design of visual encodings that transform data or information into spatial marks, channels and interactions for exploration, explanation and decision support.
Data binningA preprocessing transformation that groups values into intervals or categories and replaces or summarizes observations by their bin membership or representative value.
Data cleansingThe governed detection, diagnosis and correction, standardization, quarantine or removal of data defects so records satisfy declared quality rules while preserving provenance and uncertainty.
Data LineageRecording, for every data element, the complete sequence of sources, transformations, and movements that produced it — as edges in a queryable dependency graph — so audit questions become forward and backward traversals rather than forensic reconstruction.
Data Reporting—A governed workflow that collects scoped observations, maps them into a required representation, validates them, and submits them to an identified consumer before interpretation.
Distribution FormCatalog one logical dataset as many typed delivery artifacts — CSV, JSON, Parquet, Shapefile — each declaring its own media type, size, and checksum while identity and descriptive metadata stay fixed at the dataset.
Distributional Blind SpotThe region of a model's input space inadequately sampled during development into which the deployed model still makes confident predictions — extrapolations whose error is unknown, indistinguishable in confidence from in-distribution outputs.
Funnel ChartA chart that encodes quantities associated with successive process stages as aligned widths or areas, typically narrowing to reveal attrition, conversion, or remaining volume from one stage to the next.
Ground-Truth DriftThe model-evaluation failure in which the operational definition of the correct answer drifts on a clock the monitoring apparatus cannot see, so metrics keep scoring against a moved target while the dashboards stay green — invisible because every detector consumes current ground truth as its reference.
Horizon chartA compact quantitative graphic that folds value bands onto a shared baseline and uses color intensity and sign to preserve magnitude patterns in little vertical space.
Imputation LeakageThe model-evaluation failure in which a missing-value repair step is fit across the train/test boundary, so its parameters encode facts about the held-out rows — inflating performance that survives into the test metric, because imputation, mentally filed as data cleaning, is really a model.
Label AmbiguityDiagnose a headline accuracy figure as a blend of two measurements — model capability in the class interior where annotators agree, and mere adjudication agreement in the boundary zone where reasonable experts split — by stratifying metrics on the inter-annotator agreement rate.
Label noiseIncorrect, inconsistent, ambiguous, or corrupted target labels in supervised-learning data, arising randomly or systematically from annotators, processes, proxies, attacks, or changing definitions.
Label ShiftThe distribution shift in which the label marginal P(Y) changes between training and deployment while P(X|Y) stays fixed, so a classifier's discrimination survives but its calibration and thresholds miscalibrate — correctable by re-estimating the deployment prior rather than retraining.
Lagrangian–Eulerian advectionA flow-visualization technique that combines particle-following motion with grid-based texture updating to depict unsteady velocity fields coherently.
Local maximum intensity projectionRender volumetric data by tracing each viewing ray and selecting the first threshold-qualified local intensity maximum, preserving depth order that global maximum projection discards.
Location IntelligenceIn business intelligence, location intelligence (LI), or spatial intelligence, is the process of deriving meaningful insight from geospatial data relationships to solve a particular problem.
Log–log plotPlot positive x and y values on logarithmic axes so multiplicative ratios become equal distances and a power law y=ax^k becomes a straight line with slope k and intercept log a.
Motion chartA motion chart dynamically maps multivariate longitudinal data to position, size, color, glyph, and time for interactive exploration.
Parallel coordinatesRepresent each multivariate record as a polyline crossing one parallel axis per variable at its scaled coordinate, making high-dimensional profiles visible while exposing axis-order, scaling, and overplotting choices.
RegressionThe statistical method of modelling an outcome as a systematic function of explanatory variables plus specified noise, fit by minimising a loss — supporting three distinct uses (prediction, effect estimation, variance attribution) each gated by its own validity conditions.
Sankey diagramRepresent transfers through a directed node-link diagram whose band widths are proportional to declared extensive quantities, making dominant paths, splits, mergers, and accounted losses visually comparable.
Scatter plotA graph representing paired observations as points positioned by two quantitative variables.
Self-Similarity MatrixIn data analysis, the self-similarity matrix is a graphical representation of similar sequences in a data series.
Simulated fluorescence process algorithmA volume-rendering algorithm that models fluorescence excitation, emission, absorption and scattering to produce physically interpretable images of three-dimensional data.
Slope OneA family of item-based collaborative-filtering algorithms that predicts a user's rating from average pairwise rating differences between items and the user's ratings of neighboring items.
Spatial coverageA metadata declaration publishing the geographic region within which a resource is valid — an explicit inclusion-exclusion rule on the spatial dimension that lets the catalog check, before any analysis, whether a question's scope falls inside or outside it.
Temporal coverageA metadata declaration that a resource is valid only within a stated time interval, published as an inclusion-exclusion rule on the time dimension so consumers can check at the catalog layer whether their question's period falls inside the window — and whether it has gone stale.
Text miningText mining, text data mining (TDM) or text analytics is the process of deriving high-quality information from text.
Waterfall ChartA floating-bar visualization that reconciles an opening value to a closing value by encoding each ordered signed contribution as the step between consecutive running totals.