Skip to content

Column-Oriented Storage

A physical tabular-data layout that groups values by attribute into separately readable column runs while retaining the associations needed to reconstruct rows.

Version
v1 · 2026-10-03 · History
Domain-specific #
13070
Domain group
Applied Sciences & Engineering
Origin domain
Computer Science & Software Engineering
Subdomains
Database Systems, Analytic Storage → Computer Science & Software Engineering
Aliases
Columnar Storage, Column Store Layout

Core Idea

Column-oriented storage physically groups a table's values by attribute into separately readable runs or chunks instead of primarily placing every field of one row together. It still preserves a way to match separated field values back to their logical rows. Thus the identity is a physical layout, not merely a table that has named columns.[ref-dc569408e775][parquet]

Reading only a few fields across many records can avoid unrelated column runs; reconstructing whole records or updating them may require extra coordination. Encoding and page statistics can improve a particular implementation, but they are not required for the layout to be column-oriented.[ref-dc569408e775][ref-ff08f0a87be1]

Scope of Application

C-Store is a database-engine example: its read-store projection segments are split into columns, while storage keys and join indices allow logical rows to be reconstructed. It also has an updatable write-store component, so column orientation does not mean updates are impossible.[^ref-dc569408e775]

Apache Parquet is a file-format example. It divides rows into row groups, stores one contiguous column chunk per field within each group, and subdivides chunks into pages. The column chunk is the I/O unit; encoding and optional page indexes are further design choices. An in-memory Arrow record batch can also use column arrays, but that specific format is not the definition of all column-oriented storage.[parquet][ref-ff08f0a87be1][^ref-c170b8c36872]

Clarity

The decisive question is where the values sit physically and how separated fields are matched into rows. A logical SQL column, an optional compression codec, or a min/max index does not by itself prove column-oriented layout. Nor does a wide-column system's “column family” label: Bigtable uses families for sparse keys, access control and compression, which is a different data-model commitment unless the relevant physical column-run layout is independently shown.[ref-dc569408e775][ref-18a68a69d029]

Manages Complexity

The abstraction reduces a storage comparison to three checks: which values are grouped, which units can be read separately, and how records are reassembled. These checks explain a likely narrow-scan advantage and a possible full-row or mutation cost without pretending the layout is faster for every query. Compression, sorting, statistics and row-group size modify performance but should be assessed separately.[ref-dc569408e775][parquet]

Abstract Reasoning

Given a proposed format or engine, identify its logical records and attributes, locate its physical column runs or chunks, and verify the row-association mechanism. Then compare that layout with the workload: a query scanning two fields of a wide table may benefit from reading two runs; a workload repeatedly fetching or changing complete rows may value row locality instead.[ref-dc569408e775][parquet]

Parquet shows a second tradeoff within the layout. Larger row groups permit larger sequential I/O but require more write buffering; smaller pages permit finer-grained reads with additional header/processing cost. The appropriate unit size follows from the access profile, not from the label “columnar.”[^ref-ab26193561cd]

Knowledge Transfer

The physical attribute-grouping pattern transfers literally from C-Store's database projections to Parquet's persisted row groups, although their mechanisms for row correspondence differ. It is a strict proposed instance of live Data Structure: arranging information to favor some operations at the cost of others. The staged Parallel Array pattern is narrower because it requires same-index field arrays; column-oriented storage also permits keyed projections and grouped chunks.

[^ref-dc569408e775]: Michael Stonebraker and colleagues, “C-Store: A Column-oriented DBMS”, VLDB 2005, original author-hosted paper, §§1–4. [^parquet]: Apache Parquet project, “Concepts”, official file-format documentation. [^ref-ff08f0a87be1]: Apache Parquet project, “Page Index”, official file-format documentation. [^ref-ab26193561cd]: Apache Parquet project, “Configurations”, official file-format documentation. [^ref-18a68a69d029]: Fay Chang and colleagues, “Bigtable: A Distributed Storage System for Structured Data”, OSDI 2006, original paper, §2. [^ref-c170b8c36872]: Apache Arrow project, “Arrow Columnar Format”, official specification.

Relationships to Other Abstractions

Local relationship map for Column-Oriented StorageParents appear above the current abstraction, mutual partners to the right, and children below. Node labels state whether each abstraction is prime or domain-specific; colors identify relation types.Column-OrientedStorageDOMAINPrime abstraction: Data Structure — is a kind ofData StructurePRIME

Current abstraction Column-Oriented Storage Domain-specific

Parents (1) — more general patterns this builds on

  • Column-Oriented Storage is a kind of Data Structure Prime

    Column-oriented storage is a data arrangement whose attribute-wise layout changes the cost of projected scans and whole-row access.

Hierarchy path (1) — routes to 1 parentless root

Neighborhood in Abstraction Space

Column-Oriented Storage sits in a sparse region of the domain-specific corpus (77th percentile for distinctiveness): few abstractions share its structure, so a faithful description tends to retrieve it precisely.

Family — Digital Resource Formats & Metadata (7 abstractions)

Nearest neighbors

Computed from structural-signature embeddings · 2026-10-08