Working set size¶
Working-set size is the memory footprint of the input and intermediate data required to solve a particular problem instance.
Core Idea¶
Working-set size (WSS) is the amount of memory occupied by the input and intermediate data that a program or algorithm needs while solving a particular problem instance.[1] The working set is the data collection; its size is the corresponding memory demand.[2]
The quantity is instance- and execution-dependent.[3] The same algorithm can have different working-set sizes for differently sized inputs or for implementations that retain, stream, recompute, or discard intermediate results in different ways.[4] It is therefore not simply executable-file size, total storage consumed by the dataset, or a property of an abstract problem independent of the chosen computation.[5]
WSS becomes a capacity constraint when compared with the fast memory available to the computation.[6] If the active data cannot fit in physical memory, a virtual-memory system must repeatedly move pages between memory and a slower hierarchy level.[7] When the access pattern causes continual replacement and reload, the system can enter thrashing and its performance can collapse even though the computation is otherwise correct.[8] High-performance systems therefore size memory and choose decompositions partly from the expected WSS of target problem instances.[9]
The measurement must declare what data and execution interval count. Peak WSS, typical WSS, and the operating system's observed resident set can differ; WSS denotes the problem-solving footprint needed to complete the computation.[10] Merely allocating a large address space does not establish a large working set if those data are never required; conversely, transient intermediate data count when they are part of the memory needed to complete the computation.[11]
Structural Signature¶
Sig role-phrases:
- the problem instance — the particular input and size for which memory demand is being characterized
- the executable computation — the algorithm and implementation choices that determine which values are retained, streamed, recomputed, or discarded
- the required working data — input portions that must remain available and live intermediate results whose simultaneous necessity contributes to the declared footprint
- the measurement window — the whole run or named phase over which peak, typical, or representative demand is defined
- the inclusion convention — the rule distinguishing necessary resident data from executable bytes, stored dataset size, or merely reserved address space
- the working-set size — the byte footprint of the included input and intermediate data under that convention
- the fast-memory capacity — the cache, physical memory, or other tier against which the footprint is compared
- the capacity-regime branch — fitting within the intended fast tier permits residency, whereas overflow forces movement to a slower tier
- the thrashing limit — repeated replacement and reload dominate useful progress when footprint and access pattern defeat available memory
- the runtime boundary — working-set size alone does not determine locality, bandwidth, latency, sharing, contention, or total execution time
What It Is Not¶
- Not the working set itself. The working set is the required input and intermediate data; working-set size is the memory footprint assigned to that collection under a declared convention.
- Not executable-file or stored-dataset size. Code bytes and data at rest need not equal the material that must be available while a particular computation runs.
- Not reserved virtual address space. Allocated but untouched pages do not establish required residency, while transient intermediates can matter when they set the active footprint.
- Not a fixed property of an algorithm independent of execution. Input size and implementation choices such as retaining, streaming, tiling, recomputing, or discarding values can change simultaneous memory demand.
- Not automatically identical to an operating system's instantaneous resident-set report. Peak, phase-local, typical, and observed resident measures answer different questions and require an explicit measurement window and inclusion rule.
- Not by itself a runtime or thrashing prediction. Capacity mismatch creates pressure for slower-tier movement, but access order, locality, bandwidth, page policy, sharing, and contention determine the resulting performance.[12]
Scope of Application¶
Working-set size applies across computations wherever the input and intermediate data simultaneously needed for a declared problem instance and execution interval can be identified and assigned a byte footprint. The measure is literal only when its inclusion rule, peak or representative convention, implementation, and comparison memory tier are stated. Without a defined required-data set or interval there is no WSS, and capacity fit alone does not determine runtime without locality, bandwidth, latency, page policy, sharing, and contention.
- Algorithm and implementation analysis — WSS compares how retaining, streaming, recomputing, tiling, or discarding intermediates changes memory demand while the problem instance and result remain fixed.
- High-performance-computing system design — expected footprints of target problem instances inform physical-memory capacity, node configuration, and problem decomposition before costly runs are provisioned.[13]
- Virtual-memory capacity analysis — comparing WSS with available RAM identifies executions that must move data to a slower hierarchy tier and therefore face paging pressure.
- Thrashing diagnosis — a footprint that defeats available memory is combined with access and replacement behavior to explain repeated swap or page traffic that dominates useful computation.
- Phase-local performance engineering — peak, typical, and phase-specific working sets isolate memory-intensive kernels or stages rather than treating one instantaneous resident-set observation as the whole run.
- Memory profiling and measurement — instrumentation estimates required live input and intermediate data under a declared window while separating it from executable bytes, stored datasets, total allocation, and untouched address space.
- Memory-footprint optimization — controlled changes to data layout, blocking, streaming, lifetime, and recomputation are evaluated by whether they reduce simultaneous necessity and allow the computation to fit a faster tier.[14]
- Instance and workload comparison — input size, implementation, and execution convention are held explicit when comparing WSS across runs, machines, or algorithms.
Clarity¶
Working-set size separates the memory a computation needs for a particular problem instance from executable size, stored dataset size, or allocated address space. Input and intermediate data count when they are needed during execution; reserved but untouched pages do not. The quantity can therefore change with input size and with an implementation’s choices to retain, stream, recompute, or discard intermediates even when the underlying algorithmic task is unchanged.
The term also prevents peak, typical, and operating-system-observed resident memory from being compared without declaring the interval and measurement convention. Its capacity consequence is likewise conditional: performance deteriorates when the active footprint exceeds available fast memory and the access pattern forces repeated movement through the hierarchy. The engineering question is: which data must be simultaneously available during this execution, what is their peak or representative footprint, and does that footprint fit the memory tier whose performance the computation assumes?
Manages Complexity¶
A computation may allocate many objects, touch only part of a dataset, create transient intermediates, and move data across caches, main memory, and storage. Working-set size reduces this execution detail to the footprint of the input and intermediate data that must be available over a declared interval. The practitioner tracks the problem instance, implementation, included data, measurement window, and peak or representative footprint.
Comparing that footprint with memory-tier capacity exposes the main performance regimes. When the active set fits in the intended fast tier, repeated accesses can remain local; when it exceeds physical memory, paging becomes necessary; when replacement and reload dominate progress, the execution enters thrashing. Streaming, tiling, recomputation, and earlier disposal of intermediates are recognizable branches because each changes simultaneous data residency even if the mathematical result is unchanged.
Compression stops before predicting runtime. WSS alone does not specify access order, locality, bandwidth, latency, page-replacement policy, sharing, compression, or contention, and it need not equal address-space size or an operating system's instantaneous resident-set report. Those details must be restored to explain actual hierarchy traffic and performance.
Abstract Reasoning¶
Working-set size supports a capacity-regime inference. From the peak or representative footprint of simultaneously needed input and intermediate data to its comparison with available physical memory, an engineer can predict whether the computation can remain resident or must rely on a slower hierarchy level. If replacement and reload recur faster than useful progress, the footprint-capacity mismatch provides a diagnosis of thrashing rather than of an incorrect algorithmic result.
It also permits interventionist comparison across implementations. From retaining every intermediate to a larger simultaneous footprint, and from streaming, tiling, recomputing, or discarding intermediates earlier to a potentially smaller footprint, one can compare memory demand while holding the problem instance and mathematical output fixed. This isolates memory-liveness choices from changes in the underlying problem.
Measurement reasoning must move from an explicitly declared execution interval and inclusion rule to the reported WSS. A peak over the whole run, a phase-local working set, and an operating system's instantaneous resident set answer different questions. Allocated but untouched address space does not by itself establish required residency; conversely, transient data can matter if they set the peak.
These inferences stop short of runtime prediction. Equal WSS values can produce different performance under different access orders, locality, bandwidth, page-replacement policies, sharing, or contention, so hierarchy traffic and timing require those variables in addition to footprint.
Knowledge Transfer¶
Within computer performance engineering, working-set size transfers literally across programs, problem instances, implementations, memory hierarchies, and measurement tools when the footprint is tied to the data simultaneously needed over a declared interval. The cargo that carries intact is the instance, input and intermediate data, execution strategy, measurement window, peak or representative convention, and capacity comparison. Diagnostics transfer by changing retention, streaming, recomputation, or tiling and observing whether residency, paging, or thrashing changes.
This is (C) a performance measure wherever those operational definitions are preserved. The home-bound cargo is computation, memory residency, address-space and page behavior, and the selected hierarchy tier. Stored dataset size, executable size, reserved virtual memory, or total allocation is not interchangeable with WSS. The stopping boundary is simultaneous necessity under a measurement convention: without that relation, a byte count cannot predict the memory regime, and even a valid WSS does not by itself predict all cache locality or runtime costs.
Examples¶
Canonical¶
Suppose one run of a numerical program requires a 2 GiB input array and three 1 GiB intermediate arrays to be simultaneously live during its principal solve phase.[15] Under an inclusion rule that counts those required data but excludes the executable and untouched reserved address space, its phase-local working-set size is 2 + 1 + 1 + 1 = 5 GiB.[16] On a machine with 8 GiB of usable physical memory, that footprint can remain resident.[17] If the same instance instead retains four additional 1 GiB intermediates, the footprint becomes 9 GiB and exceeds that capacity; a virtual-memory system must move some pages to a slower tier.[18] The arithmetic identifies the capacity regime, not the eventual runtime: an access pattern that seldom revisits evicted pages may slow modestly, while continual replacement and reload can produce thrashing.[19]
Mapped back: The particular numerical input is the problem instance, and the program's retention choices define the executable computation. The input plus simultaneously live intermediates are the required working data; the principal solve phase is the measurement window; excluding code and untouched reservations is the inclusion convention. Five or nine GiB is the working-set size, 8 GiB is the fast-memory capacity, and fit versus overflow is the capacity-regime branch. Repeated paging would reach the thrashing limit, while the unresolved effect of locality and access order preserves the runtime boundary.
Applied / In Practice¶
When specifying a high-performance-computing node for a target simulation workload, an engineer inventories the input state and peak simultaneously needed intermediate fields for the largest intended problem instance. That measured or modeled peak—not the installation package or the total archived dataset—is compared with usable memory per node. If it does not fit, the engineer can increase node memory, distribute the instance across more nodes, or change the implementation to stream, tile, recompute, or release intermediates earlier. A second profile then tests whether the intervention actually reduced simultaneous necessity. This is a working-set-size decision even before a production run exists; it becomes a thrashing diagnosis only if observed hierarchy traffic shows repeated replacement and reload dominating useful work.
Mapped back: The largest target simulation is the problem instance; its decomposition and retention strategy are the executable computation; its live input state and intermediate fields are the required working data. Peak profiling supplies the measurement window and the inclusion convention, yielding the working-set size for comparison with the fast-memory capacity. Provisioning or restructuring responds to the capacity-regime branch; observed repeated swap traffic, rather than footprint alone, establishes the thrashing limit. The need to restore locality, bandwidth, latency, sharing, and contention keeps the runtime boundary intact.
Structural Tensions¶
T1: Required live data versus allocated address space.
A program may reserve or allocate a large region while touching only a small part of it, or may create short-lived intermediates whose simultaneous necessity sets a real peak despite modest steady-state allocation. Counting all reserved bytes overstates the working set; counting only a convenient snapshot can miss required transient data. The inclusion convention must therefore follow computational necessity over the declared interval rather than equate footprint with any single memory-accounting field. Diagnostic: Which input and intermediate bytes must be available together for progress, and which reported bytes are merely reserved, executable, inactive, or outside the specified working-data set?
T2: Problem-instance dependence versus algorithm comparison.
Working-set size belongs to an algorithm or program running a particular instance, not to an abstract task without input scale or implementation choices. Holding the instance fixed reveals whether streaming, tiling, retention, or recomputation changes simultaneous memory demand; changing the instance may dominate those differences. Treating WSS as a timeless algorithm label makes unlike runs appear comparable, while refusing all abstraction prevents useful scaling analysis. Diagnostic: Are the instance, implementation, and inclusion convention controlled closely enough that the observed footprint difference can be attributed to the factor under comparison?
T3: Peak footprint versus representative footprint.
Peak WSS protects capacity planning against brief memory maxima, but it may exaggerate the footprint sustained through most of the run. A typical or phase-local measure describes common demand better, yet can miss the short interval that determines whether the computation fits at all. Neither summary is intrinsically correct; each answers a different engineering question and depends on the measurement window. Diagnostic: Is the decision governed by the maximum simultaneous requirement, a named phase, or representative residency, and does the reported WSS use that same temporal convention?
T4: Lower footprint versus additional computation.
Discarding an intermediate and recomputing it later can reduce simultaneous memory demand, while retaining it spends capacity to avoid repeated work. Streaming and tiling make similar exchanges between footprint, data movement, and computation. Minimizing WSS alone can therefore worsen runtime or energy use; maximizing retention can push the execution into a slower memory regime. The optimal balance depends on hierarchy costs and access patterns that the size measure itself does not encode. Diagnostic: Does the footprint reduction preserve useful performance once recomputation and extra movement are restored to the analysis?
T5: Capacity fit versus access locality.
A working set smaller than physical memory can reside there in principle, but poor access order may still produce cache misses, bandwidth pressure, or contention. Conversely, a footprint that exceeds a fast tier need not collapse performance if accesses stream predictably and evicted data are not repeatedly revisited. Capacity comparison identifies a regime boundary, not a complete runtime model. Diagnostic: Is the performance claim supported only by WSS-to-capacity fit, or also by the locality, reuse, bandwidth, and sharing behavior that determines actual hierarchy traffic?
T6: Memory pressure versus thrashing diagnosis.
Exceeding available memory creates pressure to use a slower hierarchy level, but thrashing requires repeated replacement and reload to dominate useful progress. Labeling any overflow as thrashing overstates the consequence; treating heavy swap traffic as unrelated to WSS ignores the capacity mismatch that makes it possible. A defensible diagnosis joins footprint, available capacity, access recurrence, and observed movement. Diagnostic: Does the execution merely spill beyond a memory tier, or does its replacement pattern repeatedly evict data needed again soon enough that movement overwhelms computation?
T7: Working-Set Size autonomy versus reduction to its Measurement prerequisite.
Working-Set Size strictly presupposes the parent Prime Measurement: the target attribute is simultaneous required-data footprint, bytes supply the scale and unit, profiling or modeling supplies the instrument and procedure, and a declared execution interval and inclusion rule fix the frame. It is not a kind of Measurement—WSS is the resulting computer-performance quantity, not the operation that maps an attribute to a value. Removing the measurement chain makes the size undefined, while Measurement remains complete without a problem instance, live intermediate data, memory tier, or thrashing boundary. Diagnostic: Does the account supply the complete Measurement prerequisite and then measure simultaneously needed working data for a specified computation and interval, or does it report an unframed byte count?
Structural–Framed Character¶
Working-Set Size is mixed-structural: it has a thin, reusable footprint-and-capacity skeleton, but its identity remains fixed to the memory behavior of a specified computation. Its smallest reviewed Prime skeleton is Measurement: an identified required-data attribute is mapped to bytes through an inclusion rule, execution window, and profiling or modeling procedure. The cross-domain reach belongs to that Prime. The candidate adds a problem instance, executable retention strategy, fast-memory tier, capacity branch, and thrashing limit that Measurement does not require.
Its evaluative_weight is low because a footprint is descriptive, although its adequacy for a machine becomes consequential only relative to capacity and performance aims. Its human_practice_bound is low: programs can require data and exceed memory without human interpretation, even though engineers choose the convention and instrument. Its institutional_origin is low because the abstraction comes from computer performance engineering rather than from a rule-making institution, while implementation and operating-system conventions still frame what is counted. Its vocab_travels is medium-low: footprint, working data, residency, and capacity transfer among computing settings, but working set, paging, and thrashing retain a memory-system accent. Its import_vs_recognize judgment is mixed: recognizing simultaneously necessary data and a capacity mismatch tracks an execution property, whereas calling a byte count WSS imports a declared interval and inclusion convention.
Its character: the structural skeleton is a measured footprint compared with a capacity threshold, while the domain accent supplies input and intermediate data, implementation-dependent liveness, memory tiers, paging, and the distinction between pressure and thrashing. Removing those computing-specific roles leaves Measurement plus a generic capacity comparison, not Working-Set Size; retaining them preserves a domain-specific abstraction whose structural content is real but incomplete without its home-domain frame.
Structural Core vs. Domain Accent¶
Working-Set Size is domain-specific even though part of its organization can be isolated as a Prime-level structure: a target attribute is assigned a value under a declared procedure, and that value is compared with a capacity boundary. The first operation is supplied by Measurement; the second is a thinner capacity comparison rather than an asserted catalog parent.
What is skeletal (could lift toward a cross-domain prime). The portable skeleton is the Measurement chain: identify a target and attribute, choose a scale and unit, apply an instrument or estimation procedure under a stated frame, and report a value whose uncertainty and interpretation remain tied to those choices. That complete signature recurs in at least three unrelated domains—for example, physical metrology, clinical assessment, and economic statistics—without referring to programs or memory. Working-Set Size strictly presupposes this chain because its byte value is undefined unless the computation, required-data attribute, inclusion rule, execution window, and profiling or modeling procedure have been fixed. It is not a specialization of Measurement: the candidate is the resulting computer-performance quantity, whereas Measurement is the operation that makes such a quantity reportable.
What is domain-bound. The irreducible accent is a particular problem instance executed by an implementation whose input and intermediate values have lifetimes. It supplies the required working data, byte footprint, peak or representative interval, executable-versus-data exclusions, and comparison with a cache, physical-memory, or other fast tier. It also supplies the regime distinction between residency, slower-tier movement, and thrashing, while preserving the boundary that access order, locality, bandwidth, replacement policy, sharing, and contention are not encoded by size alone. Remove these computing roles and the name no longer distinguishes Working-Set Size from any other measured quantity or generic capacity comparison.
Why this does not clear the prime bar. Stripping the memory vocabulary leaves Measurement plus an underspecified resource-versus-capacity relation, not a self-sufficient abstraction with Working-Set Size's recognition and collapse tests. Conversely, remove the Measurement chain while retaining talk of live input, intermediates, and memory tiers, and there is no defensible WSS value—only an informal claim that a computation uses some memory. The domain accent therefore cannot be detached without destroying the candidate, while the presupposed Prime is far broader than this one performance measure; both removal directions confirm that the current node belongs in computer performance engineering rather than in the Prime inventory.
Instantiates / Related Primes¶
This entry presupposes Measurement.
Strictly presupposes — Measurement (Measurement). A working-set-size value is meaningful only through an operational chain that identifies the computation and required-data attribute, maps it to bytes through an inclusion rule and measurement window, and reports the resulting value under a stated convention. Measurement supplies that target-to-value coupling; Working-Set Size names the computing-specific quantity produced by it and adds the fast-memory-capacity and thrashing branches.
Related to — Scale (Scale). WSS can be compared across input sizes, memory tiers, and temporal windows, and those comparisons may reveal different capacity regimes. A single WSS value does not, however, require Scale's cross-band ontology, level-specific interaction laws, or coupling between adjacent bands.
Decline — Measurement (Measurement) as a subsumption parent. Working-Set Size is the instance- and convention-indexed memory-demand quantity, not the operation that couples a target attribute to a scale through an instrument and procedure. Measuring or estimating WSS establishes the value but is not identical to the value's domain identity.
Relationships to Other Abstractions¶
Current abstraction Working set size Domain-specific
Parents (1) — more general patterns this builds on
-
Working set size presupposes Measurement Prime
A working-set-size value is meaningful only through an operational chain that identifies the computation and required-data attribute, maps it to bytes through an inclusion rule and measurement window, and reports the resulting value under a stated convention.Measurement supplies that target-to-value coupling; Working-Set Size names the computing-specific quantity produced by it and adds the fast-memory-capacity and thrashing branches.
Hierarchy path (1) — routes to 1 parentless root
- Working set size → Measurement
Neighborhood in Abstraction Space¶
Working set size sits in a sparse region of the domain-specific corpus (72nd percentile for distinctiveness): few abstractions share its structure, so a faithful description tends to retrieve it precisely.
Family — Program Execution & Runtime Concepts (27 abstractions)
Nearest neighbors
- Rematerialization — 0.85
- Program Profiling — 0.85
- Buffer Overflow — 0.83
- Analysis of algorithms — 0.83
- Sun–Ni Law — 0.83
Computed from structural-signature embeddings · 2026-10-08
Not to Be Confused With¶
- The working set. The working set is the collection of input and intermediate data needed by the computation; working-set size is the byte footprint assigned to that collection. Tell: ask whether the object is a set of live data or a measured amount of memory.
- Resident set size. Resident set size reports pages currently resident in physical memory, whereas WSS characterizes required working data over a declared execution window. Tell: ask whether the value is an operating-system snapshot or a computation-indexed requirement.
- Virtual address-space size. Address-space size includes memory ranges reserved for a process even when they are untouched or not simultaneously needed. Tell: ask whether mapped or allocated bytes count automatically, or only data required for progress.
- Stored dataset size. Dataset size measures data at rest; WSS counts the input portions and intermediate results that must be available during a particular execution. Tell: ask whether the bytes describe stored material or simultaneous runtime necessity.
- Space complexity. Space complexity describes how an algorithm's memory requirement grows with input size, whereas a WSS is a footprint for a specified instance, implementation, interval, and inclusion rule. Tell: ask whether the claim is an asymptotic function or an operational byte quantity.
- Thrashing. Thrashing is the performance regime in which repeated replacement and reload dominate useful progress; a large WSS creates capacity pressure but does not establish that access behavior. Tell: ask whether only footprint exceeds a tier or recurrent page movement has been shown to cause collapse.
- Working memory. Working memory is a cognitive system for temporarily maintaining and manipulating information, not the execution-data footprint of a computer program. Tell: ask whether the subject is human cognition or computational memory capacity.
References¶
[1] Unverified encyclopedia synthesis; no authoritative source located for the claim as written. ↩
[2] Unverified encyclopedia synthesis; no authoritative source located for the claim as written. ↩
[3] Unverified encyclopedia synthesis; no authoritative source located for the claim as written. ↩
[4] Unverified encyclopedia synthesis; no authoritative source located for the claim as written. ↩
[5] Unverified encyclopedia synthesis; no authoritative source located for the claim as written. ↩
[6] Unverified encyclopedia synthesis; no authoritative source located for the claim as written. ↩
[7] Unverified encyclopedia synthesis; no authoritative source located for the claim as written. ↩
[8] Peter J. Denning, “The Working Set Model for Program Behavior,” Communications of the ACM 11 (1968) (source). registry ↩
[9] Unverified encyclopedia synthesis; no authoritative source located for the claim as written. ↩
[10] Unverified encyclopedia synthesis; no authoritative source located for the claim as written. ↩
[11] Unverified encyclopedia synthesis; no authoritative source located for the claim as written. ↩
[12] Unverified encyclopedia synthesis; no authoritative source located for the claim as written. ↩
[13] Unverified encyclopedia synthesis; no authoritative source located for the claim as written. ↩
[14] Unverified encyclopedia synthesis; no authoritative source located for the claim as written. ↩
[15] Unverified encyclopedia synthesis; no authoritative source located for the claim as written. ↩
[16] Unverified encyclopedia synthesis; no authoritative source located for the claim as written. ↩
[17] Unverified encyclopedia synthesis; no authoritative source located for the claim as written. ↩
[18] Unverified encyclopedia synthesis; no authoritative source located for the claim as written. ↩
[19] Unverified encyclopedia synthesis; no authoritative source located for the claim as written. ↩