Cache-Only Memory Architecture¶
Make distributed node-local main memories cache-like, home-free stores that migrate and replicate shared data according to demand.
Core Idea¶
A cache-only memory architecture (COMA) presents shared memory across a multiprocessor while treating each node's substantial local main memory as an attraction memory—a cache for data that may currently reside at any node. Unlike a conventional non-uniform-memory-access system, a datum has no fixed physical home determined by its address. Accesses draw data toward processors that use them, and reads may replicate copies. The system must locate valid copies, keep concurrent access coherent, and ensure replacement never discards the last remaining copy.[1][2]
This is an architecture pattern, not simply the historical Data Diffusion Machine (DDM). The DDM and Kendall Square Research KSR-1's ALLCACHE system are two different realizations with different interconnect details.[1][3]
Structural Signature¶
Sig role-phrases:
- Processor–memory nodes — multiple processors have local memories but see a shared address space.[1]
- Attraction memory — local main memory is cache-organized rather than a permanently assigned address range; a block's address does not impose one home node.[1][2]
- Demand-driven movement and replication — data migrate or copy toward processors that reference them. This is an ability of the architecture, not a guarantee that every access is local.[1]
- Location and coherence protocol — distributed state finds copies and coordinates their valid/read/write status.[1]
- Last-copy invariant — a block that exists in no fixed backing memory cannot have its only copy silently evicted.[1][2]
A hierarchical directory is one way to realize the location mechanism, as in DDM; the identity does not require that exact topology.
What It Is Not¶
It is not ordinary NUMA with a fixed home node for each address plus private caches. COMA removes fixed data-home placement at the main-memory level. It is not merely a CPU cache either: a CPU cache normally sits in front of an independent backing store, whereas COMA's attraction memories together are the shared store.[1][2]
Nor does “cache-only” mean no memory exists. The distributed local memories collectively hold data; the difficult part is retaining one valid copy when cache replacement would otherwise evict it.
Scope of Application¶
COMA addresses the locality problem in distributed shared-memory multiprocessors: programs see shared addresses, but frequent remote accesses can be expensive. If a working set follows a processor's use, automatic movement and replication can reduce later remote accesses. That benefit is workload-dependent and comes with location, metadata, replacement, and replication costs.[1][2]
The original DDM paper introduces COMA as a class and DDM as an example. KSR-1's ALLCACHE uses a ring-of-rings system in which a requested datum can be in several distributed caches; first access latency varies until the data are found and pulled local.[1][3]
Clarity¶
An address names a logical location; it need not name a permanent physical memory owner in COMA. The same data block may move or have several valid copies. Coherence decides which copies can satisfy reads and how writes alter or invalidate them. The design is not “free migration”: locating a missing block and maintaining coherence require machinery.[1]
The last-copy rule differs from ordinary cache replacement. When an ordinary cache discards a line, a backing home may still hold it. When the entire main memory is cache-like, the protocol has to preserve or transfer the sole copy before eviction.[2]
Manages Complexity¶
Fixed-home NUMA asks placement and scheduling to make use of a preassigned home. COMA moves some placement work into the memory system, allowing reference patterns to shape physical location. That can simplify the programmer's data-placement burden, but the hardware must now maintain associative tags, distributed location information, coherence states, and safe replacement.[1][2]
The resulting tradeoff is not a blanket speed claim. If a workload repeatedly reuses a moved block locally, movement can pay off. If it streams data or shares a write-heavy block across nodes, movement and invalidation can dominate.
Abstract Reasoning¶
Consider a shared datum initially held near processor A. Processor B requests it; the location protocol finds a valid copy, transfers the block to B's attraction memory, and records the new copy/coherence state. Repeated B reads can now be local. If A and B both write, the protocol must serialize or otherwise coordinate ownership. If replacement tries to evict a block whose only valid copy is at B, it must preserve that block elsewhere first.[1][2]
The invariant is independent of interconnect: DDM uses a hierarchy; another implementation may use different directory machinery. What must remain is home-free location, coherent discoverability, and last-copy preservation.
Knowledge Transfer¶
DDM's hierarchical buses and KSR-1's rings instantiate the same abstraction: distributed local memories collectively emulate shared memory while data follow use. The common roles transfer; the directory layout, latency, and migration policy do not. A software cache or ordinary demand paging is an analogy unless it also removes fixed main-memory homes and maintains this shared-store invariant.[1][3]
Examples¶
Data Diffusion Machine¶
Hagersten, Landin and Haridi's DDM organizes processor/memory pairs under a hierarchy that locates and coordinates distributed data copies.[1] Mapped back: the pairs are nodes; their large local stores are attraction memories; address and physical location are decoupled; accesses migrate or replicate blocks; hierarchical control locates copies and enforces coherence; the protocol preserves the last copy. DDM is an implementation, not the definition of every COMA directory.
KSR-1 ALLCACHE¶
The KSR-1 report describes a ring-of-rings distributed shared-memory machine with ALLCACHE; a requested datum may be found in more than one cache, with variable initial access time before it is available locally.[3] Mapped back: processor-local memories collectively provide a shared image; ALLCACHE is its cache-like memory; no permanent home is assumed; first use retrieves data to the requesting node; system-level coherence is needed for distributed copies. The cited report does not establish one particular low-level last-copy replacement algorithm, so that implementation detail is not inferred from it.
Structural Tensions¶
Local reuse versus remote discovery. Demand placement can turn later reads into local hits, but the first miss needs a distributed search and transfer. Diagnostic: Does reuse repay lookup and migration under this workload?[3]
Replication versus capacity. Extra copies help readers but consume main-memory capacity and metadata. Diagnostic: How much unique data fits after replication and directory/tag overhead?[2]
Replacement freedom versus data survival. Attraction memory is cache-like, yet no separate fixed home guarantees persistence. Diagnostic: Where does the final valid copy go when its current holder needs space?[1]
Structural–Framed Character¶
Evaluative weight. COMA's relocation can help some workloads but also incurs overhead; performance superiority is not constitutive. Human-practice bound. Architects choose granularity, directories and interconnect while the no-fixed-home and last-copy safety rules constrain valid implementations.[1][2]
Institutional origin. Multiprocessor memory research introduced the design; neither DDM nor KSR1 alone defines the architecture class. Vocabulary travel. Cache and migration are broad computing words, but shared-address-space memory blocks, tags and coherence have exact machine roles.[1][3]
Import versus recognition. A new machine qualifies when distributed memory acts as cache-only storage with migratory/replicated blocks and no fixed home under coherence and retention rules. A web content cache only borrows part of the pattern. Its character: mixed-structural—a memory-organization architecture framed by workload and implementation economics.
Structural Core vs. Domain Accent¶
Portable skeleton. Demand-based relocation with replica consistency is a future-prime candidate only, not an applied parent. The staged actual edge is composition/presupposes live Shared memory: COMA implements a shared address-space resource, but the hardware architecture is not simply a subtype of the resource.[1]
Domain-bound mechanism. Main memory on nodes acts as a cache for a shared address space; blocks migrate or replicate, coherence tracks copies, and replacement must not destroy the last valid copy. DDM and KSR1 vary implementation details while retaining those organization rules.[1][3]
Why not prime. Distributed content caches may relocate replicas, but they do not thereby implement coherent shared-memory blocks with last-copy retention. The broad cache analogy omits the machine-level invariants that define COMA; no new generic parent is asserted.
Instantiates / Related Primes¶
The proposed, unapplied edge is a composition/presupposes relation to live Shared memory: COMA implements a shared-memory address-space resource over distributed node stores. Shared Memory is not a taxonomic genus for the whole hardware architecture, so a strict subtype edge would misstate the types. Live cpu_cache, cache_coherence and memory_management are component/related identities rather than the full COMA pattern. No canonical edge is changed.
Neighborhood in Abstraction Space¶
Cache-Only Memory Architecture sits in a moderately populated region (56th percentile for distinctiveness): it has near-neighbors but no dense thicket of look-alikes.
Family — Distributed Systems Theorems & Fallacies (19 abstractions)
Nearest neighbors
- Memory Management — 0.86
- Buffer Overflow — 0.86
- Oblivious RAM — 0.86
- Fragmentation (computing) — 0.85
- Array-access analysis — 0.85
Computed from structural-signature embeddings · 2026-10-08
Not to Be Confused With¶
- Fixed-home NUMA plus ordinary processor caches.
- A cache hierarchy whose local caches all retain a separate fixed backing-memory home.
- Shared-memory programming alone, which does not determine physical data placement.
- A claim that COMA always outperforms NUMA or that every COMA uses DDM's hierarchical directory.
References¶
[1] Erik Hagersten, Anders Landin and Seif Haridi, “DDM—A Cache-Only Memory Architecture”, IEEE Computer 25(9), 44–54 (1992), abstract and DDM design. Introduces the COMA class and its DDM example. registry ↩a ↩b ↩c ↩d ↩e ↩f ↩g ↩h ↩i ↩j ↩k ↩l ↩m ↩n ↩o ↩p ↩q ↩r ↩s ↩t
[2] Truman Joe and John L. Hennessy, “Evaluating the Memory Overhead Required for COMA Architectures”, Proceedings of ISCA (1994), abstract and introduction. registry ↩a ↩b ↩c ↩d ↩e ↩f ↩g ↩h ↩i ↩j
[3] Emilia Rosti, Evgenia Smirni, T. D. Wagner, Amy W. Apon and L. W. Dowdy, “The KSR1: Experimentation and Modeling of Poststore”, Oak Ridge National Laboratory report ORNL/TM-12287 (February 1993), abstract and architecture description. registry ↩a ↩b ↩c ↩d ↩e ↩f ↩g