Spatial architecture¶
A computing organization that maps operations and dataflow across an array of processing elements with direct inter-element communication.
Core Idea¶
A spatial architecture maps a computation across an array of processing elements (PEs) and arranges data movement among them. Operations occupy positions in the array, while a dataflow schedule determines when operands and intermediate values arrive or move onward. Direct inter-element communication distinguishes this organization from merely running many processors independently through a common memory path. Which values are stored or reused locally varies by design and workload.[1][2]
The structure is visible in two unlike chips. Eyeriss maps convolutional-neural-network work onto a PE array with row-stationary dataflow. The original TPU uses a matrix-multiply systolic array with a different arrangement of inputs, weights, and accumulation. They share spatial mapping and communication, but neither chip's exact operand schedule defines every spatial architecture.[1][2]
Structural Signature¶
- Operation graph or kernel. A matrix product or convolution supplies work to distribute. Without a computation to map, the array is hardware inventory rather than an executing architecture.
- Processing-element array. Multiple compute units provide spatial positions for scheduled operations. A single serial unit does not instantiate this relation.
- Mapping and dataflow. A design or compiler assigns operations and times the movement of inputs and intermediate values. PE count alone cannot say how a particular kernel runs.
- Direct inter-element data paths. Selected operands or intermediate values communicate within the array. Exact links, broadcast patterns, local storage, and which value is reused vary; neither input nor partial-sum reuse is mandatory in every design.
- Utilization and movement cost. A workload's shape and mapping determine how many elements are productive and what data must move. This is an evaluation boundary, not a guarantee that spatial execution is always faster or less energy-intensive.[1][2]
What It Is Not¶
Spatial architecture does not mean all PEs execute the same program. A systolic array may repeat a multiply-accumulate operation, while another mapped array can assign different work. It does not demand only nearest-neighbor wiring, one stationary operand, or one fixed memory hierarchy. Those are implementation and dataflow choices.[1]
Nor is every parallel computer a spatial architecture in this sense. Many processors can run separate tasks while obtaining all operands through a shared path, without a mapped inter-PE dataflow. And a local-data-access pattern alone is not this architecture: locality may improve an ordinary processor, but the distinctive relation here includes the PE array and direct mapped communication.
Scope of Application¶
The literal setting is spatial computing hardware that maps operations and data movement onto cooperating PE arrays. Eyeriss is a convolutional accelerator: its row-stationary design aims to reuse filter weights and input-feature pixels while reducing partial-sum accumulation and movement costs. The TPU matrix unit is another setting: input data enter from the left, weights are loaded from above, and 256-element multiply-accumulate operations travel diagonally through its 256×256 array. Dedicated accumulators sit outside that matrix unit.[1][2]
These cases span different kernels and dataflows within computer architecture. Neither paper establishes that every PE array is a coarse-grained reconfigurable array, that all spatial designs use the TPU's wavefront, or that every mapped workload fills its array efficiently.
Clarity¶
The abstraction separates parallel hardware capacity from the mapping that uses it. Saying that a chip has many PEs does not explain which computation each performs, how operands reach them, or whether they sit idle. Conversely, saying that data are reused does not specify whether reuse happens in caches, registers, or an inter-PE flow. The spatial claim requires a mapped computation over communicating elements.[1][2]
It also makes “locality” more exact. Eyeriss deliberately reduces some data transfers, but reuse of filter weights and input pixels is not the same as reuse of partial sums. TPU's diagonal operation wavefront likewise should not be misdescribed as all results flowing only through neighboring PEs; its accumulators are a separate structure.[1][2]
Manages Complexity¶
Array size, wiring, dataflow, storage, kernel shape, and compiler scheduling interact. The role map reduces an initial comparison to four linked questions: what operation graph is mapped, where are its PEs, how do values move among them, and which utilization or transfer costs appear? This lets a CNN accelerator and a matrix unit be compared without flattening them into one algorithm.[1][2]
The map does not replace quantitative evaluation. A dataflow optimized for one reuse opportunity can incur other movement, while a larger matrix array can be poorly occupied by a shape that tiles awkwardly. The original TPU analysis shows why more arithmetic units do not automatically mean less elapsed time for every matrix.[1][2]
Abstract Reasoning¶
Begin with the kernel's operations and dependencies. Trace how a proposed design assigns them to PEs and moves each operand or intermediate along the physical communication paths. Then ask which values are retained, forwarded, broadcast, or accumulated, and which transfers still reach a more distant memory structure. This makes a claimed energy or speed advantage testable rather than a slogan about spatiality.[1]
Next change the workload shape. If fewer PEs are occupied or the dataflow must be retiled, the same hardware can deliver a different result. A valid inference is conditional: this mapping reduces a specified movement or improves a specified execution under measured or modeled conditions. It is not proof that local communication alone guarantees superior performance.[2]
Knowledge Transfer¶
Within accelerator design, the same analysis applies to unlike computational kernels: a CNN convolution can map to Eyeriss's row-stationary PE array, while matrix multiplication can map to TPU's systolic matrix unit. The role mapping transfers literally; the stationary values, links, and accumulation path must be redrawn for each architecture.[1][2]
Outside computing hardware, people may call an organization “spatial” because work is distributed across places. That is only analogy unless actual operations, processing elements, mapped dataflow, and direct communication retain their roles. A more general theory of distributed work may be a future-prime question; it does not make this named hardware organization a Prime.
Examples¶
Canonical: original TPU matrix-multiply unit¶
Jouppi and colleagues describe a 256×256 matrix unit in the original Tensor Processing Unit. Input data enter from the left, weights are loaded from the top, and a 256-element multiply-accumulate operation moves through the array as a diagonal wavefront. Accumulation uses dedicated storage outside the matrix unit. This example is about that original design, not a claim about all TPU generations.[2]
Mapped back: matrix multiplication is the operation graph or kernel; the 256×256 matrix unit is the PE array; top-loaded weights and left-entering inputs establish mapping and dataflow; the wavefront is direct inter-element data movement; matrix shape and tiling bound utilization and movement cost. The paper's occupancy analysis prevents an unconditional speed claim.
Applied: Eyeriss convolution accelerator¶
Chen and colleagues map CNN convolution onto a PE array using a row-stationary dataflow. The design seeks local reuse of filter weights and input-feature pixels while limiting the cost of accumulating or moving partial sums. Its mapping differs from TPU's matrix wavefront, yet both organize computation and data movement spatially.[1]
Mapped back: CNN convolution loops are the operation graph or kernel; Eyeriss supplies the PE array; row-stationary scheduling is mapping and dataflow; movement and reuse among PEs use direct inter-element data paths; filter/input reuse, partial-sum cost, and occupancy set utilization and movement cost. This does not make Eyeriss a generic CGRA or its dataflow universal.
Structural Tensions¶
T1: Increase local reuse vs preserve useful occupancy and cheap intermediate movement. Keeping filter weights or input pixels near PEs can reduce distant traffic, but a dataflow that maximizes one reuse objective can raise partial-sum handling costs. A larger array can also leave more elements idle or require awkward tiling for a given matrix shape. Favoring one side without examining the other can make an apparently efficient mapping slower or more costly for the actual workload. Diagnostic: Which operand transfer is saved, what intermediate transfer or idle capacity replaces it, and for which kernel shape?[1][2]
Structural–Framed Character¶
The entry is mixed, leaning structural. The operation-to-PE mapping and communication relation can be checked on unlike chips. Evaluative weight enters when calling one implementation “efficient”: the result depends on workload, energy accounting, occupancy, and the selected baseline. Human practice chooses workloads, compiler mappings, and design objectives; physical circuits constrain which movements and schedules are possible. The term is an engineering architecture category, not membership granted by a standards body.[1][2]
Its vocabulary travels literally between CNN and matrix accelerators when the same roles are present. Importing “spatial architecture” into a social organization merely because work is geographically spread would import the term by metaphor. The possible portable skeleton of mapped work across communicating units is a future-prime question, not a proven current parent. Its character: a reusable hardware-organization pattern whose measured benefits depend on a particular computation and implementation.
Structural Core vs. Domain Accent¶
The skeletal relation is assignment of work to multiple positions with flows among them. A broader Prime might capture such mapped distribution, but none passed the typed test as this entry's necessary parent. The domain accent is computational operations, PE hardware, dataflow schedules, operands and partial sums, and the cost of memory and array occupancy.[1]
The named entry does not clear the Prime bar. If PEs, computational mappings, and value movement are removed, “spatial” becomes a vague statement about location. Computer Architecture, Reconfigurable Computing, and Locality of Reference are live comparison points, and their live definitions do not supply a necessary typed parent.
Instantiates / Related Primes¶
Spatial architecture has no broader abstraction in the encyclopedia yet. Computer Architecture is not a broader category of it: its entry demands whole-system organization including ISA, microarchitecture, memory, and input/output, while a mapped PE array need not supply that identity. Reconfigurable Computing overlaps configurable designs but does not cover a fixed systolic array, so it is not a necessary broader category. Locality of Reference addresses an access pattern that may help explain reuse; it is not the architecture itself.
Neighborhood in Abstraction Space¶
Spatial architecture sits in a sparse region of the domain-specific corpus (94th percentile for distinctiveness): few abstractions share its structure, so a faithful description tends to retrieve it precisely.
Family — Processor Architecture & Instruction Sets (8 abstractions)
Nearest neighbors
- Parallel computing — 0.79
- Analysis of parallel algorithms — 0.78
- Matrix-Free Methods — 0.78
- Unique set size — 0.78
- Single instruction, multiple data — 0.78
Computed from structural-signature embeddings · 2026-10-08
Not to Be Confused With¶
- Generic parallel processing: many operations can run at once without a mapped direct inter-PE dataflow.
- A coarse-grained reconfigurable array: some spatial architectures are programmable, but the class also includes fixed systolic designs.
- One dataflow policy: row-stationary mapping and a TPU-style systolic wavefront are alternatives within the larger spatial organization.[1][2]
- A guarantee of lower energy or latency: workload shape and data movement decide measured benefit.[1][2]
References¶
[1] Yu-Hsin Chen, Joel S. Emer, and Vivienne Sze, “Eyeriss: A Spatial Architecture for Energy-Efficient Dataflow for Convolutional Neural Networks,” Proceedings of ISCA (2016), DOI 10.1109/ISCA.2016.40, especially printed pp. 367–369 and §V/Fig. 6, pp. 371–373. https://www.cs.cmu.edu/~18742/papers/Chen2016.pdf registry ↩a ↩b ↩c ↩d ↩e ↩f ↩g ↩h ↩i ↩j ↩k ↩l ↩m ↩n ↩o ↩p ↩q
[2] Norman P. Jouppi et al., “In-Datacenter Performance Analysis of a Tensor Processing Unit,” Proceedings of ISCA (2017), especially §2/Fig. 4, PDF p. 4, §7 p. 8 and conclusion p. 10. https://web.stanford.edu/class/cs114/readings/tpu.pdf registry ↩a ↩b ↩c ↩d ↩e ↩f ↩g ↩h ↩i ↩j ↩k ↩l ↩m ↩n ↩o