Skip to content

Spatial architecture

A computing organization that maps operations and dataflow across an array of processing elements with direct inter-element communication.

Core Idea

A spatial architecture maps a computation across an array of processing elements (PEs) and arranges data movement among them. Operations occupy positions in the array, while a dataflow schedule determines when operands and intermediate values arrive or move onward. Direct inter-element communication distinguishes this organization from merely running many processors independently through a common memory path. Which values are stored or reused locally varies by design and workload.[ref-dde928a32df9][ref-e9ac03dae815]

The structure is visible in two unlike chips. Eyeriss maps convolutional-neural-network work onto a PE array with row-stationary dataflow. The original TPU uses a matrix-multiply systolic array with a different arrangement of inputs, weights, and accumulation. They share spatial mapping and communication, but neither chip's exact operand schedule defines every spatial architecture.[ref-dde928a32df9][ref-e9ac03dae815]

Scope of Application

The literal setting is spatial computing hardware that maps operations and data movement onto cooperating PE arrays. Eyeriss is a convolutional accelerator: its row-stationary design aims to reuse filter weights and input-feature pixels while reducing partial-sum accumulation and movement costs. The TPU matrix unit is another setting: input data enter from the left, weights are loaded from above, and 256-element multiply-accumulate operations travel diagonally through its 256×256 array. Dedicated accumulators sit outside that matrix unit.[ref-dde928a32df9][ref-e9ac03dae815]

These cases span different kernels and dataflows within computer architecture. Neither paper establishes that every PE array is a coarse-grained reconfigurable array, that all spatial designs use the TPU's wavefront, or that every mapped workload fills its array efficiently.

Clarity

The abstraction separates parallel hardware capacity from the mapping that uses it. Saying that a chip has many PEs does not explain which computation each performs, how operands reach them, or whether they sit idle. Conversely, saying that data are reused does not specify whether reuse happens in caches, registers, or an inter-PE flow. The spatial claim requires a mapped computation over communicating elements.[ref-dde928a32df9][ref-e9ac03dae815]

It also makes “locality” more exact. Eyeriss deliberately reduces some data transfers, but reuse of filter weights and input pixels is not the same as reuse of partial sums. TPU's diagonal operation wavefront likewise should not be misdescribed as all results flowing only through neighboring PEs; its accumulators are a separate structure.[ref-dde928a32df9][ref-e9ac03dae815]

Manages Complexity

Array size, wiring, dataflow, storage, kernel shape, and compiler scheduling interact. The role map reduces an initial comparison to four linked questions: what operation graph is mapped, where are its PEs, how do values move among them, and which utilization or transfer costs appear? This lets a CNN accelerator and a matrix unit be compared without flattening them into one algorithm.[ref-dde928a32df9][ref-e9ac03dae815]

The map does not replace quantitative evaluation. A dataflow optimized for one reuse opportunity can incur other movement, while a larger matrix array can be poorly occupied by a shape that tiles awkwardly. The original TPU analysis shows why more arithmetic units do not automatically mean less elapsed time for every matrix.[ref-dde928a32df9][ref-e9ac03dae815]

Abstract Reasoning

Begin with the kernel's operations and dependencies. Trace how a proposed design assigns them to PEs and moves each operand or intermediate along the physical communication paths. Then ask which values are retained, forwarded, broadcast, or accumulated, and which transfers still reach a more distant memory structure. This makes a claimed energy or speed advantage testable rather than a slogan about spatiality.[^ref-dde928a32df9]

Next change the workload shape. If fewer PEs are occupied or the dataflow must be retiled, the same hardware can deliver a different result. A valid inference is conditional: this mapping reduces a specified movement or improves a specified execution under measured or modeled conditions. It is not proof that local communication alone guarantees superior performance.[^ref-e9ac03dae815]

Knowledge Transfer

Within accelerator design, the same analysis applies to unlike computational kernels: a CNN convolution can map to Eyeriss's row-stationary PE array, while matrix multiplication can map to TPU's systolic matrix unit. The role mapping transfers literally; the stationary values, links, and accumulation path must be redrawn for each architecture.[ref-dde928a32df9][ref-e9ac03dae815]

Outside computing hardware, people may call an organization “spatial” because work is distributed across places. That is only analogy unless actual operations, processing elements, mapped dataflow, and direct communication retain their roles. A more general theory of distributed work may be a future-prime question; it does not make this named hardware organization a Prime.

Example

Jouppi and colleagues describe a 256×256 matrix unit in the original Tensor Processing Unit. Input data enter from the left, weights are loaded from the top, and a 256-element multiply-accumulate operation moves through the array as a diagonal wavefront. Accumulation uses dedicated storage outside the matrix unit. This example is about that original design, not a claim about all TPU generations.[^ref-e9ac03dae815]

Mapped back: matrix multiplication is the operation graph or kernel; the 256×256 matrix unit is the PE array; top-loaded weights and left-entering inputs establish mapping and dataflow; the wavefront is direct inter-element data movement; matrix shape and tiling bound utilization and movement cost. The paper's occupancy analysis prevents an unconditional speed claim.

Neighborhood in Abstraction Space

Spatial architecture sits in a sparse region of the domain-specific corpus (94th percentile for distinctiveness): few abstractions share its structure, so a faithful description tends to retrieve it precisely.

Family — Processor Architecture & Instruction Sets (8 abstractions)

Nearest neighbors

Computed from structural-signature embeddings · 2026-10-08

Not to Be Confused With

  • Generic parallel processing: many operations can run at once without a mapped direct inter-PE dataflow.
  • A coarse-grained reconfigurable array: some spatial architectures are programmable, but the class also includes fixed systolic designs.
  • One dataflow policy: row-stationary mapping and a TPU-style systolic wavefront are alternatives within the larger spatial organization.[ref-dde928a32df9][ref-e9ac03dae815]
  • A guarantee of lower energy or latency: workload shape and data movement decide measured benefit.[ref-dde928a32df9][ref-e9ac03dae815]

The typed hierarchy review leaves this entry a provisional unparented root. Computer Architecture requires whole-system organization, Reconfigurable Computing requires retargetable fabric, and Locality of Reference describes an access regularity rather than the mapped PE-array organization.

References

[^ref-dde928a32df9]: Yu-Hsin Chen, Joel S. Emer, and Vivienne Sze, “Eyeriss: A Spatial Architecture for Energy-Efficient Dataflow for Convolutional Neural Networks,” Proceedings of ISCA (2016), DOI 10.1109/ISCA.2016.40, especially printed pp. 367–369 and §V/Fig. 6, pp. 371–373. https://www.cs.cmu.edu/~18742/papers/Chen2016.pdf [^ref-e9ac03dae815]: Norman P. Jouppi et al., “In-Datacenter Performance Analysis of a Tensor Processing Unit,” Proceedings of ISCA (2017), especially §2/Fig. 4, PDF p. 4, §7 p. 8 and conclusion p. 10. https://web.stanford.edu/class/cs114/readings/tpu.pdf