Skip to content

Tensions in Practice: Narrow datapath reuse in tension with word completion time

Four-bit operation · stipulated equal unit delay

Two four-bit words, 1010 and 1100, are combined with XOR: each output bit is 1 when its two input bits differ and 0 when they match. One unit can process the four positions in sequence, or four identical units can process them together. The output is 0110 either way. In this declared model a unit takes one tick and all input bits are already available.

Reuse one narrow operation unit

Spend less replicated bit-operation hardware on a word.

Finish the word sooner

Compute independent positions together so all output bits are ready after one tick.

Why these aims pull against each other

Time sharing moves bit positions through one unit. Parallel width duplicates the unit so those positions need not wait for each other.

Compare the arrangements

One unit

Process bit positions 1 through 4 in sequence, keeping each result until the word is complete.

Bitwise XOR · 1010 with 1100 gives 0110
InputsReady tickOutput
Bit 11, 110
Bit 20, 121
Bit 31, 031
Bit 40, 040
What it protects
Only one XOR operation unit is needed.
What it costs
The complete four-bit output is ready at tick 4; storage and sequencing are also needed.
When it fits
A narrow operation unit is worth the longer completion time and inputs can be supplied in the declared order.

Illustration note: Ticks are identical ideal unit-delay steps. This arithmetic counts operation units, not total transistor area, power, external wires or measured clock cycles of a real processor.

Four units

Run four independent XOR units at the first tick and collect their outputs.

Bitwise XOR · 1010 with 1100 gives 0110
InputsReady tickOutput
Bit 11, 110
Bit 20, 111
Bit 31, 011
Bit 40, 010
What it protects
The complete word is ready at tick 1 in the stipulated model.
What it costs
Four operation units and concurrent input/output paths replace one reused unit.
When it fits
The word is available in parallel and the added width is worthwhile for completion time.

Illustration note: XOR positions are independent; this is not a ripple-carry adder or a claim that every word operation parallelizes into one real clock cycle.

What this illustration does—and does not—establish

Bit Serial Architecture supplies reuse of a narrow datapath and warns against universal N-cycle or area laws. The XOR model makes its own fixed unit-delay assumptions explicit.

  • The comparison is internal computation, not serial versus parallel transmission of an already computed word.
  • Actual area, energy, clock rate, storage and data bandwidth can change the practical trade.
  • Unlike a pipeline overlapping different words, this arrangement changes how many positions within one word execute together.

Source entries

Bit-Serial Architecture

Domain-specific abstraction · Source of the tension

Bit-Serial Architecture: Narrow hardware per lane vs. longer per-word latency supplies the conflict examined here.

Narrow hardware per lane vs. longer per-word latency

Reusing a small datapath can save logic and wiring, while a word operation needs multiple internal steps. Pipelining and clock choices change the result without dissolving the trade.

Read the source section