Tensions in Practice: Narrow datapath reuse in tension with word completion time¶
Four-bit operation · stipulated equal unit delay
Two four-bit words, 1010 and 1100, are combined with XOR: each output bit is 1 when its two input bits differ and 0 when they match. One unit can process the four positions in sequence, or four identical units can process them together. The output is 0110 either way. In this declared model a unit takes one tick and all input bits are already available.
Reuse one narrow operation unit
Spend less replicated bit-operation hardware on a word.
Finish the word sooner
Compute independent positions together so all output bits are ready after one tick.
Why these aims pull against each other
Time sharing moves bit positions through one unit. Parallel width duplicates the unit so those positions need not wait for each other.
Choose an arrangement to see what changes and what remains difficult.
Finite illustrative comparisons. Text states carry the meaning; color is not a measured score or universal preference.
What this choice protects
What it costs
When it fits
Compare the arrangements
One unit
Process bit positions 1 through 4 in sequence, keeping each result until the word is complete.
| Inputs | Ready tick | Output | |
|---|---|---|---|
| Bit 1 | 1, 1 | 1 | 0 |
| Bit 2 | 0, 1 | 2 | 1 |
| Bit 3 | 1, 0 | 3 | 1 |
| Bit 4 | 0, 0 | 4 | 0 |
- What it protects
- Only one XOR operation unit is needed.
- What it costs
- The complete four-bit output is ready at tick 4; storage and sequencing are also needed.
- When it fits
- A narrow operation unit is worth the longer completion time and inputs can be supplied in the declared order.
Illustration note: Ticks are identical ideal unit-delay steps. This arithmetic counts operation units, not total transistor area, power, external wires or measured clock cycles of a real processor.
Four units
Run four independent XOR units at the first tick and collect their outputs.
| Inputs | Ready tick | Output | |
|---|---|---|---|
| Bit 1 | 1, 1 | 1 | 0 |
| Bit 2 | 0, 1 | 1 | 1 |
| Bit 3 | 1, 0 | 1 | 1 |
| Bit 4 | 0, 0 | 1 | 0 |
- What it protects
- The complete word is ready at tick 1 in the stipulated model.
- What it costs
- Four operation units and concurrent input/output paths replace one reused unit.
- When it fits
- The word is available in parallel and the added width is worthwhile for completion time.
Illustration note: XOR positions are independent; this is not a ripple-carry adder or a claim that every word operation parallelizes into one real clock cycle.
What this illustration does—and does not—establish
Bit Serial Architecture supplies reuse of a narrow datapath and warns against universal N-cycle or area laws. The XOR model makes its own fixed unit-delay assumptions explicit.
- The comparison is internal computation, not serial versus parallel transmission of an already computed word.
- Actual area, energy, clock rate, storage and data bandwidth can change the practical trade.
- Unlike a pipeline overlapping different words, this arrangement changes how many positions within one word execute together.
Source entries
Bit-Serial Architecture
Bit-Serial Architecture: Narrow hardware per lane vs. longer per-word latency supplies the conflict examined here.
Narrow hardware per lane vs. longer per-word latency
Reusing a small datapath can save logic and wiring, while a word operation needs multiple internal steps. Pipelining and clock choices change the result without dissolving the trade.