Pipeline Hazards, Forwarding, Register Timing, and the x0 Trap

RISC-V

Posted by Bruce Lee on 2024-01-26

About Me

Welcome to my blog! This is where I collect my observations and notes on programming and technology. The main subjects range from implementation details to broader ideas about programming.

Main Topics

  • Engineering Projects: Exploring implementation details and how technical systems work.
  • C/C++: Notes on language features and programming techniques.
  • The Programmer’s Perspective: Ideas about developing a career and a way of thinking as a programmer.

For more, visit the categories page.

Contact

If you have questions or would like to discuss something, please get in touch through the About page.

Thank you for reading and for your support. I hope these notes help you on your own technical journey!


Avoiding structural hazards

A structural hazard occurs when simultaneous instructions need a hardware resource that cannot serve both. In a five-stage pipeline with a single-ported unified memory, one instruction’s data access can conflict with another instruction’s fetch.

When deriving a pipeline from a single-cycle datapath, identify resource use and stage order. Duplicating resources or adding ports can remove conflicts. Instructions can omit some work while still flowing through a consistent stage structure.

Data hazards

Independent instructions pose no data dependency, but real programs frequently consume a preceding result:

1
2
1   add x19, x0, x1
2 sub x2, x19, x3

The subtraction needs the new x19 produced by the addition. Reading the register file before the addition writes back would obtain the old value. Waiting for writeback is a possible solution, but forwarding can often supply the result earlier.

Stage movement and stalls

Pipeline registers transfer state at clock boundaries, with the period constrained by stage timing. A consumer may need a value that a producer has computed but has not yet written to the register file. A direct path from the producer’s intermediate result to the consumer avoids waiting for architectural writeback.

Editorial clarification: stalling does not have to freeze the whole pipeline. A load-use interlock normally holds earlier stages, inserts a bubble, and lets older instructions advance. Freezing the producer along with the consumer would prevent the needed result from becoming available.

RAW dependencies resolved by forwarding

A read-after-write dependency arises when a later instruction reads the register an earlier instruction writes. If the value is available in time at an intermediate pipeline register, forwarding can resolve the dependency without a stall.

Load-use dependencies

A load result becomes available later than an ordinary ALU result. An immediately dependent instruction can need it before it exists at a usable forwarding point. In the usual five-stage example, a bubble plus forwarding resolves that case. Exact stall counts depend on stage placement and timing; forwarding cannot send a value backward in time.

When does the PC advance?

Textbooks often say that after fetching an instruction the PC advances to the next instruction. The physical timing depends on the datapath: a combinational next-PC path computes the sequential address, and the PC register captures the selected next value at its update edge. The original notes’ placement of this operation in a particular decode/execute sequence describes a proposed implementation rather than a universal rule.

Thinking about instruction flow

It helps to imagine instruction data and control moving through fixed hardware stages. A stall controls which pipeline registers advance and which retain their state; the stages themselves do not move. This viewpoint makes dependencies and bubbles easier to follow than treating each instruction as an isolated single-cycle operation.

Simple approaches to control hazards

A branch may not have resolved before later instructions enter the pipeline. Waiting until its outcome is known is simple but costly when branches are frequent.

One improvement moves branch comparison and target calculation earlier, potentially reducing the delay to one cycle in a suitable design. Dependencies may still require additional stalls.

Another historical technique, used by MIPS, is a branch delay slot: an instruction after the branch executes regardless of the branch outcome. A compiler or assembler attempts to place useful safe work there. Such a slot is architecturally relevant and is not necessarily invisible to assembly programmers. RISC-V does not use these delay slots.

Branch prediction instead selects a likely next address and fetches speculatively, with recovery when the prediction is wrong.

Pipeline registers

Registers between stages isolate successive combinational computations. Each stage simultaneously handles a different instruction, so every value needed later—including destination-register identifiers and control signals—must travel with that instruction. Connections suitable for a single-cycle datapath often need adjustment for this reason.

Register-file reads and writes in one cycle

A common model permits combinational reads and an enabled write at a clock edge. If a read and write address match, the design must define what the reader sees. Possible implementations arrange write/read timing within the cycle or bypass the incoming write data directly to the read port. The choice must satisfy the actual register-file timing contract.

Forward to the point of use

When an ALU operand depends on an older in-flight instruction, select the forwarded value just before the ALU input. This illustrates a broader design principle: resolve the dependency where the value is required.

The special x0 trap

Suppose an older instruction names x0 as its destination and a later instruction reads x0. Ordinary register-number matching might forward the older instruction’s nonzero intermediate ALU result. That would be wrong: the later instruction expects the architectural constant zero, and writes to x0 are discarded.

Forwarding conditions must therefore exclude destination register zero. A matching register number alone is insufficient; the producer must actually write a nonzero register.

Broadcast values, select at consumers

A simple organization exposes potential producer values to consumers, which select the appropriate input. Forwarding multiplexers sit near ALU inputs. This resembles a bus whose consumers take the fields or values they need, rather than a producer deciding every possible use.

Ordering the operand multiplexers

The second ALU input may be either a register value or an immediate. Forwarding adds a selector among the original register value, an EX/MEM result, and a MEM/WB result.

One might initially put immediate selection before forwarding and special-case the controls. A clearer structure resolves the current register operand first, then selects between that value and the immediate. Forwarding candidates all represent alternative versions of a register value; the immediate belongs to a different operand category. This ordering also makes the intended control semantics easier to reason about.

Simultaneous EX and MEM matches

1
2
3
add x1, x1, x2
add x1, x1, x3
add x1, x1, x4

The second instruction depends on the previous x1. The third can match both the immediately preceding writer in EX/MEM and an older writer in MEM/WB. The newer result is required by program order.

Priority for EX/MEM forwarding

1
2
3
4
5
6
7
8
9
10
//A input
if EX/MEM.RegWrite
and EX/MEM.RegisterRd == ID/EXE.RegisterRs1
and EX/MEM.RegisterRd != 0 //避免x0陷阱
ForwardA = 10
//B input
if EX/MEM.RegWrite
and EX/MEM.RegisterRd == ID/EXE.RegisterRs2
and EX/MEM.RegisterRd != 0 //避免x0陷阱
ForwardB = 10

The EX/MEM match takes priority when it is a valid available producer. MEM/WB forwarding is selected only when the newer matching producer does not supersede it:

1
2
3
4
5
6
7
8
9
10
11
12
//A input
if MEM/WB.RegWrite
and MEM/WB.RegisterRd == ID/EXRegisterRs1
and MEM/WB.RegisterRd != 0
and not (EX/MEM.RegWrite and (EX/MEM.RegisterRd == ID/EX.RegisterRs1) and (EX/MEM.RegisterRd != 0))
Forward = 01
//B input
if MEM/WB.RegWrite
and MEM/WB.RegisterRd == ID/EXRegisterRs2
and MEM/WB.RegisterRd != 0 //避免x0陷阱
and not (EX/MEM.RegWrite and (EX/MEM.RegisterRd == ID/EX.RegisterRd2) and (EX/MEM.RegisterRd != 0)) //避免与EX冒险冲突
Forward = 01

These conditions express age priority as well as register matching. Load availability and the separate stall logic must still be respected.


If you like this blog or find it useful for you, you are welcome to comment on it. You are also welcome to share this blog, so that more people can participate in it. All the images used in the blog are my original works or AI works, if you want to take it,don't hesitate. Thank you !