About Me
Welcome to my blog! This is where I collect my observations and notes on programming and technology. The main subjects range from implementation details to broader ideas about programming.
Main Topics
- Engineering Projects: Exploring implementation details and how technical systems work.
- C/C++: Notes on language features and programming techniques.
- The Programmer’s Perspective: Ideas about developing a career and a way of thinking as a programmer.
For more, visit the categories page.
Contact
If you have questions or would like to discuss something, please get in touch through the About page.
Thank you for reading and for your support. I hope these notes help you on your own technical journey!
Load-use dependencies at different stages
Consider a load followed by arithmetic or a store. An immediately following arithmetic instruction needs the load result in EX while the load is still in MEM. In the usual five-stage pipeline, a stall and subsequent forwarding are necessary.
Stores have two different uses of register data. Editorial correction: the Chinese notes interchange rs1 and rs2 in this discussion. A store’s rs1 is the address base, needed in EX; rs2 supplies the data written in MEM. A load-to-store-base dependency can therefore require a stall, while a load-to-store-data dependency can be handled without one if a forwarding path supplies the data at MEM. The value moves from the load’s MEM/WB register to the store-data input when required.
A store followed by a load does not produce the same register dependency. However, the original statement that it can never cause any hazard is too broad: memory address dependencies, buffering, and resource conflicts depend on the implementation.
Inserting a load-use bubble
A stage can perform a computation whose result is intentionally discarded. To make that bubble harmless, its control signals must prevent architectural register or memory writes. Intermediate pipeline registers may still carry irrelevant values as long as their accompanying write controls are disabled.
At the same time, younger instructions must be held. Detect the dependency while the consumer is in IF/ID and the load is in ID/EX. Check the load’s MemRead and destination register against the source registers actually used by the consumer. Then freeze the PC and IF/ID, and write zeroed control fields into ID/EX while allowing older instructions to advance.
The PC already identifies the next fetch when the consumer sits in IF/ID. Holding both prevents losing or duplicating that instruction. Write-enable controls provide the hold mechanism. The original reference to freezing ID/EX should be read in light of the required bubble insertion: ID/EX receives harmless controls, whereas IF/ID retains the consumer.
Stalls cost cycles and detection hardware, so recognizing the dependency early matters.
Normal and redirected PC updates
During ordinary fetch, PC + 4 is computed and selected as the next address in this base-instruction example. The fetched instruction and its address enter IF/ID. A later branch decision can select a target instead.
In the simple datapath considered here, branch information reaches EX/MEM before the redirect is acted on. The branch control must combine instruction type and the actual comparison result. This timing is specific to the design, not imposed by the ISA.
Reducing branch delay
If a five-stage pipeline waits until MEM to redirect, several younger instructions may already have entered. A design that simply waits can waste roughly three fetch opportunities. Two improvements are earlier branch resolution and prediction.
Move target calculation into ID
The branch’s PC is already in IF/ID, and the immediate generator already operates in decode. A dedicated adder can combine the PC and sign-extended branch offset without waiting for the EX ALU. This adds an adder but can reuse the existing immediate logic.
Compare operands in ID
An equality comparison can use bitwise XOR followed by an OR reduction. The comparison is simple, but operand availability still matters. Moving a consumer earlier can create new forwarding and stall requirements.
An ID-stage forwarding network must use results that are actually available at that time. The original suggestion of forwarding indiscriminately from ID/EX is insufficient: a producer still computing in EX may not supply a timely value for that cycle’s ID comparison. Pipeline timing determines the legal sources and any additional stall. Register and mux delays are small relative to some operations, but not literally zero.
Flush a wrong-path fetch
When a branch resolves in ID, one younger sequential instruction may already have been fetched. If that path is wrong, an IF.Flush operation replaces or invalidates the IF/ID instruction so it cannot produce side effects. It has not yet acquired decoded control fields, so clearing a valid bit or inserting an encoded no-op is a natural implementation.
Avoid resolving the same branch twice
Moving branch handling earlier requires disabling the old later redirect path for that branch. The instruction can continue through later stages with harmless controls; an implementation may inject a bubble into ID/EX. Older instructions must remain unaffected. The essential requirement is one architectural branch decision, not a second PC redirect when the branch reaches MEM.
Static and dynamic prediction
Simple static policies predict always not taken or always taken. Their quality depends on the code, and they remain useful in some designs. Dynamic prediction adapts to observed behavior.
A one-bit history entry predicts the next outcome from the previous outcome. A taken bit selects a predicted target; a clear bit selects the sequential path. Two-bit state, correlated predictors, and tournament schemes extend this idea by retaining richer history or choosing among predictors.
Predicting before decode
If recognizing a branch requires waiting for ID, fetch has already spent a cycle on the sequential path. A useful front end needs information earlier. A branch target buffer indexed by fetch address can recognize a previously encountered control-flow instruction and supply its prior target. Additional predecode or prediction structures may also help. Direction prediction alone is insufficient without a timely target address.
Reducing branches with conditional selection
ARMv8-style conditional selection can replace a short branch whose only purpose is choosing a register value. This can avoid the control-flow cost for a small amount of useful work. Whether the replacement is profitable depends on the instruction sequence and processor.
Recovering from a misprediction
When the branch resolves, the actual next address is compared with the prediction associated with that in-flight instruction. If they differ, younger wrong-path instructions are flushed and fetch is redirected. Prediction history is updated. The relevant comparison is with the prediction used for this branch, not merely whatever value a shared predictor entry happens to contain later.
If you like this blog or find it useful for you, you are welcome to comment on it. You are also welcome to share this blog, so that more people can participate in it. All the images used in the blog are my original works or AI works, if you want to take it,don't hesitate. Thank you !