About Me
Welcome to my blog! This is where I collect my observations and notes on programming and technology. The main subjects range from implementation details to broader ideas about programming.
Main Topics
- Engineering Projects: Exploring implementation details and how technical systems work.
- C/C++: Notes on language features and programming techniques.
- The Programmer’s Perspective: Ideas about developing a career and a way of thinking as a programmer.
For more, visit the categories page.
Contact
If you have questions or would like to discuss something, please get in touch through the About page.
Thank you for reading and for your support. I hope these notes help you on your own technical journey!
Exceptions, interrupts, and direct trap entry
RISC-V uses trap as the general transfer to a handler. An exception is synchronous with instruction execution; an interrupt is an asynchronous event. Editorial clarification: the original notes use exception more broadly and mix machine- and supervisor-mode register names. A machine-mode discussion normally pairs mcause with mepc; a supervisor-mode discussion pairs scause with sepc.
In direct mode, traps routed to a particular privilege level enter a common handler address. Software reads the cause and dispatches to the appropriate routine.
A simplified pipeline trap implementation
The cause register records why the trap occurred, and the exception PC records the relevant instruction address so that software can decide whether and where to resume. These are architectural control/status registers rather than ordinary program globals.
The original implementation exercise restricts detection to EX. One could detect conditions in decode or other stages, and a complete processor must handle faults wherever they arise, including instruction fetch and memory access. The EX-only restriction simplifies the example; it is not a RISC-V requirement.
For an EX-stage fault, the design suppresses the faulting instruction’s side effects, invalidates younger instructions, records cause and PC, and redirects execution. IF/ID, ID/EX, and the faulting result entering EX/MEM may need clearing or invalidation, producing bubbles. Older instructions must complete consistently with precise trap semantics.
Hardware records the architectural trap state and redirects control. Handler software normally saves the general registers it needs, interprets the cause, and either resumes, repairs the condition, or terminates the affected program. Saving every general register is not an automatic consequence of entering a RISC-V trap.
Keeping the faulting PC precise
By the time an instruction reaches EX, the fetch PC has advanced, potentially by several instructions. Saving the current fetch PC would identify the wrong instruction. Carry the instruction’s own PC through the pipeline and use that value when recording its trap.
The original PC + 4 example illustrates the mismatch but is not a universal offset. Precise traps also require older instructions to have completed and younger instructions to have made no architectural changes. Deeper pipelines make this bookkeeping more important, and restartable faults are especially useful for virtual memory.
Two ways to expose instruction parallelism
A deeper pipeline overlaps more stages from different instructions. Multiple issue allows more than one instruction to enter execution in a cycle when resources and dependencies permit. The resulting hardware may resemble several logical pipelines, though practical implementations share and combine resources.
Compiler scheduling and hardware scheduling are complementary approaches. A compiler can reorder instructions and unroll loops to expose independent work. Dynamic scheduling can make decisions using runtime readiness information. Neither approach can discard true dependencies.
Static issue packets
In a statically scheduled multiple-issue design, the compiler groups compatible operations. A two-slot packet may contain a no-op when no suitable second operation is available. Packet-to-packet hazards still require the mechanisms defined by that architecture, such as forwarding or stalls.
Loop unrolling increases the scheduling region and can reveal work that would otherwise sit behind a stalled instruction. Treating a packet as one instruction containing several predefined operations resembles VLIW. This is a general architecture example, not a property mandated by standard RISC-V.
Where out-of-order parallelism occurs
A common dynamically scheduled processor fetches and decodes in program order, identifies operands and dependencies, and places instructions into scheduling structures. Ready operations can then run concurrently in different functional units. Those whose operands or resources are unavailable wait in reservation stations or analogous queues.
This is an implementation pattern rather than a universal requirement that every processor decode exactly one instruction at a time. Several instructions may be decoded together while preserving their order for bookkeeping.
Why a reorder buffer matters
Execution may finish out of order, but architectural retirement can remain ordered. A reorder buffer tracks this distinction and helps preserve precise exceptions. Results may be forwarded or written to speculative physical storage before retirement; writeback and architectural commitment are not necessarily the same event.
The original notes credit RISC-V’s regular instruction structure for simplifying this organization. The idea is useful, but the ISA does not mandate a particular five-stage writeback schedule or a reorder-buffer implementation.
Static and dynamic multiple issue
A simple single-issue pipeline has one instruction entering each stage per cycle. Multiple issue increases ports and execution resources. Static issue relies on compiler-formed groups; dynamic issue chooses combinations in hardware.
Issue width and execution ordering are separate dimensions. A processor can issue several instructions while preserving in-order execution, or combine multiple issue with dynamic out-of-order scheduling. Describing these dimensions separately avoids treating every wide pipeline as the same architecture.
False dependencies from name reuse
Using as few variables or registers as possible does not necessarily improve performance. Reusing a name for unrelated values can create anti-dependencies or output dependencies that force ordering even when no value truly flows between operations.
Register renaming assigns different storage to different versions of a value, removing these name dependencies while preserving true read-after-write dependencies. Compilers can use additional architectural registers; dynamic processors can also rename into physical registers. Loop unrolling provides a larger region in which such opportunities become visible.
Memory aliasing
Different pointers may refer to the same memory location. The compiler cannot freely reorder accesses unless it can establish that doing so preserves behavior. Unlike simple register names, addresses may depend on runtime values. Register renaming alone does not resolve this uncertainty. Alias analysis and appropriate runtime mechanisms address a distinct dependency problem.
Performance and energy
Pipelining improves resource utilization, and speculation tries to uncover more instruction-level parallelism. Additional transistors enabled increasingly elaborate designs, but speculation, deep pipelines, and recovery machinery also consume energy. Power limits encouraged multicore architectures and renewed interest in efficiency.
The original notes argue for simpler cores under an energy budget. That is a design hypothesis, not a universal ranking: workload parallelism, latency requirements, process technology, utilization, and memory behavior all affect whether several simpler cores or a more capable core provide better results.
Pipeline design remains difficult
Pipeline depth, dependency detection, forwarding, prediction, and control require careful analysis. The claim that pipelining is simple is misleading. The claim that it is independent of semiconductor technology is also misleading: available delay, area, wiring, and power shape feasible structures. Historical techniques such as architectural delay slots can become awkward as implementations evolve.
ISA decisions also influence implementation difficulty. Variable instruction lengths and complicated addressing modes can complicate decoding and dependency tracking. Register-updating addressing modes add destinations; instructions with several memory accesses complicate control. x86 implementations use internal micro-operations and sophisticated front ends to manage these issues, but that does not make their cost disappear.
If you like this blog or find it useful for you, you are welcome to comment on it. You are also welcome to share this blog, so that more people can participate in it. All the images used in the blog are my original works or AI works, if you want to take it,don't hesitate. Thank you !