About Me
Welcome to my blog! This is where I collect my observations and notes on programming and technology. The main subjects range from implementation details to broader ideas about programming.
Main Topics
- Engineering Projects: Exploring implementation details and how technical systems work.
- C/C++: Notes on language features and programming techniques.
- The Programmer’s Perspective: Ideas about developing a career and a way of thinking as a programmer.
For more, visit the categories page.
Contact
If you have questions or would like to discuss something, please get in touch through the About page.
Thank you for reading and for your support. I hope these notes help you on your own technical journey!
One unsigned comparison for a bounds check
Suppose the valid array bound y is a nonnegative signed value and we need 0 <= x < y. An unsigned comparison can reject both a negative index and an index at or above the bound:
1 | bgeu x, y, error //伪代码 |
A bit pattern with its most significant bit set represents a negative signed integer but a large unsigned integer. Under the stated bound assumption, that interpretation combines the two checks.
Supporting switch
A switch can become a chain of comparisons, but a jump table can be more efficient for suitable cases. RISC-V’s jalr supports the indirect jump needed to select an address from such a table.
Procedure calls
The standard integer calling convention uses x10–x17 (a0–a7) for arguments, with return values primarily in a0 and a1. Register x1 is the return address. A direct call can use jal x1, procedureAddress; jalr x0, 0(x1) returns without retaining another link.
Unconditional control transfer
A condition that is always true can express an unconditional branch:
1 | beq x0, x0, Lable //其他变体 |
A jump-and-link instruction can instead discard its link by naming x0:
1 | jal x0, Lable //unconditionally |
1 | jalr x0, 0(x1) //寄存器跳转的样例 |
The dedicated unconditional jump form is the normal choice when its range fits.
How might x0 be implemented?
Possible implementations include returning a constant zero for reads, suppressing writes whose destination is zero, or representing the zero entry without ordinary storage. These are implementation choices serving the same architectural rule. Forwarding logic must respect that rule too.
Historical notes: program counter might more literally be called instruction-address register; x2 is the stack pointer; the conventional stack grows toward lower addresses.
Temporary and saved registers
x5–x7 and x28–x31 are temporaries. x8–x9 and x18–x27 are callee-saved registers. Using temporaries for short-lived values can reduce save/restore work, but values live across calls need appropriate preservation.
The caller preserves caller-saved values it still needs, including argument or temporary registers. The callee preserves any callee-saved registers it modifies. A non-leaf procedure must also preserve its own return address when another call would overwrite it. When writing a reusable subroutine, reason from the callee’s obligations and maintain the stack convention.
Global and frame pointers
Register x3 is conventionally gp, supporting access patterns for global or small data in an ABI-dependent way. Static storage duration does not itself imply global visibility; a block-scope static variable is a counterexample.
Register x8 can serve as frame pointer fp/s0. A frame can contain saved registers, local variables, structures, and outgoing arguments. A stable frame pointer can simplify references when the stack pointer changes.
The original notes link omission of the frame pointer to absence of local variables. In practice, frame-pointer use depends on compiler options, ABI rules, and function requirements; functions with locals can also omit it. Arguments beyond the available argument registers are passed on the stack according to the calling convention.
Stack, static storage, and heap
A simplified address-space diagram places code and static data below a dynamically allocated heap, with a downward-growing stack above. Real layouts vary. Static variables belong to static storage, while malloc obtains dynamic storage. Arrays and linked structures may reside in either region depending on their lifetime and allocation; they are not inherently heap objects.
Dangling pointers and memory leaks are opposite lifetime-management failures: using storage after its lifetime, and losing the ability to reclaim storage that remains allocated.
Recursion and iteration
Iteration repeatedly processes data through a loop. A suitable tail-recursive call can be transformed into iteration, avoiding growth of recursive call frames. Whether a compiler performs that transformation depends on the language and compilation conditions.
Constructing a large immediate
lui supplies an upper immediate with twelve low zero bits. In RV64 its 32-bit result is sign-extended. An addi can combine a signed 12-bit low part with that upper value.
The low part is not simply an unsigned bit-field insertion. If its top bit is set, sign extension makes it negative and the chosen upper part must compensate.
The original toy example uses an eight-bit register and four-bit halves to construct 1110 1010. The intended sequence is:
1 | lui rd, 1110 |
At this point the register contains 1110 0000.
1 | addi rd, rd, 1010 |
Treating the low part as unsigned would produce 1110 0000 + 0000 1010 = 1110 1010. But the actual signed-immediate interpretation is:
1 | lui rd, 1110 |
1 | addi rd, rd, 1010 |
The sum is 1110 0000 + 1111 1010 = 1101 1010 modulo eight bits, one upper unit too low. Incrementing the upper immediate compensates. In the real 12-bit case, a negative low immediate requires the corresponding 2^12 adjustment. Arbitrary 64-bit constants can require additional instructions; this example explains the upper/lower split.
Branch and jump displacement units
The base instruction examples use 32-bit instructions, but RISC-V also supports 16-bit compressed encodings. Branch and direct-jump offsets are therefore encoded in multiples of two bytes.
B-type, historically called SB-type, carries twelve explicit displacement bits with an implicit low zero, covering roughly ±4 KiB. J-type, historically UJ-type, carries twenty explicit bits plus the low zero, covering roughly ±1 MiB. The upper positive endpoint is one two-byte step below the corresponding power of two. These offsets are relative to the address of the branch or jump instruction.
If you like this blog or find it useful for you, you are welcome to comment on it. You are also welcome to share this blog, so that more people can participate in it. All the images used in the blog are my original works or AI works, if you want to take it,don't hesitate. Thank you !