About Me
Welcome to my blog! This is where I collect my observations and notes on programming and technology. The main subjects range from implementation details to broader ideas about programming.
Main Topics
- Engineering Projects: Exploring implementation details and how technical systems work.
- C/C++: Notes on language features and programming techniques.
- The Programmer’s Perspective: Ideas about developing a career and a way of thinking as a programmer.
For more, visit the categories page.
Contact
If you have questions or would like to discuss something, please get in touch through the About page.
Thank you for reading and for your support. I hope these notes help you on your own technical journey!
A conditional target beyond branch range
When a conditional destination lies outside the B-type range, assembly expansion can invert the condition and branch over a longer-range unconditional jump. This handles an uncommon case in software instead of expanding every conditional-branch encoding. A direct J-type jump has a wider displacement range, roughly ±1 MiB. Still more distant targets may require a register-based sequence.
Recovering assembly from machine code
Disassembly converts machine-code bytes into assembly representations. It is useful when inspecting a core dump and reconstructing what the processor was executing. Correct interpretation still depends on the architecture, mode, and instruction boundaries.
Synchronization primitives
RISC-V’s atomic extension provides load-reserved/store-conditional pairs, such as lr.d and sc.d, for constructing atomic updates:
1 | again: lr.d x10, (x20) //load-reserved |
A lock can give one participating thread exclusive access to a protected region. The lock variable itself needs an atomic acquisition protocol:
1 | addi x12, x0, 1 //copy locked value |
The original release example is:
1 | sd x0, 0(x20) /free lock by writing 0 |
Clarification: another processor merely reading a location does not in general make every reservation fail. Reservation validity follows the architecture’s rules, and store-conditional can fail for several reasons. Real locks also require appropriate acquire/release ordering, not just an atomic test-and-update. These snippets are learning examples, not complete portable synchronization implementations.
Indirection as a design tool
A familiar saying in computer science is that another level of indirection can solve many problems. It is a useful design prompt, though not a literal theorem covering every problem or the cost of additional indirection.
Array indices and pointers
A straightforward indexed loop may repeatedly calculate base + index × element size. A pointer loop can maintain the current address directly. This can reduce arithmetic in a naive translation.
Modern optimizing compilers often transform indexed loops into equivalent address induction, so writing a pointer does not by itself guarantee better performance. Choose a clear representation and inspect or measure the compiled result when performance matters.
Common compiler optimizations
Procedure inlining removes some call, save, and return overhead by placing the body at the call site. Repeated inlining can enlarge code and harm instruction-cache behavior, so it is a tradeoff.
Register allocation keeps useful values in registers and spills when necessary. The selection among temporary registers, saved registers, and memory depends on lifetimes and calling costs.
Strength reduction replaces expensive operations with cheaper equivalents when semantics permit, such as some multiplication-by-constant patterns becoming shifts and adds. Induction-variable elimination can replace repeated array address calculations with a running pointer. Constant folding evaluates suitable expressions during compilation.
A description of object-oriented programming
One way to describe the style is organizing a program around objects and data together with their operations, rather than around a sequence of actions alone. This is a perspective on program organization, not a prohibition on logic or functions.
RISC-V observations and qualifications
The base integer architecture follows a load/store organization for ordinary memory transfers. Its simple scalar subset does not offer an arbitrary multi-register load/store instruction. Extensions, including atomics and vectors, broaden the available operations.
The original notes say register comparisons exist only inside branches. That needs qualification: slt, sltu, and their immediate forms produce comparison results without branching.
x86 and the cost of compatibility
The original discussion argues that binary compatibility was commercially valuable but accumulated architectural complexity. Early x86 registers had many specialized roles, and modern instructions retain some implicit operand constraints. Variable instruction lengths, numerous addressing modes, prefixes, and a large encoding space make decoding more involved than a small regular instruction subset.
Simple RISC-V sequences can express many operations packaged into one complex x86 instruction. A more powerful instruction does not automatically execute faster than those sequences, because implementation techniques and algorithms continue to change. Conversely, claims that particular x86 string operations are always slower are too broad and can become outdated; performance depends on the processor and operation.
The source also credits x86’s early market entry and large ecosystem with funding engineering work to manage this complexity. Its prediction that x86 would lack future competitiveness is the author’s historical opinion, not an established result. Likewise, the cited total of 184 instructions plus 13 system instructions describes the material being studied, not a permanent count for an expanding RISC-V ecosystem.
RISC-V also has qualifications to the simple comparison: x0 is fixed at zero, register conventions constrain software, and compressed encodings mean instruction length is not universally 32 bits. Architectural regularity remains useful without treating either family as free of tradeoffs.
Assembly does not guarantee maximum performance
The original notes describe optimizing compilers as an early form of powerful machine intelligence, particularly in register allocation and instruction selection. Whether that merits the label artificial intelligence is a conceptual question left for reflection.
The practical point is that handwritten assembly takes more time to write and debug, is harder to maintain and port, and does not automatically outperform optimized higher-level code. Specialized kernels can still benefit from expert assembly when measurements justify it.
Remaining observations
Binary compatibility does not mean a successful ISA never evolves. Byte addressing means adjacent words or doublewords differ by their byte widths, not by one. A pointer can refer to an existing object independently of the statement that defined it, but must respect scope and lifetime. Many ISA operations support familiar language constructs. Multimedia extensions often include saturating arithmetic, whose overflow behavior differs from ordinary modular integer arithmetic.
For reservation and ordering rules, see the RISC-V atomic extension.
If you like this blog or find it useful for you, you are welcome to comment on it. You are also welcome to share this blog, so that more people can participate in it. All the images used in the blog are my original works or AI works, if you want to take it,don't hesitate. Thank you !