About Me
Welcome to my blog! This is where I collect my observations and notes on programming and technology. The main subjects range from implementation details to broader ideas about programming.
Main Topics
- Engineering Projects: Exploring implementation details and how technical systems work.
- C/C++: Notes on language features and programming techniques.
- The Programmer’s Perspective: Ideas about developing a career and a way of thinking as a programmer.
For more, visit the categories page.
Contact
If you have questions or would like to discuss something, please get in touch through the About page.
Thank you for reading and for your support. I hope these notes help you on your own technical journey!
Rounding and intermediate precision
Floating-point formats represent only finitely many values, so intermediate results often need rounding. Binary32 has 23 stored fraction bits and 24 bits of precision for normal numbers; binary64 has 53 bits including its implicit leading bit. The original reference to fewer than 23 significant bits confuses stored fraction width with precision. These format details are described in the Numerical Computation Guide.
Guard and round bits retain extra information while an intermediate result is reduced to the destination format. Familiar rounding directions include toward positive infinity, toward negative infinity, toward zero, and nearest with ties to even. The original Java comparison refers to its usual nearest rounding behavior.
Why a sticky bit is needed
Shifts during arithmetic can discard significant information beyond the guard and round positions. A sticky bit records whether any discarded bit was nonzero. When the retained bits look exactly halfway between two representable values, this additional information distinguishes an exact tie from a value lying beyond it. Ties-to-even handles the exact tie; the sign and rounding mode determine the general direction, so a set sticky bit does not universally mean round toward positive infinity.
Fused multiply-add
Rounding errors can accumulate over many operations. Fused multiply-add computes an expression such as a + b * c with one final rounding, rather than rounding the product separately. It can improve both accuracy and performance in workloads where this pattern is common.
Subnormal performance
Subnormal numbers extend representation toward zero, but their processing cost varies by hardware. Some implementations use slower paths or software assistance. They are part of floating-point representation, not outside it; a performance claim must be tied to the processor in question.
AVX and subword parallelism
A 256-bit YMM register can hold four 64-bit floating-point values. In a suitable DGEMM implementation, operations on several elements of C = C + A * B can therefore run in parallel. SSE and AVX are forms of SIMD: a wide register holds several independent lanes.
An innermost vectorized loop can load four values from A and B, multiply corresponding lanes, and accumulate into C, advancing by four elements. This does not imply that every outer loop should advance by four or that every workload will achieve exactly four times the speed.
Many arithmetic traps arise from applying unlimited-precision mathematical intuition to finite-precision machine operations.
Shifts do not always replace signed division
Left shifts can implement multiplication by powers of two within the relevant representation constraints. Arithmetic right shift is not generally interchangeable with signed division that truncates toward zero. For example, shifting -5 right by two commonly gives -2, whereas -5 / 4 with truncation gives -1.
Editorial clarification: for truncating signed division, a nonzero remainder has the dividend’s sign, not the divisor’s.
Floating-point addition is not associative
(c + a) + b can differ from c + (a + b). A small addend can disappear when aligned with a much larger one because the result lacks enough precision to retain it. The precise boundary depends on the significand, rounding mode, and discarded bits.
For binary32, take c = 2^123, a = -2^123, and b = 1. The first grouping gives 1; the second can give 0. Cancellation makes the effect especially visible.
Parallel reductions and reproducibility
Parallel reductions regroup additions, and floating-point regrouping can change results. The number of workers and the reduction schedule may vary between runs, so a program can produce small numerical differences even when each execution follows the same high-level algorithm.
This does not make floating-point parallelism impossible. It means accuracy and reproducibility require numerical analysis, appropriate tolerances, and sometimes a deterministic reduction order. A useful development sequence is to establish a serial reference, then validate the accelerated implementation against that reference and the problem’s numerical requirements. Floating-point accuracy concerns practical programmers as well as mathematicians.
Beginning a processor design
An ISA influences implementation complexity. RISC-V’s regular fields make a small datapath approachable: a program counter supplies the fetch address, instruction bits select registers and immediates, and instruction type controls the required operations.
A useful first subset includes arithmetic and logic, a conditional branch such as beq, and loads/stores such as ld and sd. Instruction fetch and explicit data loads are distinct operations; loads are not themselves responsible for fetching every instruction.
Limitations of a simple single-cycle processor
Every instruction must complete between the relevant state-update edges, so the clock period must accommodate the longest combinational path. Short instructions then inherit the cost of long ones.
Edge-triggered state provides clear update boundaries. Register outputs can be read while a write occurs at a designated edge. A functional unit that must do two independent operations during one single-cycle instruction may need to be duplicated or given additional ports. If the required value has not settled by the update edge, the design fails its timing requirement.
Instruction and data access
Bits have no intrinsic instruction or data meaning; interpretation supplies that meaning. Nevertheless, a single-cycle load instruction can require instruction fetch and data-memory access in the same cycle. One single-ported memory cannot independently perform both accesses. Separate instruction/data paths, or suitable multiport storage, address that structural requirement.
Clock discipline
An open note in the original warns against casually routing clocks through ordinary logic gates. Uncontrolled gating can create timing problems. Practical clock gating needs purpose-designed clocking structures and timing analysis.
Main control and ALU control
The main controller decodes instruction classes and drives enables and selectors; a subordinate ALU controller determines the arithmetic operation. This decomposition can reduce control complexity. Correct use of don’t-care conditions requires understanding the entire datapath, including which shared units matter for each instruction.
A 32-bit instruction bus makes the instruction bits available to several consumers. The register file and immediate generator select the fields they need. There is no need for an imagined sequential instruction-field distribution process: the relevant wiring carries the bits concurrently.
Why pipelining helps
The single-cycle approach makes common short operations wait for the slowest instruction class. Pipelining overlaps stages from different instructions. In an ideal balanced four-stage pipeline, steady-state throughput can approach four times that of the corresponding non-overlapped execution. Fill/drain time, unequal stages, register overhead, and hazards reduce the actual speedup. The gain concerns throughput; it does not mean one instruction’s latency automatically becomes four times shorter.
If you like this blog or find it useful for you, you are welcome to comment on it. You are also welcome to share this blog, so that more people can participate in it. All the images used in the blog are my original works or AI works, if you want to take it,don't hesitate. Thank you !