x86 Programming Notes, Part 3: Execution, Descriptors, and Mode Switching

x86

Posted by Bruce Lee on 2023-10-11

About Me

Welcome to my blog! This is where I collect my observations and notes on programming and technology. The main subjects range from implementation details to broader ideas about programming.

Main Topics

  • Engineering Projects: Exploring implementation details and how technical systems work.
  • C/C++: Notes on language features and programming techniques.
  • The Programmer’s Perspective: Ideas about developing a career and a way of thinking as a programmer.

For more, visit the categories page.

Contact

If you have questions or would like to discuss something, please get in touch through the About page.

Thank you for reading and for your support. I hope these notes help you on your own technical journey!


Address spaces and execution

A 32-bit linear address space contains up to 4 GiB of addresses. How much is available to a particular task, and how it maps to physical memory, depends on the operating system.

Pipelining divides instruction processing into stages and overlaps work from different instructions. SRAM uses feedback-based storage cells, while DRAM stores charge requiring refresh. Caches exploit temporal and spatial locality.

Out-of-order execution is more than arbitrary fetch order or merely overlapping stages. A processor identifies dependencies and allows ready operations to execute before older blocked ones while preserving architectural behavior. Register renaming maps logical names such as EAX to different physical storage for different versions of a value, removing false name dependencies.

A modern pipeline can include fetch, decode, allocation and renaming, scheduling, execution, and retirement. Branch prediction and target buffers help fetch continue before a branch’s final outcome is known; mistakes require recovery.

Address and operand sizes

The 16-bit addressing forms permit certain base-plus-index combinations with an optional displacement. The same instruction bytes can be interpreted differently under different default operand and address sizes. NASM directives such as BITS 16 and BITS 32 tell the assembler which default encoding environment to target; they do not themselves switch the CPU’s operating mode.

For the classic 32-bit shift forms, a register count comes from CL, and the effective count is masked to five bits. A 32-bit push stores a 32-bit stack operand; a compact immediate may be sign-extended to that operand width. These details depend on the selected instruction form.

Segment descriptors

A conventional IA-32 code/data descriptor occupies eight bytes, with a 32-bit base, a 20-bit limit field, and access/attribute fields. The base need not be 16-byte aligned, though aligning data and code appropriately can help access performance. No single alignment rule guarantees maximum performance for every object or processor.

Privilege levels and descriptor attributes constrain access. Some instructions require the most privileged level. Segment checks consider more than a simple numeric minimum: selector privilege, descriptor privilege, segment type, and the operation all participate.

The present bit marks whether a descriptor’s segment is available. A not-present condition can raise an exception, allowing an OS to load or otherwise manage the segment. This is one historical virtual-memory technique; page-based virtual memory is a separate and widely used mechanism.

GDT and GDTR

The loader in this exercise creates a global descriptor table and loads GDTR with LGDT. In the relevant legacy form, the memory operand contains a 16-bit limit followed by a 32-bit base. The limit is the table size minus one. This is the six-byte pseudo-descriptor used by the instruction, not a segment descriptor itself.

The visible segment register is a selector plus hidden cached state in the processor. Treating that state as exactly a universally visible 80-bit register is an oversimplification; the architectural programming interface exposes the selector while hardware retains decoded descriptor information.

Switching to protected mode

The exercise must establish a GDT, arrange A20 behavior for the intended addresses, set the protected-mode enable bit in CR0, and transfer control appropriately to load the new code-segment state. Data and stack selectors then need suitable initialization.

Real mode and protected mode differ in how segment values are interpreted, how access checks work, and which default sizes the code segment establishes. A boot program can therefore contain both 16-bit setup code and 32-bit code, with assembler directives matching each part.

In ordinary real-mode segment loading, the base is formed from the segment value shifted left by four. On later processors, hidden segment state and A20 behavior make blanket statements about all real-mode physical addresses being limited to exactly twenty bits too simplistic.

Operand-size overrides

The 66h prefix changes the operand size relative to the code segment’s default for many legacy instruction forms: 16 to 32 or 32 to 16. An address-size override is a different prefix. A 386 or later CPU can use 32-bit registers in a 16-bit code environment with appropriate encodings. The original 8086 cannot.

Null selectors and aliases

The first GDT entry is conventionally reserved as the null descriptor. A null data selector can be loaded in permitted cases, but accessing memory through it causes a fault. This helps catch uninitialized segment references. Loading CS or an ordinary protected-mode stack selector has stricter requirements.

Several descriptors can refer to the same memory with different access roles. Such aliases can support shared regions or controlled access to code as data. They need careful permission design.

When a selector is loaded, the processor checks table bounds, descriptor type, and applicable privilege rules. An out-of-range or otherwise invalid selection can produce a general-protection exception. Matching the descriptor to the intended segment register is part of this process.

Expand-down stack segments

The source discusses an expand-down stack arrangement, in which offsets above a configured limit are valid up to the size-dependent maximum. Stack checks concern SP/ESP, not EIP; the original use of EIP is a typo.

With a 20-bit limit of FFFFEh and page granularity, the effective limit is FFFFEFFFh. In a suitable 32-bit expand-down segment, offsets FFFFF000h through FFFFFFFFh form a 4 KiB valid range. The base and modular address arithmetic then determine which linear region those offsets reach.

This is a specialized layout technique, not a general rule that every stack descriptor base must equal the top of an allocated buffer. In the referenced allocator example, adding allocation size to the low address is part of that particular construction. Both the segment limit and actual stack pointer must be initialized consistently.

Additional instruction notes

An address-size override can use 16-bit address forms in a 32-bit code environment. XCHG exchanges operands. IA-32 addressing supports a scaled index, illustrated by:

1
mov [es:0x0b80a0+ebx*2],ax

ROL and BSWAP can help rearrange fields while constructing descriptors. CMOVcc, available on appropriate later processors, can replace some small conditional branches, though performance depends on dependencies and workload.

MOVZX zero-extends a smaller source; MOVSX sign-extends it. Appropriate natural alignment can avoid split accesses, but four-byte alignment is not the only relevant boundary in a 32-bit system.

CMPS compares string elements and updates flags and indices according to its form and direction flag. It is a convenient instruction family, not an unconditional guarantee of faster comparison than every alternative implementation.

The assembler mode and prefix behavior is documented in the NASM BITS directive reference.


If you like this blog or find it useful for you, you are welcome to comment on it. You are also welcome to share this blog, so that more people can participate in it. All the images used in the blog are my original works or AI works, if you want to take it,don't hesitate. Thank you !