A Simple 8086 Assembly Lab: Displaying a Number on Screen

An introductory assembly experiment

Posted by Bruce Lee on 2023-07-30

About Me

Welcome to my blog! This is where I collect my observations and notes on programming and technology. The main subjects range from implementation details to broader ideas about programming.

Main Topics

  • Engineering Projects: Exploring implementation details and how technical systems work.
  • C/C++: Notes on language features and programming techniques.
  • The Programmer’s Perspective: Ideas about developing a career and a way of thinking as a programmer.

For more, visit the categories page.

Contact

If you have questions or would like to discuss something, please get in touch through the About page.

Thank you for reading and for your support. I hope these notes help you on your own technical journey!


Begin with the high-level operation

In C, displaying an integer can be hidden behind a small function:

1
2
3
4
5
bool print_data(int a)
{
printf("%d", a);
return true;
}

A C++ template can work with different types that support stream insertion:

1
2
3
4
5
6
template <class T>
bool print_data(T a)
{
std::cout << a;
return true;
}

The task here is to expose some of the lower-level work behind displaying a number. We need to know both how characters reach the display and how an integer becomes a sequence of character codes.

Text video memory

The CPU communicates with display hardware through its interface and mapped storage. In the color text-mode environment used by this experiment, video memory begins at physical 0xB8000. Each screen cell uses two bytes: a character byte followed by an attribute byte.

An 80-column, 25-row screen uses 80 × 25 × 2 = 4000 visible bytes. Historical adapters can provide several pages, often at 4 KiB page intervals, with the active page selected by the display configuration. The source describes a 32 KiB region but gives B900:0000 as its end; that address is only 4 KiB beyond B800:0000, so the two claims are inconsistent. This exercise only needs the first text page.

We will display at row 8, column 3 using zero-based row and column parameters. The byte offset is row × 160 + column × 2.

A number is not its printed digits

The integer 12666 is stored as a binary value, whereas the visible text requires character codes for 1, 2, 6, 6, and 6: 31h, 32h, 36h, 36h, and 36h. The correct hexadecimal integer value is 317Ah; the binary transcription in the source is inconsistent with it.

ASCII digit codes are consecutive, so add 30h to a digit between zero and nine. First, however, the integer must be separated into decimal digits.

Repeated division by ten

Divide by ten, retain the remainder, and repeat with the quotient. Each remainder becomes one digit after adding 30h. The remainders arrive from least significant digit to most significant digit, so their order must be reversed when forming the string.

From a number to a string

For 12666, five divisions suffice, but the general loop stops when the quotient becomes zero. A conditional branch expresses that condition more directly than LOOP, which decrements a counter. CX can be changed by ordinary instructions too; its special role in LOOP does not make it otherwise immutable. A complete conversion also handles the input zero explicitly.

Division and quotient overflow

With an 8-bit divisor, unsigned DIV divides AX, returning quotient in AL and remainder in AH. With a 16-bit divisor, it divides DX:AX, returning quotient in AX and remainder in DX. The Chinese notes use BX in several places where the hardware instruction requires DX; a custom subroutine can choose a different interface, but DIV itself cannot.

For example, AX = 0100h divided by an 8-bit value one would need a quotient of 256, which cannot fit in AL. That operation raises divide error. To demonstrate the 8-bit case the operand is cl, not cx; operand width determines which division form is used.

This motivates a wider division subroutine that returns a quotient in two registers. Its implementation still needs to reject a zero divisor and keep each intermediate hardware division within range.

Divide the program into three routines

dtos converts a number into a string. The original proposed buffer is addressed by DS:SI, has ten zero-initialized bytes, and allows nine characters plus a terminator. That is enough for some inputs but not every unsigned 32-bit value, which may require ten digits plus the terminator. The eventual interface should state the input width and buffer capacity explicitly.

divdw performs a doubleword-by-word division. The proposed custom interface uses BX:AX for the input dividend, CX for the divisor, BX:AX for the quotient, and CX for the remainder. Internally it must adapt to the hardware’s DX:AX convention. Returning a wider quotient avoids the particular single-word quotient limit.

show_str copies a zero-terminated string into text video memory. Its parameters are DH for row 0–24, DL for column 0–79, CL for the color attribute, and DS:SI for the string address. Its visible screen output is the result.

The routines have distinct responsibilities. show_str can be developed independently with a sample string; divdw should be established before dtos, which repeatedly needs division. Clear parameter and return conventions keep their coupling manageable.

Testing show_str

Start with a fixed string and known location:

1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
assume cs:code
data segment
db 'Hello,i am Bruce Lee',0
data ends
code segment
start:
mov dh,8
mov dl,3
mov cl,2
mov ax,data
mov ds,ax
mov si,0
call show_str
mov ah,4ch
mov al,00h
int 21h
show_str:
; Implement the routine here.
ret
code ends
end start

The original skeleton’s moc is a typo for mov. It is also reasonable to write and debug the body inline first, then extract it into a callable routine once it works. A DOS executable using calls and pushes must have a valid stack supplied or initialized by its environment.

The recorded implementation

The original code is retained below, including its Chinese comments and punctuation. As the source warns, full-width semicolons are not valid MASM comment delimiters; replace them with ASCII ; or remove those annotations before assembling.

1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
25
26
27
28
29
30
31
32
33
34
35
36
37
38
39
40
41
42
43
44
45
46
47
48
49
50
51
52
53
54
55
56
57
58
59
assume cs:code
data segment
db 'Welcome to masm!',0
data ends

code segment
start: mov dh,8
mov dl,3
mov cl,2
mov ax,data
mov ds,ax
mov si,0 ;设置参数
call show_str

mov ah,4ch ;程序返回,使用4ch功能号
mov al,00h ;程序返回值为0
int 21h ;调用21h号中断例程

show_str:
push dx ;将用到的所有寄存器保留入栈
push cx
push ax
push es
push si
push di

push cx ;尽量用较少的寄存器,所以将cx入栈,在后面还有出栈
mov al,dh
mov ah,0
mov cl,160
mul cl ;ax=dh*160
mov di,ax ;di=dh*160
mov cl,2
mov al,dl
mov ah,0
mul cl ;ax=dl*2
add di,ax ;di=dh*160+dl*2
pop cx ;在这里出栈,这样,cx又存储的是原来应该存储的颜色值

mov ax,0b800h ;设置显存地址
mov es,ax
show: mov al,ds:[si] ;注意是al,不是ax,是8位的内存单元
cmp al,0
je show_ret ;相等就表示结束了,可以return了,否者一直进行跳转指令来覆写显存内容
mov es:[di],al
mov es:[di+1],cl ;每一个偶数内存单元后面的奇数内存单元就是该偶数单元的属性单元
inc si
add di,2 ;注意si,di的增长值不一样
jmp show
show_ret:
pop di ;对应出栈
pop si
pop es
pop ax
pop cx
pop dx
ret
code ends
end start

The routine preserves registers it uses, temporarily saves the color while calculating row × 160 + column × 2, sets ES to B800h, and copies each byte until the zero terminator. SI advances by one source byte, while DI advances by two video-memory bytes.

At an even cell offset, the first byte is the character; the next byte is its display attribute:

Text-mode character attributes

This installment implements the display routine. The wider division and numeric conversion routines remain the next parts of the original exercise; it does not yet contain a complete general-purpose numeric printer.


If you like this blog or find it useful for you, you are welcome to comment on it. You are also welcome to share this blog, so that more people can participate in it. All the images used in the blog are my original works or AI works, if you want to take it,don't hesitate. Thank you !