A hand-written x86 operating system,
booting in your browser.

Boot sector to isolated userspace processes: every layer in C and x86 assembly, no Linux and no libraries underneath. It runs live below via v86, an x86 PC emulated in WebAssembly right in this tab; the exact same disk image boots in QEMU on a desktop.

kernel · 32-bit x86 · v86

Three worker processes run at once, pre-empted by the timer. Each reads the same virtual address (0xB0000000) but it maps to a different physical frame per process: proof they're truly isolated. Below them, a fourth process is a shell. Click the screen and type help, then try poke to watch the CPU stop it writing to kernel memory. The clock top-right ticks on the timer IRQ.

Where this code actually runs

There's no operating system underneath it: this is the operating system. The source is compiled and stitched into one raw disk image, and that single image is what boots, whether on a real desktop or right here in your browser. Same bytes, both places.

1 · source
kernel/*.c + *.asm, hand-written C and x86 assembly, no libraries.
→
2 · build
make: a cross-compiler + NASM + a linker script flatten it all to raw machine code.
→
3 · disk image
os-image.bin: boot sector + kernel, one 1 MiB file.
→
on a desktop
make run boots it in QEMU.
=
in this tab
v86 boots the identical image, live above.

How it works: the actual code

Every snippet below is really in the kernel booting above. Follow the machine come up, from the first byte the BIOS runs all the way to isolated processes.

Boot
01 · boot

512 bytes and a magic number

The BIOS loads the first disk sector to 0x7C00 and jumps there in 16-bit real mode, but only if the sector ends in the signature 0x55 0xAA. That tiny stub is all we get to bootstrap everything else.

boot/boot.asm
[bits 16]
[org 0x7C00]            ; the BIOS drops us here, still a 1978-era 8086

    mov ax, 0x0003     ; ask the video BIOS for 80x25 colour text mode
    int 0x10
    call load_kernel   ; read the C kernel off the disk into memory
    call enable_a20    ; unlock addressing above 1 MiB
    call switch_to_pm  ; leave real mode behind; never returns

times 510 - ($ - $$) db 0   ; pad to 510 bytes...
dw 0xAA55                    ; ...then the signature the BIOS checks for
02 · protected mode

Escaping real mode by hand

We load a flat GDT, set the protection-enable bit, then far-jump, which is what actually reloads CS and drops the CPU into 32-bit protected mode.

boot/boot.asm
switch_to_pm:
    cli
    lgdt [gdt_descriptor]   ; describe our flat segment layout to the CPU
    mov eax, cr0
    or  eax, 0x1            ; CR0.PE = 1: turn on protected mode
    mov cr0, eax
    jmp CODE_SEG:init_pm    ; far jump flushes the pipeline, loads a 32-bit CS
03 · output

The screen is just memory

No driver stack, no BIOS call. VGA text mode is memory-mapped at 0xB8000: each cell is a character byte plus a colour attribute byte. Write a 16-bit word and a character appears.

kernel/vga.c
#define VGA_MEMORY ((volatile uint16_t *)0xB8000)

// low byte = character, high byte = attribute (background << 4 | foreground)
static inline uint16_t vga_entry(char c, uint8_t attr) {
    return (uint16_t)c | ((uint16_t)attr << 8);
}
VGA_MEMORY[row * 80 + col] = vga_entry(c, color);
Interrupts
04 · interrupts

A jump table for the CPU

The IDT tells the CPU where to go when interrupt N fires. Each gate stores a handler address split across two halves; 0x8E marks it a present, ring-0, 32-bit interrupt gate.

kernel/idt.c
void idt_set_gate(uint8_t num, uint32_t base, uint16_t sel, uint8_t flags) {
    idt[num].base_low  = base & 0xFFFF;
    idt[num].base_high = (base >> 16) & 0xFFFF;
    idt[num].selector  = sel;      // our kernel code segment, 0x08
    idt[num].flags     = flags;    // 0x8E = present, ring 0, interrupt gate
}
05 · interrupts

Defusing a legacy landmine

The 8259 PIC delivers hardware IRQs on vectors that, by default, collide with the CPU's own exception vectors, so a timer tick would look like a double fault. We remap it so IRQs land safely on vectors 32–47.

kernel/isr.c
void pic_remap(void) {
    outb(PIC1_CMD, 0x11);  outb(PIC2_CMD, 0x11);   // begin init
    outb(PIC1_DATA, 0x20); outb(PIC2_DATA, 0x28);  // master->32, slave->40
    outb(PIC1_DATA, 0x04); outb(PIC2_DATA, 0x02);  // wire the cascade
    outb(PIC1_DATA, 0x01); outb(PIC2_DATA, 0x01);  // 8086 mode
    outb(PIC1_DATA, 0x00); outb(PIC2_DATA, 0x00);  // unmask every line
}
06 · keyboard

Every keystroke is an interrupt

Press a key and the controller raises IRQ1. Our handler reads the raw scancode from port 0x60, ignores key-release codes, translates to ASCII and buffers it for a reader.

kernel/keyboard.c
static void keyboard_callback(registers_t *r) {
    uint8_t scancode = inb(0x60);       // read the key from the controller
    if (scancode & 0x80) return;        // top bit set = key released; ignore
    char c = scancode_ascii[scancode];  // scancode -> ASCII
    if (c) buffer_push(c);
}
Memory
07 · physical memory

Handing out RAM, one frame at a time

Physical memory is carved into 4 KiB frames, one bit each in a bitmap: 1 used, 0 free. Everything above (page tables, the heap) ultimately draws its memory from here.

kernel/pmm.c
uint32_t pmm_alloc_frame(void) {
    for (uint32_t f = 0; f < TOTAL_FRAMES; f++) {
        if (!is_used(f)) {
            set_used(f);
            return f * FRAME_SIZE;   // the physical address of a free 4 KiB frame
        }
    }
    return 0;   // out of physical memory
}
08 · virtual memory

Turning on the illusion

map_page points a virtual page at a physical frame through the two-level page tables; invlpg drops the stale cached translation. Then we load CR3 and set CR0.PG, and every address the CPU touches becomes virtual.

kernel/paging.c
void map_page(uint32_t virt, uint32_t phys, uint32_t flags) {
    uint32_t pd = virt >> 22, pt = (virt >> 12) & 0x3FF;
    // ...allocate a page table for this region if none exists yet...
    page_table[pt] = (phys & ~0xFFF) | flags | PAGE_PRESENT;
    invlpg(virt);                 // forget any cached translation for this page
}

asm("mov %0, %%cr3" :: "r"(page_directory));   // point the MMU at our tables
cr0 |= 0x80000000;                             // CR0.PG: paging is now on
asm("mov %0, %%cr0" :: "r"(cr0));
09 · heap

kmalloc, from nothing

A first-fit walk over a linked list of blocks: find a free one big enough, split off the surplus, hand back the payload. kfree marks it free and coalesces with its neighbour so the heap doesn't fragment into confetti.

kernel/heap.c
void *kmalloc(size_t size) {
    for (block_t *b = head; b; b = b->next) {
        if (b->free && b->size >= size) {
            split(b, size);       // carve off any surplus as a new free block
            b->free = 0;
            return (uint8_t *)b + sizeof(block_t);
        }
    }
    return NULL;   // out of heap
}
Multitasking
10 · multitasking

What a context switch actually is

Startlingly small. Save the outgoing task's registers, stash its stack pointer, load the next task's, restore its registers. The final ret resumes it exactly where it left off. Driven by the timer IRQ, this is pre-emptive multitasking.

kernel/switch.asm
context_switch:            ; (uint32_t *save_old_esp, uint32_t new_esp)
    mov eax, [esp + 4]     ; where to save the outgoing esp
    mov edx, [esp + 8]     ; the incoming task's saved esp
    push ebp               ; preserve the callee-saved registers
    push ebx
    push esi
    push edi
    mov [eax], esp         ; freeze the outgoing task
    mov esp, edx           ; adopt the incoming task's stack
    pop edi                ; restore its registers
    pop esi
    pop ebx
    pop ebp
    ret                    ; ...and resume it where it last left off
Userspace
11 · userspace

The one guarded doorway

User programs run in ring 3, unprivileged. Their only channel into the kernel is int 0x80, a gate that traps into the kernel, which does the privileged work and hands back a result. That trap boundary is what a system call actually is.

kernel/shell.c
// running in ring 3: the ONLY way to reach the kernel
static char s_read(void) {
    uint32_t c;
    asm volatile("int $0x80" : "=a"(c) : "a"(SYS_READ) : "memory");
    return (char)c;               // the kernel read a key on our behalf
}
RING 3 · user processes (unprivileged) RING 0 the kernel ↩ int 0x80 syscall

Ring 3 can't touch hardware or kernel memory; the only way in is the syscall gate.

12 · protection

Catching a memory violation

Kernel pages are mapped supervisor-only; only a process's own code and data carry the user bit. If ring-3 code reaches for memory it doesn't own, the CPU raises a page fault with the offending address in CR2. The handler reads it and reports which process misbehaved. If it was the shell, the kernel restarts the shell on a fresh stack and the other processes never notice.

kernel/fault.c
static void page_fault(registers_t *r) {
    uint32_t cr2;                              // the address that faulted
    asm volatile("mov %%cr2, %0" : "=r"(cr2));

    if ((r->cs & 3) == 3 && sched_current_id() == SHELL_PID) {
        vga_puts("[page fault] the CPU blocked process 3 at ");
        r->eip = (uint32_t)shell_loop;        // resume the shell's command loop
        r->useresp = USER_STACK_TOP;          // on a fresh user stack
        return;                                // iret straight back to ring 3
    }
    // any other fault: report cr2 / eip / error code, then halt.
}
13 · isolation

What a process really is

Each process gets its own page directory: its own map from virtual to physical memory. On every context switch the scheduler loads that process's CR3 (and its kernel stack via the TSS). That's why the three worker processes above read the identical virtual address 0xB0000000 yet each sees a different physical frame.

kernel/sched.c
void schedule(void) {                    // called from the timer IRQ
    task_t *prev = current;
    current = current->next;              // round-robin to the next process
    if (prev == current) return;

    tss_set_kernel_stack(current->kstack_top);  // where its ring-3 traps land
    paging_switch(current->page_dir);           // its address space (load CR3)
    context_switch(&prev->esp, current->esp);   // its saved kernel context
}
same virtual address 0xB0000000 → a different physical frame per process physical RAM proc 0 VA 0xB0000000 phys 0x6E000 proc 1 VA 0xB0000000 phys 0x76000 proc 2 VA 0xB0000000 phys 0x7E000

Hover a process to trace its mapping: three separate address spaces, one CR3 each.

● KERNEL · live