Booting Python + NumPy (Pyodide)…
Runs our hand-written cnn/conv.py with the trained weights: the real from-scratch CNN, classifying in your browser.
M5 · real NumPy CNN

A CNN, Written From Scratch

The target discriminator is a small convolutional net built in pure NumPy: the conv, pool and fully-connected layers, and the backprop that trained it, all hand-written with zero autograd. Draw a patch below and watch the real network classify it, live.

The target

Draw a 32×32 patch: pick a colour and shape. The designated target is the red balloon. This same patch feeds the panels alongside and below.
32×32 input (scaled up)
colour
shape

① What a convolution does

A conv layer slides a small kernel over the image. Pick one, applied by our conv2d, and watch the feature map it produces (edges, blur, …).
input
feature map
kernel

② The trained CNN decides

The full forward pass (conv → relu → pool → conv → relu → pool → fc → fc), run in pure NumPy with the trained weights. Left: what the first conv layer detects. Right: the verdict.
conv-1 feature maps (8 channels)
…

Things to notice

It learned shape, not just colour

Draw a red square: it scores near zero. Only the red balloon fires the TARGET verdict: the net needs the right colour and the right shape.

Why the negative set matters

Channels specialise

The eight conv-1 feature maps don't all do the same thing: some respond to edges, others to colour blobs. Stacked, they give later layers a richer description to classify.

What a feature map is

Pooling buys shift-tolerance

Max-pooling keeps only the strongest response in each block, so nudging the shape a few pixels barely changes the verdict: the net recognises the target wherever it sits.

How pooling helps

Trained without autograd

Every gradient in the backward pass is hand-derived, with no framework. A finite-difference check in the tests confirms the analytic gradients are correct.

The hand-written backprop