A Vision Transformer, running in your browser.
Pick a galaxy cutout. A ViT-Small exported to ONNX runs entirely in your browser (no server) and classifies its morphology, while the attention map shows where the transformer looked. The metrics below are honest, cross-validated, and reported against the majority-class baseline. loading model…
What you're seeing. Real JWST/NIRCam cutouts (F200W+F356W+F444W of GOODS-S, COSMOS and UDS galaxies), pulled live from the DAWN JWST Archive cutout service. The labels are real human visual classifications: the canonical Galaxy Zoo: CANDELS split into featured/disk, smooth and merger (majority volunteer vote, ≥20 classifiers). Galaxy Zoo judged HST imaging; the ViT here is shown the JWST view of the same galaxies and learns morphology straight from the pixels. Next milestone: a self-supervised atlas of JWST galaxies you can fly through, and the hunt for the Little Red Dots.
Tick attention overlay: the transformer concentrates on the source and ignores the empty sky around it.
how that's drawnIt chains the ViT's self-attention across all layers back to the image, showing which patches the network actually leaned on.
Nobody told it where the galaxy was; it learned to weight the structured pixels.
In the confusion matrix, smooth and merger trip each other up far more than disks do.
why they blurA relaxed elliptical and a settled merger remnant can look identical in a single image; even human classifiers disagree.
The model reports the confusion rather than hiding it; that's the honest 81% vs a 33% baseline.
A clean disk scores ~99%; a faint or messy cutout splits its vote across the three classes.
what the bars meanThe bar heights are the softmax probabilities. A near-even split is the model honestly saying this one is genuinely ambiguous, not guessing.