image-classification Β· vision Β· WASM Β· ~110 MB
Same job as the ViT demo β name what's in an image across ImageNet's 1,000 classes β and the same ViT-Base backbone (16Γ16 patches, a transformer, a [CLS] token). What's different is how it learned. Before ever seeing a label, BEiT was pre-trained the way BERT learns language: mask out a chunk of image patches and make the model predict them. That self-supervised masked-image-modeling step is the whole story. Drop an image, read the labels, then run it head-to-head against the supervised ViT on the same picture.
| Model | Xenova/beit-base-patch16-224 |
|---|---|
| Task / pipeline | image-classification |
| Architecture | BEiT-Base β Vision Transformer, ~86M params, 16Γ16 patches, 12 layers; self-supervised masked-image-modeling pretraining (predict masked patches' visual tokens), then fine-tuned on ImageNet |
| Labels | ImageNet-1k β 1,000 fixed classes |
| Quantization | q8 (8-bit) ONNX β verified non-degenerate |
| Backend | WebAssembly (runs anywhere; no WebGPU needed) |
| Download | ~110 MB, cached after first load |
| License | Apache-2.0 (Microsoft Research; Bao, Dong & Wei) |
| Web APIs | WebAssembly, Web Workers, FileReader β all Baseline widely available |
Pick or drop an image and hit Classify. The bars are the softmax over ImageNet's 1,000 classes β only the top few are shown, drawn from the full distribution.

5 labels
The whole point of this page: run the same image through two Vision Transformers with the same architecture but opposite pretraining. BEiT learned self-supervised first β masking and predicting image patches, no labels β then got fine-tuned on ImageNet. The supervised ViT learned straight from labelled images. Where they agree, the label is solid; where they disagree, you're watching two different training recipes pull the same network toward different answers. Loading ViT downloads ~88 MB the first time.
Self-supervised (MIM) Β· β
Supervised Β· β
BEiT's final layer produces one raw score β a logit β per class. A softmax turns those into probabilities that sum to 100%. The table shows the top classes with their raw logit next to the softmax probability. The two numbers below describe the shape of the whole distribution: the top-1 margin (how far the winner leads the runner-up) and the normalised entropy (0 = certain, 1 = spread thin).
| Class | logit | softmax |
|---|
Confidence (low β β high):
BERT learns language by hiding words and predicting them. BEiT does the same to pictures. First, a small learned tokenizer (a discrete VAE, borrowed from DALLΒ·E) turns every 16Γ16 image patch into one of a few thousand visual tokens β a vocabulary for images. Then pretraining is a fill-in-the-blank game: mask out ~40% of the patches, feed the rest to the transformer, and make it predict the visual token of each hidden patch. No class labels are involved β the supervision comes entirely from the image itself, which is what "self-supervised" means. To do that well, the model has to learn how shapes, textures, and objects hang together, so it builds rich general-purpose features. Only afterwards is a classification head bolted on and fine-tuned on ImageNet's labels. The payoff: self-supervised pretraining can use oceans of unlabelled images and often transfers better than purely supervised training. The head-to-head above lets you feel where that different schooling actually changes the answer β same ViT wiring, different upbringing.
BEiT was a landmark: it showed that BERT-style self-supervised pretraining works for vision, kicking off the masked-image-modeling wave (MAE, SimMIM, and more) that now underpins a lot of modern vision. For a product, the lesson is that you can pretrain strong image features on unlabelled data you already have, then fine-tune on a small labelled set.
Classify a sample and read the top-k β feel a self-supervised transformer name what's in an image.
On-device tagging: classify a batch of images into a searchable tag index.
Mask patches like BEiT's own pretraining and watch how gracefully it degrades.
Self-supervised vs supervised: BEiT and ViT agreement on the same image.
The convenience path is two lines with Transformers.js:
import { pipeline } from "@huggingface/transformers";
const classify = await pipeline(
"image-classification",
"Xenova/beit-base-patch16-224",
{ dtype: "q8" },
); // downloads + caches ~110 MB once
const output = await classify(imageURL, { top_k: 5 });
// β [{ label: "Egyptian cat", score: 0.55 }, { label: "tabby, tabby cat", score: 0.26 }, β¦]
This page goes one level deeper for "See inside": instead of the pipeline it runs the model's own
processor then model() in a single forward pass, so it can read the raw
logits tensor (all 1,000 classes) and compute the softmax, entropy, and margin
itself. Same weights, same result β it just keeps the intermediate numbers. Inference runs in a
Web Worker so the page never janks. Note the masked-image-modeling pretraining is BEiT's
history; the shipped model here is the fine-tuned classifier, so classification is a
plain forward pass β no masking at inference time.