← All models

image-classification Β· vision Β· WASM Β· ~110 MB

BEiT β€” BERT, but for images

Same job as the ViT demo β€” name what's in an image across ImageNet's 1,000 classes β€” and the same ViT-Base backbone (16Γ—16 patches, a transformer, a [CLS] token). What's different is how it learned. Before ever seeing a label, BEiT was pre-trained the way BERT learns language: mask out a chunk of image patches and make the model predict them. That self-supervised masked-image-modeling step is the whole story. Drop an image, read the labels, then run it head-to-head against the supervised ViT on the same picture.

At a glance

ModelXenova/beit-base-patch16-224
Task / pipelineimage-classification
ArchitectureBEiT-Base β€” Vision Transformer, ~86M params, 16Γ—16 patches, 12 layers; self-supervised masked-image-modeling pretraining (predict masked patches' visual tokens), then fine-tuned on ImageNet
LabelsImageNet-1k β€” 1,000 fixed classes
Quantizationq8 (8-bit) ONNX β€” verified non-degenerate
BackendWebAssembly (runs anywhere; no WebGPU needed)
Download~110 MB, cached after first load
LicenseApache-2.0 (Microsoft Research; Bao, Dong & Wei)
Web APIsWebAssembly, Web Workers, FileReader β€” all Baseline widely available

Run it

Pick or drop an image and hit Classify. The bars are the softmax over ImageNet's 1,000 classes β€” only the top few are shown, drawn from the full distribution.

Drop an image here, or click to choose
Two cats sleeping on a pink couch A cozy sunlit bedroom with a bookshelf A riverside at golden hour with a swan
Selected image preview

5 labels

BEiT vs ViT β€” same backbone, different schooling

The whole point of this page: run the same image through two Vision Transformers with the same architecture but opposite pretraining. BEiT learned self-supervised first β€” masking and predicting image patches, no labels β€” then got fine-tuned on ImageNet. The supervised ViT learned straight from labelled images. Where they agree, the label is solid; where they disagree, you're watching two different training recipes pull the same network toward different answers. Loading ViT downloads ~88 MB the first time.

See inside

BEiT's final layer produces one raw score β€” a logit β€” per class. A softmax turns those into probabilities that sum to 100%. The table shows the top classes with their raw logit next to the softmax probability. The two numbers below describe the shape of the whole distribution: the top-1 margin (how far the winner leads the runner-up) and the normalised entropy (0 = certain, 1 = spread thin).

Run a classification to see the raw numbers.
Masked image modeling β€” how BEiT learns without labels

BERT learns language by hiding words and predicting them. BEiT does the same to pictures. First, a small learned tokenizer (a discrete VAE, borrowed from DALLΒ·E) turns every 16Γ—16 image patch into one of a few thousand visual tokens β€” a vocabulary for images. Then pretraining is a fill-in-the-blank game: mask out ~40% of the patches, feed the rest to the transformer, and make it predict the visual token of each hidden patch. No class labels are involved β€” the supervision comes entirely from the image itself, which is what "self-supervised" means. To do that well, the model has to learn how shapes, textures, and objects hang together, so it builds rich general-purpose features. Only afterwards is a classification head bolted on and fine-tuned on ImageNet's labels. The payoff: self-supervised pretraining can use oceans of unlabelled images and often transfers better than purely supervised training. The head-to-head above lets you feel where that different schooling actually changes the answer β€” same ViT wiring, different upbringing.

Why it matters

BEiT was a landmark: it showed that BERT-style self-supervised pretraining works for vision, kicking off the masked-image-modeling wave (MAE, SimMIM, and more) that now underpins a lot of modern vision. For a product, the lesson is that you can pretrain strong image features on unlabelled data you already have, then fine-tune on a small labelled set.

Use cases

Basics

Classify a sample and read the top-k β€” feel a self-supervised transformer name what's in an image.

Practical

On-device tagging: classify a batch of images into a searchable tag index.

Wild

Mask patches like BEiT's own pretraining and watch how gracefully it degrades.

Multi-model

Self-supervised vs supervised: BEiT and ViT agreement on the same image.

How the API works

The convenience path is two lines with Transformers.js:

import { pipeline } from "@huggingface/transformers";

const classify = await pipeline(
  "image-classification",
  "Xenova/beit-base-patch16-224",
  { dtype: "q8" },
); // downloads + caches ~110 MB once

const output = await classify(imageURL, { top_k: 5 });
// β†’ [{ label: "Egyptian cat", score: 0.55 }, { label: "tabby, tabby cat", score: 0.26 }, …]

This page goes one level deeper for "See inside": instead of the pipeline it runs the model's own processor then model() in a single forward pass, so it can read the raw logits tensor (all 1,000 classes) and compute the softmax, entropy, and margin itself. Same weights, same result β€” it just keeps the intermediate numbers. Inference runs in a Web Worker so the page never janks. Note the masked-image-modeling pretraining is BEiT's history; the shipped model here is the fine-tuned classifier, so classification is a plain forward pass β€” no masking at inference time.

References