image-classification Β· vision Β· WASM Β· ~88 MB
The Vision Transformer chops your picture into 16Γ16 patches, treats each like a word in a
sentence, and reads the whole thing at once. Out comes a probability for every one of the 1,000
ImageNet categories β tabby cat, espresso, street sign β
computed entirely on your device.
| Model | Xenova/vit-base-patch16-224 |
|---|---|
| Task / pipeline | image-classification |
| Architecture | ViT-B/16 β 12-layer transformer, 86M params, 197 tokens (196 patches + [CLS]) |
| Labels | ImageNet-1k β 1,000 fixed classes |
| Quantization | q8 (8-bit) ONNX |
| Backend | WebAssembly (runs anywhere; no WebGPU needed) |
| Download | ~88 MB, cached after first load |
| License | Apache-2.0 (Google Research) |
| Web APIs | WebAssembly, Web Workers, FileReader β all Baseline widely available |
Pick or drop an image and hit Classify. The bars are the softmax over ImageNet's 1,000 classes β only the top few are shown, but they're drawn from the full distribution.

5 labels
ViT's final layer produces one raw score β a logit β for each of the 1,000 classes. A softmax turns those logits into probabilities that sum to 100%. The table shows the top classes with their raw logit next to the softmax probability. Below it, two numbers describe the shape of the whole distribution: the top-1 margin (how far the winner leads the runner-up) and the normalised entropy (0 = the model is certain, 1 = it's spread thin across many classes). A confident cat photo is peaky and low-entropy; an ambiguous crop is flat and high-entropy.
| Class | logit | softmax |
|---|
Confidence (low β β high):
A fast, private, offline "what is this?" is a building block for a surprising amount of product. ViT gives you a 1,000-way answer with calibrated-ish probabilities in tens of milliseconds, with no API bill and no image ever leaving the tab.
Classify a sample and read the top-k β feel image classification first-hand.
A batch auto-tagging pipeline: classify many images into a tag index you can search.
Fool the model: crop, rotate, and cover the image and watch confidence collapse.
Classify β caption: chain ViT's labels into a full-sentence description.
The convenience path is two lines with Transformers.js:
import { pipeline } from "@huggingface/transformers";
const classify = await pipeline(
"image-classification",
"Xenova/vit-base-patch16-224",
); // downloads + caches ~88 MB once
const output = await classify(imageURL, { top_k: 5 });
// β [{ label: "tabby, tabby cat", score: 0.94 }, { label: "tiger cat", score: 0.03 }, β¦]
This page goes one level deeper for "See inside": instead of the pipeline it runs the model's own
processor then model() in a single forward pass, so it can read the raw
logits tensor (all 1,000 classes) and compute the softmax, entropy, and margin
itself. Same weights, same result β just keeping the intermediate numbers. Inference runs in a
Web Worker so the page never janks.