← All models

image-classification Β· vision Β· WASM Β· ~88 MB

ViT β€” name what's in an image, in 1,000 ways

The Vision Transformer chops your picture into 16Γ—16 patches, treats each like a word in a sentence, and reads the whole thing at once. Out comes a probability for every one of the 1,000 ImageNet categories β€” tabby cat, espresso, street sign β€” computed entirely on your device.

At a glance

ModelXenova/vit-base-patch16-224
Task / pipelineimage-classification
ArchitectureViT-B/16 β€” 12-layer transformer, 86M params, 197 tokens (196 patches + [CLS])
LabelsImageNet-1k β€” 1,000 fixed classes
Quantizationq8 (8-bit) ONNX
BackendWebAssembly (runs anywhere; no WebGPU needed)
Download~88 MB, cached after first load
LicenseApache-2.0 (Google Research)
Web APIsWebAssembly, Web Workers, FileReader β€” all Baseline widely available

Run it

Pick or drop an image and hit Classify. The bars are the softmax over ImageNet's 1,000 classes β€” only the top few are shown, but they're drawn from the full distribution.

Drop an image here, or click to choose
Two cats sleeping on a pink couch A cozy sunlit bedroom with a bookshelf A riverside at golden hour with a swan
Selected image preview

5 labels

See inside

ViT's final layer produces one raw score β€” a logit β€” for each of the 1,000 classes. A softmax turns those logits into probabilities that sum to 100%. The table shows the top classes with their raw logit next to the softmax probability. Below it, two numbers describe the shape of the whole distribution: the top-1 margin (how far the winner leads the runner-up) and the normalised entropy (0 = the model is certain, 1 = it's spread thin across many classes). A confident cat photo is peaky and low-entropy; an ambiguous crop is flat and high-entropy.

Run a classification to see the raw numbers.

Why it matters

A fast, private, offline "what is this?" is a building block for a surprising amount of product. ViT gives you a 1,000-way answer with calibrated-ish probabilities in tens of milliseconds, with no API bill and no image ever leaving the tab.

Use cases

Basics

Classify a sample and read the top-k β€” feel image classification first-hand.

Practical

A batch auto-tagging pipeline: classify many images into a tag index you can search.

Wild

Fool the model: crop, rotate, and cover the image and watch confidence collapse.

Multi-model

Classify β†’ caption: chain ViT's labels into a full-sentence description.

How the API works

The convenience path is two lines with Transformers.js:

import { pipeline } from "@huggingface/transformers";

const classify = await pipeline(
  "image-classification",
  "Xenova/vit-base-patch16-224",
); // downloads + caches ~88 MB once

const output = await classify(imageURL, { top_k: 5 });
// β†’ [{ label: "tabby, tabby cat", score: 0.94 }, { label: "tiger cat", score: 0.03 }, …]

This page goes one level deeper for "See inside": instead of the pipeline it runs the model's own processor then model() in a single forward pass, so it can read the raw logits tensor (all 1,000 classes) and compute the softmax, entropy, and margin itself. Same weights, same result β€” just keeping the intermediate numbers. Inference runs in a Web Worker so the page never janks.

References