image-to-text Β· vision β language Β· WASM Β· ~250 MB
A Vision Transformer reads the image; a GPT-2 decoder writes about it, one word at a time. The
result is a plain-English caption β a cat laying on top of a bed β generated
entirely on your device, no captioning API in sight.
| Model | Xenova/vit-gpt2-image-captioning |
|---|---|
| Task / pipeline | image-to-text |
| Architecture | ViT-B/16 encoder β GPT-2 decoder (encoderβdecoder, cross-attention) |
| Output | Free-form English caption, generated token by token |
| Quantization | q8 (8-bit) ONNX β separate encoder + decoder graphs |
| Backend | WebAssembly (runs anywhere; no WebGPU needed) |
| Download | ~250 MB, cached after first load |
| License | Apache-2.0 (nlpconnect/vit-gpt2-image-captioning) |
| Web APIs | WebAssembly, Web Workers, FileReader β all Baseline widely available |
Pick or drop an image and hit Caption. Words appear as the decoder generates them β that streaming is the real generation, not an animation.

Captioning is autoregressive: after the ViT encoder turns the image into patch embeddings once, the GPT-2 decoder produces the caption one token at a time, each token conditioned on the image and everything written so far. The trace below shows every token as it arrived, with the elapsed time β you can literally watch the sentence being built and see where the model hesitates. The first token is the slow one (it includes the image encode); the rest come faster.
A caption is structured meaning you can act on β search it, read it aloud, translate it, gate on it. Doing it locally makes it free, private, and offline-capable, which changes where you can put it.
alt text for user-uploaded images so screen-reader users aren't left with "image".Caption a sample and watch the sentence stream in token by token.
An alt-text generator with a copy-ready <img alt> snippet β accessibility, done client-side.
Caption a live webcam frame β point your camera and let the model narrate it.
Caption β sentiment: describe the image, then read the mood of the description.
The base call is two lines with Transformers.js:
import { pipeline } from "@huggingface/transformers";
const captioner = await pipeline(
"image-to-text",
"Xenova/vit-gpt2-image-captioning",
); // downloads + caches ~250 MB once
const output = await captioner(imageURL, { max_new_tokens: 30 });
// β [{ generated_text: "a cat laying on top of a bed" }]
For "See inside", this page passes a TextStreamer into the same call. The streamer's
callback fires once per decoded token, so the page can render the caption as it forms and record
each token's arrival time β all off the main thread in a Web Worker, so the page never janks.