← All models

image-to-text Β· vision β†’ language Β· WASM Β· ~250 MB

ViT-GPT2 β€” write a sentence about any picture

A Vision Transformer reads the image; a GPT-2 decoder writes about it, one word at a time. The result is a plain-English caption β€” a cat laying on top of a bed β€” generated entirely on your device, no captioning API in sight.

At a glance

ModelXenova/vit-gpt2-image-captioning
Task / pipelineimage-to-text
ArchitectureViT-B/16 encoder β†’ GPT-2 decoder (encoder–decoder, cross-attention)
OutputFree-form English caption, generated token by token
Quantizationq8 (8-bit) ONNX β€” separate encoder + decoder graphs
BackendWebAssembly (runs anywhere; no WebGPU needed)
Download~250 MB, cached after first load
LicenseApache-2.0 (nlpconnect/vit-gpt2-image-captioning)
Web APIsWebAssembly, Web Workers, FileReader β€” all Baseline widely available

Run it

Pick or drop an image and hit Caption. Words appear as the decoder generates them β€” that streaming is the real generation, not an animation.

Drop an image here, or click to choose
Two cats sleeping on a pink couch A cozy sunlit bedroom with a bookshelf A riverside at golden hour with a swan
Selected image preview

See inside

Captioning is autoregressive: after the ViT encoder turns the image into patch embeddings once, the GPT-2 decoder produces the caption one token at a time, each token conditioned on the image and everything written so far. The trace below shows every token as it arrived, with the elapsed time β€” you can literally watch the sentence being built and see where the model hesitates. The first token is the slow one (it includes the image encode); the rest come faster.

Generate a caption to see the token-by-token trace.

Why it matters

A caption is structured meaning you can act on β€” search it, read it aloud, translate it, gate on it. Doing it locally makes it free, private, and offline-capable, which changes where you can put it.

Use cases

Basics

Caption a sample and watch the sentence stream in token by token.

Practical

An alt-text generator with a copy-ready <img alt> snippet β€” accessibility, done client-side.

Wild

Caption a live webcam frame β€” point your camera and let the model narrate it.

Multi-model

Caption β†’ sentiment: describe the image, then read the mood of the description.

How the API works

The base call is two lines with Transformers.js:

import { pipeline } from "@huggingface/transformers";

const captioner = await pipeline(
  "image-to-text",
  "Xenova/vit-gpt2-image-captioning",
); // downloads + caches ~250 MB once

const output = await captioner(imageURL, { max_new_tokens: 30 });
// β†’ [{ generated_text: "a cat laying on top of a bed" }]

For "See inside", this page passes a TextStreamer into the same call. The streamer's callback fires once per decoded token, so the page can render the caption as it forms and record each token's arrival time β€” all off the main thread in a Web Worker, so the page never janks.

References