← All models · ViT-GPT2

use case · basics

Basics: describe an image

The whole idea, plainly. Give the model a picture; it writes a sentence. The words appear as they're generated — that's the decoder actually running, not a typing animation.

Drop an image, or click to choose
Two cats on a couch A cozy bedroom A riverside at golden hour
Selected image

← Back to ViT-GPT2 · Next: Practical →