use case · basics
The whole idea, plainly. Give the model a picture; it writes a sentence. The words appear as they're generated — that's the decoder actually running, not a typing animation.

← Back to ViT-GPT2 · Next: Practical →