← All models · ViT

use case · multi-model

Multi-model: classify → caption

Labels tell you what; a caption tells you what's going on. Here ViT classifies the image into ImageNet tags, and ViT-GPT2 writes a sentence about the same picture — giving you structured tags and a human-readable description from one drop, both computed on your device.

Stage 1 — classifier (ViT, ~88 MB)

Stage 2 — captioner (ViT-GPT2, ~250 MB)

Drop an image, or click to choose
Two cats on a couch A cozy bedroom A riverside at golden hour
Selected image

1 · Tags (ViT)

2 · Caption (ViT-GPT2)

Tags and captions complement each other: tags are great for filtering and faceted search; the caption is what you show a person or feed to a screen reader. Producing both in one pass is a tidy little enrichment pipeline.

← Back to ViT