use case · multi-model
Labels tell you what; a caption tells you what's going on. Here ViT classifies the image into ImageNet tags, and ViT-GPT2 writes a sentence about the same picture — giving you structured tags and a human-readable description from one drop, both computed on your device.
Stage 1 — classifier (ViT, ~88 MB)
Stage 2 — captioner (ViT-GPT2, ~250 MB)

1 · Tags (ViT)
2 · Caption (ViT-GPT2)
Tags and captions complement each other: tags are great for filtering and faceted search; the caption is what you show a person or feed to a screen reader. Producing both in one pass is a tidy little enrichment pipeline.