← All models · ViT-GPT2

use case · multi-model

Multi-model: caption → sentiment

One model can't do this alone. First ViT-GPT2 turns the image into a sentence; then DistilBERT reads that sentence and scores its mood. Two models, two modalities, chained end-to-end — and both run entirely on your device.

Stage 1 — captioner (ViT-GPT2, ~250 MB)

Stage 2 — sentiment (DistilBERT SST-2, ~67 MB)

Drop an image, or click to choose
Two cats on a couch A cozy bedroom A riverside at golden hour
Selected image

1 · Caption

2 · Sentiment of that caption

Why chain them? A caption is a compact, searchable summary of an image; sentiment turns that into a signal you can sort or filter on — e.g. surface the "happiest-looking" photos in a library, all without a server.

← Back to ViT-GPT2