use case · multi-model
One model can't do this alone. First ViT-GPT2 turns the image into a sentence; then DistilBERT reads that sentence and scores its mood. Two models, two modalities, chained end-to-end — and both run entirely on your device.
Stage 1 — captioner (ViT-GPT2, ~250 MB)
Stage 2 — sentiment (DistilBERT SST-2, ~67 MB)

1 · Caption
2 · Sentiment of that caption
Why chain them? A caption is a compact, searchable summary of an image; sentiment turns that into a signal you can sort or filter on — e.g. surface the "happiest-looking" photos in a library, all without a server.