use case · wild
Turn the camera on and the model describes whatever it's pointed at. Capture a single frame, or let it narrate continuously. Every pixel and every word stays on your device — nothing is uploaded. It's a glimpse of on-device assistive vision.
Latest caption
Captioning takes a few seconds per frame on WASM, so "auto-narrate" waits for each caption to finish before grabbing the next frame — it's a narrator, not a video stream.
← Back to ViT-GPT2 · Next: Multi-model →