← All models · BEiT

use case · wild

Wild: mask the patches — play BEiT's own game

BEiT pretrained itself by masking image patches and predicting them. So let's feed it exactly that: hide a slice of the 16×16 patches — the same grid BEiT sees — and re-classify. Push the mask higher and watch how far the label survives before the picture becomes too gappy to read. Every mask is drawn on a <canvas>; every re-classify is a real on-device run.

Two cats on a couch A cozy bedroom A riverside at golden hour
…or drop your own image
40%

Note: the shipped BEiT is the fine-tuned classifier, so it isn't reconstructing the masked patches here — you're just seeing how robust a masked-pretrained transformer is to missing patches at inference time. The 14×14 grid matches its patch size exactly.

← Practical · Back to BEiT · Next: Multi-model →