use case · wild
BEiT pretrained itself by masking image patches and predicting them. So let's
feed it exactly that: hide a slice of the 16×16 patches — the same grid BEiT sees — and
re-classify. Push the mask higher and watch how far the label survives before the picture becomes
too gappy to read. Every mask is drawn on a <canvas>; every re-classify is a
real on-device run.
Note: the shipped BEiT is the fine-tuned classifier, so it isn't reconstructing the masked patches here — you're just seeing how robust a masked-pretrained transformer is to missing patches at inference time. The 14×14 grid matches its patch size exactly.
← Practical · Back to BEiT · Next: Multi-model →