use case · basics
The whole idea in its simplest form. Give BEiT a picture; the self-supervised transformer hands back its best guesses from the 1,000 ImageNet categories, each with a probability. Bigger bar = more sure.

The bars are a softmax over all 1,000 classes, so the ones shown here rarely add to 100% — the rest of the probability is spread across the classes that didn't make the cut.
← Back to BEiT · Next: Practical →