use case · multi-model
Same ViT backbone, opposite schooling. Run the same image through BEiT (pretrained self-supervised by masking patches) and the supervised ViT, average their softmax distributions, and the ensemble is usually more robust than either alone. Because their only real difference is how they were pretrained, their agreement is an unusually clean signal — when two training recipes on identical wiring concur, trust it. Loading ViT downloads ~88 MB the first time.

Self-supervised (MIM) · –
Supervised · –