87M BEiT scale
Captured config identifies BeitForImageClassification and Safetensors metadata reports 86996692 parameters.
Open Source Model Profile · microsoft
beit-base-patch16-224 is an 87M-parameter BEiT image-classification fine-tune from Microsoft. Its card documents ImageNet-21k pretraining and ImageNet-1k fine-tuning at 224x224.
beit-base-patch16-224 is published by Microsoft as an image-classification model. The captured configuration identifies BeitForImageClassification with a beit model type, and Safetensors metadata reports 86996692 parameters. The model card describes a base-sized BEiT model pretrained on ImageNet-21k and fine-tuned on ImageNet-1k, both at 224x224 resolution.
Captured config identifies BeitForImageClassification and Safetensors metadata reports 86996692 parameters.
According to the model card, the model was pretrained on ImageNet-21k with 14 million images and 21,841 classes, then fine-tuned on ImageNet 2012 with 1 million images and 1,000 classes.
The card describes self-supervised pretraining that predicts visual tokens from OpenAI DALL-E VQ-VAE encodings for masked patches.
According to the card, images use 16x16 patches with relative position embeddings and mean-pooled classification rather than a CLS linear head.
The card says images are resized to 224x224 and normalized with mean and standard deviation of 0.5 across RGB channels.
Source: microsoft/beit-base-patch16-224
Captured: Unknown. Processed: 2026-09-07T19:34:51.899074+00:00.
BEiT (base-sized model, fine-tuned on ImageNet-1k) BEiT model pre-trained in a self-supervised fashion on ImageNet-21k (14 million images, 21,841 classes) at resolution 224x224, and fine-tuned on ImageNet 2012 (1 million images, 1,000 classes) at resolution 224x224. It was introduced in the paper BEIT: BERT Pre-Training of Image Transformers by Hangbo Bao, Li Dong and Furu Wei and first released in this repository . Disclaimer: The team releasing BEiT did not write a model card for this model so this model card has been written by the Hugging Face team. Model description The BEiT model is a Vision Transformer (ViT), which is a trans…
F001F002F003F004F005F006F007F008F010F011F012F013F015F016F017F018F020F021F022