Skip to content

EthenEthenEthen

Open Source Model Profile · microsoft

beit-base-patch16-224

beit-base-patch16-224 is an 87M-parameter BEiT image-classification fine-tune from Microsoft. Its card documents ImageNet-21k pretraining and ImageNet-1k fine-tuning at 224x224.

Publisher
microsoft
Task
image-classification
Model type
beit
License
apache-2.0
Library
transformers
Publication status
Accepted · not indexed

Model overview

beit-base-patch16-224 is published by Microsoft as an image-classification model. The captured configuration identifies BeitForImageClassification with a beit model type, and Safetensors metadata reports 86996692 parameters. The model card describes a base-sized BEiT model pretrained on ImageNet-21k and fine-tuned on ImageNet-1k, both at 224x224 resolution.

Recorded capabilities

87M BEiT scale

Captured config identifies BeitForImageClassification and Safetensors metadata reports 86996692 parameters.

ImageNet-21k to ImageNet-1k lineage

According to the model card, the model was pretrained on ImageNet-21k with 14 million images and 21,841 classes, then fine-tuned on ImageNet 2012 with 1 million images and 1,000 classes.

Masked visual-token objective

The card describes self-supervised pretraining that predicts visual tokens from OpenAI DALL-E VQ-VAE encodings for masked patches.

16x16 patches and relative positions

According to the card, images use 16x16 patches with relative position embeddings and mean-pooled classification rather than a CLS linear head.

Documented 224 preprocessing

The card says images are resized to 224x224 and normalized with mean and standard deviation of 0.5 across RGB channels.

Use cases in the source record

  • Image classification into 1,000 ImageNet classes using the documented BeitImageProcessor and BeitForImageClassification workflow.
  • Vision feature extraction by placing a classifier over the pretrained encoder or mean-pooled patch states, as described in the card.

Limitations and unknowns

  • No evaluation scores were extracted from this record; the card refers to tables 1 and 2 of the original paper.
  • No context-window value applies to this vision record as extracted.
  • Provider state is historical snapshot data, not independently refreshed current availability.
  • Training and evaluation details come from a publisher card written by the Hugging Face team and the cited paper, not independently verified by Ethen.

Source and provenance

Source: microsoft/beit-base-patch16-224

Captured: Unknown. Processed: 2026-09-07T19:34:51.899074+00:00.

BEiT (base-sized model, fine-tuned on ImageNet-1k) BEiT model pre-trained in a self-supervised fashion on ImageNet-21k (14 million images, 21,841 classes) at resolution 224x224, and fine-tuned on ImageNet 2012 (1 million images, 1,000 classes) at resolution 224x224. It was introduced in the paper BEIT: BERT Pre-Training of Image Transformers by Hangbo Bao, Li Dong and Furu Wei and first released in this repository . Disclaimer: The team releasing BEiT did not write a model card for this model so this model card has been written by the Hugging Face team. Model description The BEiT model is a Vision Transformer (ViT), which is a trans…

F001F002F003F004F005F006F007F008F010F011F012F013F015F016F017F018F020F021F022