Skip to content

EthenEthenEthen

Open Source Model Profile · microsoft

beit-large-finetuned-ade-640-640

beit-large-finetuned-ade-640-640 is a 502.82M-parameter BEiT image-segmentation model from microsoft. According to the model card, it is fine-tuned on ADE20k at 640x640 resolution.

Publisher
microsoft
Task
image-segmentation
Model type
beit
License
apache-2.0
Library
transformers
Publication status
Accepted · not indexed

Model overview

beit-large-finetuned-ade-640-640 is published by microsoft as a BEiT image-segmentation model. The captured configuration identifies BeitForSemanticSegmentation and Safetensors metadata reports 502,815,044 parameters. According to the model card, it was pretrained on ImageNet-21k and fine-tuned on ADE20k at 640x640 resolution.

Recorded capabilities

502.82M BEiT with Transformers

The captured configuration reports BeitForSemanticSegmentation with model type beit and 502,815,044 parameters, with Transformers support.

ADE20k tuning at 640x640

According to the model card, the large-sized model is fine-tuned on ADE20k at resolution 640x640.

ImageNet-21k pretraining

According to the model card, the model was pretrained in self-supervised fashion on ImageNet-21k with 14 million images and 21,841 classes at 224x224.

Documented preprocessing

According to the model card, images are cropped and padded to 640x640 and normalized with ImageNet mean and standard deviation.

Use cases in the source record

  • Semantic segmentation of images into ADE20k categories using the documented 640x640 workflow.
  • Vision-Transformer feature experiments building on ImageNet-21k pretraining with an mmseg-style decode head.

Limitations and unknowns

  • No Ethen-measured evaluation results were extracted; the card refers evaluation to tables 1 and 2 of the original paper.
  • According to the model card disclaimer, the releasing team did not write the card and the Hugging Face team wrote it instead.
  • Provider state is historical snapshot data, not independently refreshed current availability.

Source and provenance

Source: microsoft/beit-large-finetuned-ade-640-640

Captured: Unknown. Processed: 2026-09-07T19:34:51.948546+00:00.

BEiT (large-sized model, fine-tuned on ADE20k) BEiT model pre-trained in a self-supervised fashion on ImageNet-21k (14 million images, 21,841 classes) at resolution 224x224, and fine-tuned on ADE20k (an important benchmark for semantic segmentation of images) at resolution 640x640. It was introduced in the paper BEIT: BERT Pre-Training of Image Transformers by Hangbo Bao, Li Dong and Furu Wei and first released in this repository . Disclaimer: The team releasing BEiT did not write a model card for this model so this model card has been written by the Hugging Face team. Model description The BEiT model is a Vision Transformer (ViT),…

F001F002F003F004F005F006F007F008F009F010F011F012F013F016F017F019F021