Skip to content

EthenEthenEthen

Open Source Model Profile · nvidia

mit-b2

mit-b2 is a SegFormer b2-sized image-classification encoder from Nvidia. According to the model card, it is a pre-trained-only encoder fine-tuned on ImageNet-1k for 1,000-class classification.

Publisher
nvidia
Task
image-classification
Model type
segformer
License
other
Library
transformers
Publication status
Accepted · not indexed

Model overview

mit-b2 is published by Nvidia as an image-classification checkpoint. Captured configuration identifies SegformerForImageClassification with model type segformer, and hub tags record ImageNet-1k association. According to the model card, it is the b2-sized SegFormer encoder introduced in SegFormer: Simple and Efficient Design for Semantic Segmentation with Transformers by Xie et al.

Recorded capabilities

B2 SegFormer encoder

According to the model card, this repository holds the b2-sized SegFormer encoder fine-tuned on ImageNet-1k.

Hierarchical Transformer plus MLP head

According to the model card, SegFormer pairs a hierarchical Transformer encoder with a lightweight all-MLP decode head, with ImageNet-1k pre-training preceding downstream fine-tuning.

Documented classification workflow

According to the model card, the model classifies a COCO 2017 image into one of 1,000 ImageNet classes with SegformerFeatureExtractor and SegformerForImageClassification.

Transformers stack

The hub record lists the Transformers library, and tags record PyTorch, TensorFlow, and endpoints compatibility.

Use cases in the source record

  • Image classification of an input image into one of 1,000 ImageNet classes using the documented SegformerFeatureExtractor and SegformerForImageClassification workflow.
  • According to the model card, the repository contains only the pre-trained hierarchical Transformer and can therefore be used for fine-tuning purposes.

Limitations and unknowns

  • The model card carries a disclaimer that the releasing team did not write it and that it was written by the Hugging Face team, so publisher attribution should be read with that provenance.
  • The captured license value is other, and the model card points to an external location for the license text; exact license terms were not extracted.
  • No evaluation results were extracted from this record.
  • No parameter count, quantization, or context-window value was extracted.
  • Provider state is historical snapshot data, not independently refreshed current availability.

Source and provenance

Source: nvidia/mit-b2

Captured: Unknown. Processed: 2026-09-07T19:34:54.976684+00:00.

SegFormer (b2-sized) encoder pre-trained-only SegFormer encoder fine-tuned on Imagenet-1k. It was introduced in the paper SegFormer: Simple and Efficient Design for Semantic Segmentation with Transformers by Xie et al. and first released in this repository . Disclaimer: The team releasing SegFormer did not write a model card for this model so this model card has been written by the Hugging Face team. Model description SegFormer consists of a hierarchical Transformer encoder and a lightweight all-MLP decode head to achieve great results on semantic segmentation benchmarks such as ADE20K and Cityscapes. The hierarchical Transformer is…

F001F002F003F004F005F006F007F008F009F010F011F012F013F014F015