Skip to content

EthenEthenEthen

Open Source Model Profile · nvidia

mit-b3

mit-b3 is an nvidia SegFormer model for image classification. The model card describes a b3-sized encoder fine-tuned on ImageNet-1k for Transformers use.

Publisher
nvidia
Task
image-classification
Model type
segformer
License
other
Library
transformers
Publication status
Accepted · not indexed

Model overview

mit-b3 is published by nvidia as a SegFormer image-classification model. The captured configuration identifies SegformerForImageClassification with model type segformer. According to the model card, it is a b3-sized encoder fine-tuned on ImageNet-1k and introduced in the SegFormer paper by Xie et al. The card notes it was written by the Hugging Face team, not the team releasing SegFormer.

Recorded capabilities

B3 ImageNet-1k encoder

According to the model card, this is a SegFormer b3-sized encoder fine-tuned on ImageNet-1k.

Hierarchical Transformer design

According to the model card, SegFormer combines a hierarchical Transformer encoder with a lightweight all-MLP decode head.

Fine-tuning-oriented checkpoint

The model card says this repository contains only the pre-trained hierarchical Transformer for fine-tuning purposes.

Documented classification example

The model card documents classifying an image into one of the 1,000 ImageNet classes with SegformerFeatureExtractor and SegformerForImageClassification.

Use cases in the source record

  • Encoder fine-tuning workflows using the repository's pre-trained hierarchical Transformer.
  • Image classification into one of the 1,000 ImageNet classes using the card's documented Transformers feature-extractor workflow.

Limitations and unknowns

  • No parameter count was extracted from this record.
  • No evaluation results were extracted from this record.
  • The hub license value is other; the model card points to an external license reference whose full terms were not extracted.
  • Provider state is historical snapshot data, not independently refreshed current availability.

Source and provenance

Source: nvidia/mit-b3

Captured: Unknown. Processed: 2026-09-07T19:34:54.069375+00:00.

SegFormer (b3-sized) encoder pre-trained-only SegFormer encoder fine-tuned on Imagenet-1k. It was introduced in the paper SegFormer: Simple and Efficient Design for Semantic Segmentation with Transformers by Xie et al. and first released in this repository . Disclaimer: The team releasing SegFormer did not write a model card for this model so this model card has been written by the Hugging Face team. Model description SegFormer consists of a hierarchical Transformer encoder and a lightweight all-MLP decode head to achieve great results on semantic segmentation benchmarks such as ADE20K and Cityscapes. The hierarchical Transformer is…

F001F002F003F004F005F006F007F008F009F010F011F012F013F014F015