Skip to content

EthenEthenEthen

Open Source Model Profile · nvidia

mit-b0

mit-b0 is an nvidia SegFormer image-classification encoder. According to the model card, the b0-sized encoder was finetuned on ImageNet-1k.

Publisher
nvidia
Task
image-classification
Model type
segformer
License
other
Library
transformers
Publication status
Accepted · not indexed

Model overview

mit-b0 is published by nvidia as a segformer-based image-classification model. The captured configuration identifies SegformerForImageClassification. According to the model card, it is a SegFormer b0-sized encoder finetuned on ImageNet-1k, with a hierarchical Transformer plus lightweight all-MLP decode-head design.

Recorded capabilities

SegFormer b0 encoder

Captured config identifies SegformerForImageClassification, described in the card as a b0-sized SegFormer encoder.

ImageNet-1k hierarchical Transformer

According to the card, the hierarchical Transformer encoder was pretrained on ImageNet-1k and introduced in the SegFormer paper by Xie et al.

Fine-tuning-oriented release

The card states this repository contains the pretrained hierarchical Transformer for fine-tuning, with a Transformers workflow for 1,000-class ImageNet classification.

Use cases in the source record

  • ImageNet 1,000-class classification experiments using the card's documented Transformers workflow.
  • Fine-tuning starting points for vision work using the publisher-described pretrained hierarchical Transformer.

Limitations and unknowns

  • No parameter count was extracted from this record.
  • No evaluation results were extracted from this record.
  • The card notes it was written by the Hugging Face team because the releasing team did not provide one, so publisher attribution should be treated with care.
  • Provider state is historical snapshot data, not independently refreshed current availability.

Source and provenance

Source: nvidia/mit-b0

Captured: Unknown. Processed: 2026-09-07T19:34:54.957119+00:00.

SegFormer (b0-sized) encoder pre-trained-only SegFormer encoder fine-tuned on Imagenet-1k. It was introduced in the paper SegFormer: Simple and Efficient Design for Semantic Segmentation with Transformers by Xie et al. and first released in this repository . Disclaimer: The team releasing SegFormer did not write a model card for this model so this model card has been written by the Hugging Face team. Model description SegFormer consists of a hierarchical Transformer encoder and a lightweight all-MLP decode head to achieve great results on semantic segmentation benchmarks such as ADE20K and Cityscapes. The hierarchical Transformer is…

F001F002F003F004F005F008F009F010F011F012F013F014