Skip to content

EthenEthenEthen

Open Source Model Profile · apple

mobilevitv2-1.0-imagenet1k-256

mobilevitv2-1.0-imagenet1k-256 is a MobileViTv2 image classifier from Apple. According to the model card, it replaces MobileViT attention with separable self-attention and targets ImageNet-1k classification.

Publisher
apple
Task
image-classification
Model type
mobilevitv2
License
other
Library
transformers
Publication status
Accepted · not indexed

Model overview

mobilevitv2-1.0-imagenet1k-256 is published by Apple as an image-classification model. The captured configuration identifies MobileViTv2ForImageClassification with a mobilevitv2 model type. According to the model card, MobileViTv2 is the second MobileViT version, proposed by Sachin Mehta and Mohammad Rastegari.

Recorded capabilities

Separable self-attention

According to the model card, MobileViTv2 replaces multi-headed self-attention in MobileViT with separable self-attention.

ImageNet-1k pretraining

According to the model card, the MobileViT model was pretrained on ImageNet-1k with 1 million images and 1,000 classes.

Paper and code provenance

According to the model card, the model was proposed in Separable Self-attention for Mobile Vision Transformers and first released in the linked repository.

Transformers usage path

According to the model card, classification of a COCO 2017 image uses MobileViTImageProcessor with MobileViTV2ForImageClassification.

Use cases in the source record

  • Raw image classification into 1,000 ImageNet classes using the documented MobileViT processor and classifier path, according to the model card.
  • Starting point for selecting task-specific fine-tunes through the model hub, as suggested in the model card.

Limitations and unknowns

  • The model card notes it was written by the Hugging Face team because the releasing team did not provide one, so details are secondary documentation rather than publisher-authored claims.
  • The card states the license used is Apple sample code license; captured metadata records a generic other value, so exact terms need confirmation.
  • No parameter count was extracted from this record.
  • No evaluation results were extracted from this record.
  • Provider state is historical snapshot data, not independently refreshed current availability.

Source and provenance

Source: apple/mobilevitv2-1.0-imagenet1k-256

Captured: Unknown. Processed: 2026-09-07T19:34:40.184062+00:00.

MobileViTv2 (mobilevitv2-1.0-imagenet1k-256) MobileViTv2 is the second version of MobileViT. It was proposed in Separable Self-attention for Mobile Vision Transformers by Sachin Mehta and Mohammad Rastegari, and first released in this repository. The license used is Apple sample code license . Disclaimer: The team releasing MobileViT did not write a model card for this model so this model card has been written by the Hugging Face team. Model Description MobileViTv2 is constructed by replacing the multi-headed self-attention in MobileViT with separable self-attention. Intended uses & limitations You can use the raw model for image cl…

F001F002F003F004F005F006F007F008F009F010F011F012F013F014F015