Skip to content

EthenEthenEthen

Open Source Model Profile · microsoft

cvt-13

cvt-13 is a 20.02M-parameter convolutional vision transformer for image classification from microsoft. According to the model card, it was pre-trained on ImageNet-1k at 224x224 resolution.

Publisher
microsoft
Task
image-classification
Model type
cvt
License
apache-2.0
Library
transformers
Publication status
Accepted · not indexed

Model overview

cvt-13 is published by microsoft as an image-classification model. The captured configuration identifies CvtForImageClassification with model type cvt, and Safetensors metadata reports 20,023,208 parameters. According to the model card, it is the CvT-13 model pre-trained on ImageNet-1k at 224x224.

Recorded capabilities

Convolutional vision transformer

According to the model card, CvT-13 introduces convolutions to vision transformers, captured as CvtForImageClassification with model type cvt.

ImageNet-1k at 224x224

According to the model card, the model was pre-trained on ImageNet-1k at 224x224 resolution, consistent with its dataset:imagenet-1k tag.

Transformers-compatible release

Hub metadata identifies transformers library support with pytorch, tf, safetensors, and endpoints compatibility.

Documented classification workflow

According to the model card, the card shows classifying a COCO 2017 image into one of the 1,000 ImageNet classes with AutoFeatureExtractor and CvtForImageClassification.

Use cases in the source record

  • Image classification of inputs into one of the 1,000 ImageNet classes using the model card's documented Transformers workflow.
  • Architecture study of convolution-augmented transformers referencing the card's cited CvT paper and ImageNet-1k setup.

Limitations and unknowns

  • No independently measured evaluation results were extracted from this record.
  • No training hyperparameters, compute, or dataset size beyond ImageNet-1k naming were extracted.
  • According to the card disclaimer, the releasing team did not write the model card; the Hugging Face team wrote it instead.
  • Provider state is historical snapshot data, not independently refreshed current availability.

Source and provenance

Source: microsoft/cvt-13

Captured: Unknown. Processed: 2026-09-07T19:34:52.031479+00:00.

Convolutional Vision Transformer (CvT) CvT-13 model pre-trained on ImageNet-1k at resolution 224x224. It was introduced in the paper CvT: Introducing Convolutions to Vision Transformers by Wu et al. and first released in this repository . Disclaimer: The team releasing CvT did not write a model card for this model so this model card has been written by the Hugging Face team. Usage Here is how to use this model to classify an image of the COCO 2017 dataset into one of the 1,000 ImageNet classes: from transformers import AutoFeatureExtractor, CvtForImageClassification from PIL import Image import requests url = 'http://images.cocodata…

F001F002F003F004F005F006F007F008F009F010F011F012F013F014