Skip to content

EthenEthenEthen

Open Source Model Profile · facebook

deit-small-patch16-224

deit-small-patch16-224 is a facebook DeiT image-classification model for ImageNet-1k at 224x224 resolution. Its model card describes a data-efficient Vision Transformer with 79.9% top-1 accuracy.

Publisher
facebook
Task
image-classification
Model type
vit
License
apache-2.0
Library
transformers
Publication status
Accepted · not indexed

Model overview

deit-small-patch16-224 is published by facebook as an image-classification model. The captured configuration identifies ViTForImageClassification with model type vit. According to the model card, it was pre-trained and fine-tuned on ImageNet-1k with 1 million images and 1,000 classes.

Recorded capabilities

Data-efficient ViT design

According to the model card, this is a more efficiently trained Vision Transformer rather than a standard ViT run.

16x16 patch input

The publisher describes fixed 16x16 patches with linear embeddings, a CLS token, and absolute position embeddings.

Reported ImageNet scores

The card table reports 79.9% top-1 and 95.0% top-5 accuracy with a 22M parameter figure for DeiT-small.

Use cases in the source record

  • Image classification into the 1,000 ImageNet classes using the card's DeiT feature-extractor path.
  • Feature extraction by placing a linear classifier over the pre-trained encoder's CLS representation.

Limitations and unknowns

  • The model card notes it was written by the Hugging Face team because the releasing team did not supply one.
  • Accuracy and parameter figures above come from the card's comparison table, not Ethen-measured results.
  • No Safetensors parameter count was extracted from this record.
  • Provider state is historical snapshot data, not independently refreshed current availability.

Source and provenance

Source: facebook/deit-small-patch16-224

Captured: Unknown. Processed: 2026-09-07T19:34:44.412523+00:00.

Data-efficient Image Transformer (small-sized model) Data-efficient Image Transformer (DeiT) model pre-trained and fine-tuned on ImageNet-1k (1 million images, 1,000 classes) at resolution 224x224. It was first introduced in the paper Training data-efficient image transformers & distillation through attention by Touvron et al. and first released in this repository . However, the weights were converted from the timm repository by Ross Wightman. Disclaimer: The team releasing DeiT did not write a model card for this model so this model card has been written by the Hugging Face team. Model description This model is actually a more effi…

F001F002F003F004F005F006F007F008F009F010F011F012F013F014F015F016F017F020F021F022