Distillation-token design
The card describes a distillation token that interacts with class and patch tokens through self-attention to learn from the teacher.
Open Source Model Profile · facebook
deit-small-distilled-patch16-224 is a distilled DeiT image classifier from Facebook trained on ImageNet-1k at 224x224. Its card documents distillation-token learning and a published 81.2% top-1 accuracy row.
deit-small-distilled-patch16-224 is published by Facebook as an image-classification model with DeiTForImageClassificationWithTeacher architecture and a deit model type. According to the model card, it is a small distilled Data-efficient Image Transformer pre-trained and fine-tuned with distillation on ImageNet-1k at 224x224. Captured metadata records an Apache-2.0 license and Transformers library compatibility.
The card describes a distillation token that interacts with class and patch tokens through self-attention to learn from the teacher.
The card specifies 256x256 resizing, 224x224 center-cropping, and ImageNet mean and standard-deviation normalization.
The card's evaluation table attributes 81.2% top-1 and 95.4% top-5 ImageNet accuracy to this distilled small variant.
Source: facebook/deit-small-distilled-patch16-224
Captured: Unknown. Processed: 2026-09-07T19:34:44.394385+00:00.
Distilled Data-efficient Image Transformer (small-sized model) Distilled data-efficient Image Transformer (DeiT) model pre-trained and fine-tuned on ImageNet-1k (1 million images, 1,000 classes) at resolution 224x224. It was first introduced in the paper Training data-efficient image transformers & distillation through attention by Touvron et al. and first released in this repository . However, the weights were converted from the timm repository by Ross Wightman. Disclaimer: The team releasing DeiT did not write a model card for this model so this model card has been written by the Hugging Face team. Model description This model is…
F001F002F003F004F005F006F009F010F011F012F014F015F017F019