Skip to content

EthenEthenEthen

Open Source Model Profile · facebook

deit-base-distilled-patch16-384

deit-base-distilled-patch16-384 is a distilled DeiT image classifier from facebook. According to the model card, it was pretrained at 224x224 and fine-tuned at 384x384 on ImageNet-1k.

Publisher
facebook
Task
image-classification
Model type
deit
License
apache-2.0
Library
transformers
Publication status
Accepted · not indexed

Model overview

deit-base-distilled-patch16-384 is published by facebook as an image-classification model. The captured configuration identifies DeiTForImageClassificationWithTeacher with a deit model type, and Safetensors metadata reports 87,630,032 parameters. According to the model card, it is a base-sized distilled DeiT model fine-tuned at 384x384 on ImageNet-1k.

Recorded capabilities

Distilled DeiT design

According to the model card, this distilled Vision Transformer uses a distillation token plus class token to learn from a CNN teacher during pretraining and fine-tuning.

384x384 fine-tuning

According to the model card, the model was pretrained at 224x224 and fine-tuned at 384x384 on ImageNet-1k with 1 million images and 1,000 classes.

Reported 85.2% top-1 accuracy

According to the model card's comparison table, the 384 distilled base variant reports 85.2% ImageNet top-1 and 97.2% top-5 accuracy at about 88M parameters.

16x16 patch input

According to the model card, images enter as 16x16 patches that are linearly embedded.

Use cases in the source record

  • Image classification into 1,000 ImageNet classes using the documented Transformers classifier and feature-extractor path, according to the model card.
  • Distillation-based Vision Transformer experiments that use the extra distillation token alongside the class token, as described in the model card.

Limitations and unknowns

  • The model card notes it was written by the Hugging Face team because the releasing team did not provide one, so details are secondary documentation rather than publisher-authored claims.
  • Publisher-reported accuracy figures are model-card claims and have not been independently verified by Ethen.
  • No current provider availability has been independently refreshed; snapshot state is historical.

Source and provenance

Source: facebook/deit-base-distilled-patch16-384

Captured: Unknown. Processed: 2026-09-07T19:34:44.374220+00:00.

Distilled Data-efficient Image Transformer (base-sized model) Distilled data-efficient Image Transformer (DeiT) model pre-trained at resolution 224x224 and fine-tuned at resolution 384x384 on ImageNet-1k (1 million images, 1,000 classes). It was first introduced in the paper Training data-efficient image transformers & distillation through attention by Touvron et al. and first released in this repository . However, the weights were converted from the timm repository by Ross Wightman. Disclaimer: The team releasing DeiT did not write a model card for this model so this model card has been written by the Hugging Face team. Model descr…

F001F002F003F004F005F006F007F008F009F010F011F012F013F014F015F016F019F020