Skip to content

EthenEthenEthen

Open Source Model Profile · facebook

data2vec-vision-base-ft1k

data2vec-vision-base-ft1k is a base data2vec vision classifier from facebook. According to the model card, it is fine-tuned on ImageNet-1k at 224x224.

Publisher
facebook
Task
image-classification
Model type
data2vec-vision
License
apache-2.0
Library
transformers
Publication status
Approved for indexing

Model overview

data2vec-vision-base-ft1k is published by facebook as an image-classification model. The captured configuration identifies Data2VecVisionForImageClassification with model type data2vec-vision. According to the model card, it applies the data2vec self-supervised framework to vision and is fine-tuned on ImageNet-1k at 224x224 resolution.

Recorded capabilities

Self-supervised data2vec backbone

According to the card, the same data2vec learning method spans speech, language, and vision by predicting latent representations of full inputs.

ImageNet-1k fine-tune at 224

According to the card, the base model was fine-tuned on ImageNet-1k at 224x224 resolution with documented RGB normalization.

Documented Transformers recipe

According to the card, inference uses BeitFeatureExtractor with Data2VecVisionForImageClassification over 1,000 ImageNet classes.

Use cases in the source record

  • Image classification into 1,000 ImageNet classes at 224x224 resolution through the documented Transformers recipe.
  • Representation-reuse experiments starting from the ImageNet-1k fine-tuned vision backbone.

Limitations and unknowns

  • No parameter count was extracted from this record.
  • No evaluation scores were extracted; the card defers evaluation to Table 1 of the original paper.
  • Provider state is historical snapshot data, not independently refreshed current availability.

Source and provenance

Source: facebook/data2vec-vision-base-ft1k

Captured: Unknown. Processed: 2026-09-07T19:34:44.350683+00:00.

Data2Vec-Vision (base-sized model, fine-tuned on ImageNet-1k) BEiT model pre-trained in a self-supervised fashion and fine-tuned on ImageNet-1k (1,2 million images, 1000 classes) at resolution 224x224. It was introduced in the paper data2vec: A General Framework for Self-supervised Learning in Speech, Vision and Language by Alexei Baevski, Wei-Ning Hsu, Qiantong Xu, Arun Babu, Jiatao Gu, Michael Auli and first released in this repository . Disclaimer: The team releasing Facebook team did not write a model card for this model so this model card has been written by the Hugging Face team. Pre-Training method For more information, pleas…

F001F002F003F004F005F006F009F010F013F014F015F017F018