Self-supervised data2vec backbone
According to the card, the same data2vec learning method spans speech, language, and vision by predicting latent representations of full inputs.
Open Source Model Profile · facebook
data2vec-vision-base-ft1k is a base data2vec vision classifier from facebook. According to the model card, it is fine-tuned on ImageNet-1k at 224x224.
data2vec-vision-base-ft1k is published by facebook as an image-classification model. The captured configuration identifies Data2VecVisionForImageClassification with model type data2vec-vision. According to the model card, it applies the data2vec self-supervised framework to vision and is fine-tuned on ImageNet-1k at 224x224 resolution.
According to the card, the same data2vec learning method spans speech, language, and vision by predicting latent representations of full inputs.
According to the card, the base model was fine-tuned on ImageNet-1k at 224x224 resolution with documented RGB normalization.
According to the card, inference uses BeitFeatureExtractor with Data2VecVisionForImageClassification over 1,000 ImageNet classes.
Source: facebook/data2vec-vision-base-ft1k
Captured: Unknown. Processed: 2026-09-07T19:34:44.350683+00:00.
Data2Vec-Vision (base-sized model, fine-tuned on ImageNet-1k) BEiT model pre-trained in a self-supervised fashion and fine-tuned on ImageNet-1k (1,2 million images, 1000 classes) at resolution 224x224. It was introduced in the paper data2vec: A General Framework for Self-supervised Learning in Speech, Vision and Language by Alexei Baevski, Wei-Ning Hsu, Qiantong Xu, Arun Babu, Jiatao Gu, Michael Auli and first released in this repository . Disclaimer: The team releasing Facebook team did not write a model card for this model so this model card has been written by the Hugging Face team. Pre-Training method For more information, pleas…
F001F002F003F004F005F006F009F010F013F014F015F017F018