Latent cross-attention design
According to the card, self-attention runs over a compact latent set while inputs supply cross-attention, so attention cost does not depend on input size.
Open Source Model Profile · deepmind
vision-perceiver-conv is a Perceiver IO image-classification model from deepmind. According to the card, it was pre-trained on ImageNet at 224x224 resolution.
vision-perceiver-conv is published by deepmind as an image-classification model. The captured configuration identifies PerceiverForImageClassificationConvProcessing with model type perceiver. According to the card, it is a Perceiver IO vision model pre-trained on ImageNet with 14 million images and 1,000 classes at 224x224 resolution.
According to the card, self-attention runs over a compact latent set while inputs supply cross-attention, so attention cost does not depend on input size.
According to the card, decoder queries decode latent states into outputs of flexible size, with image-classification output as logits of shape (batch_size, num_labels).
According to the card, the pre-trained image representation supports downstream tasks, including training a classifier by replacing the classification decoder.
Source: deepmind/vision-perceiver-conv
Captured: Unknown. Processed: 2026-09-07T19:34:43.549949+00:00.
Perceiver IO for vision (convolutional processing) Perceiver IO model pre-trained on ImageNet (14 million images, 1,000 classes) at resolution 224x224. It was introduced in the paper Perceiver IO: A General Architecture for Structured Inputs & Outputs by Jaegle et al. and first released in this repository . Disclaimer: The team releasing Perceiver IO did not write a model card for this model so this model card has been written by the Hugging Face team. Model description Perceiver IO is a transformer encoder model that can be applied on any modality (text, images, audio, video, ...). The core idea is to employ the self-attention mech…
F001F002F003F004F005F006F009F010F012F013F015F016F017