Skip to content

EthenEthenEthen

Open Source Model Profile · MBZUAI

swiftformer-xs

swiftformer-xs is a 3.48M-parameter MBZUAI image classifier. Its card traces it to the SwiftFormer paper on efficient additive attention for mobile vision.

Publisher
MBZUAI
Task
image-classification
Model type
swiftformer
License
Unknown
Library
transformers
Publication status
Accepted · not indexed

Model overview

swiftformer-xs is published by MBZUAI as an image-classification model with SwiftFormerForImageClassification architecture and a swiftformer model type. Safetensors metadata reports 3,481,320 parameters, or about 3.48M. According to the model card, the classification model is trained on ImageNet-1K and served through the Transformers stack.

Recorded capabilities

Efficient additive attention

The card describes an additive attention mechanism that swaps quadratic matrix multiplication for linear element-wise operations.

Published paper attribution

The card names the SwiftFormer paper on efficient additive attention for real-time mobile vision applications.

ImageNet-1K training

The card states the classification model is trained on the ImageNet-1K dataset.

Use cases in the source record

  • Lightweight ImageNet-1K image classification through the captured inference endpoint or Transformers workflows.
  • Mobile-oriented vision experiments building on the card's efficient additive-attention design.

Limitations and unknowns

  • No license value was extracted from this record.
  • No evaluation results specific to the xs variant were extracted; the card's 78.5% figure refers to the paper's small variant.
  • Provider state is historical snapshot data, not independently refreshed current availability.
  • Design and performance claims come from the publisher card's paper summary and have not been independently verified by Ethen.

Source and provenance

Source: MBZUAI/swiftformer-xs

Captured: Unknown. Processed: 2026-09-07T19:34:33.803084+00:00.

SwiftFormer (swiftformer-xs) Model description The SwiftFormer model was proposed in SwiftFormer: Efficient Additive Attention for Transformer-based Real-time Mobile Vision Applications by Abdelrahman Shaker, Muhammad Maaz, Hanoona Rasheed, Salman Khan, Ming-Hsuan Yang, Fahad Shahbaz Khan. SwiftFormer paper introduces a novel efficient additive attention mechanism that effectively replaces the quadratic matrix multiplication operations in the self-attention computation with linear element-wise multiplications. A series of models called 'SwiftFormer' is built based on this, which achieves state-of-the-art performance in terms of both…

F001F002F003F004F005F006F007F009F010F011F012F014