Tiny Swin v2 at 256x256
According to the model card, this is the tiny-sized Swin Transformer v2 variant pre-trained on ImageNet-1k at 256x256 resolution.
Open Source Model Profile · microsoft
swinv2-tiny-patch4-window16-256 is a tiny Swin Transformer v2 image-classification model from Microsoft. According to the model card, it is pre-trained on ImageNet-1k at 256x256 resolution for 1,000-class classification.
swinv2-tiny-patch4-window16-256 is published by Microsoft as an image-classification checkpoint. Captured configuration identifies Swinv2ForImageClassification with model type swinv2, and hub tags record ImageNet-1k association. According to the model card, it is a tiny Swin Transformer v2 model pre-trained on ImageNet-1k at 256x256, introduced in the Swin Transformer V2 paper by Liu et al.
According to the model card, this is the tiny-sized Swin Transformer v2 variant pre-trained on ImageNet-1k at 256x256 resolution.
According to the model card, the architecture builds hierarchical feature maps with local-window self-attention for linear complexity relative to image size.
According to the model card, Swin v2 adds residual post-norm with cosine attention, log-spaced continuous position bias, and the SimMIM self-supervised pre-training method.
The hub record lists the Transformers library with an Apache-2.0 license, and tags record PyTorch and endpoints compatibility.
Source: microsoft/swinv2-tiny-patch4-window16-256
Captured: Unknown. Processed: 2026-09-07T19:34:51.955524+00:00.
Swin Transformer v2 (tiny-sized model) Swin Transformer v2 model pre-trained on ImageNet-1k at resolution 256x256. It was introduced in the paper Swin Transformer V2: Scaling Up Capacity and Resolution by Liu et al. and first released in this repository . Disclaimer: The team releasing Swin Transformer v2 did not write a model card for this model so this model card has been written by the Hugging Face team. Model description The Swin Transformer is a type of Vision Transformer. It builds hierarchical feature maps by merging image patches (shown in gray) in deeper layers and has linear computation complexity to input image size due t…
F001F002F003F004F005F006F007F008F009F010F011F012F013F014