Hierarchical local-window design
According to the card, hierarchical patch merging with self-attention inside local windows gives linear complexity in image size and supports classification and dense recognition.
Open Source Model Profile · microsoft
This Microsoft Swin V2 large classifier targets ImageNet-1k at 384x384 resolution. According to the model card, it follows ImageNet-21k pre-training.
This Microsoft release is an image-classification model in the Swin Transformer V2 large line. The captured configuration identifies Swinv2ForImageClassification with model type swinv2. According to the model card, it was pre-trained on ImageNet-21k and fine-tuned on ImageNet-1k at 384x384 resolution.
According to the card, hierarchical patch merging with self-attention inside local windows gives linear complexity in image size and supports classification and dense recognition.
According to the card, log-spaced continuous position bias transfers models pre-trained at low resolution to higher-resolution downstream tasks.
According to the card, inference uses AutoImageProcessor with AutoModelForImageClassification over 1,000 ImageNet classes, illustrated on COCO 2017 images.
Source: microsoft/swinv2-large-patch4-window12to24-192to384-22kto1k-ft
Captured: Unknown. Processed: 2026-09-07T19:34:51.941430+00:00.
Swin Transformer v2 (large-sized model) Swin Transformer v2 model pre-trained on ImageNet-21k and fine-tuned on ImageNet-1k at resolution 384x384. It was introduced in the paper Swin Transformer V2: Scaling Up Capacity and Resolution by Liu et al. and first released in this repository . Disclaimer: The team releasing Swin Transformer v2 did not write a model card for this model so this model card has been written by the Hugging Face team. Model description The Swin Transformer is a type of Vision Transformer. It builds hierarchical feature maps by merging image patches (shown in gray) in deeper layers and has linear computation comp…
F001F002F003F004F005F006F009F010F012F013F014