Swin hierarchical windows
According to the model card, the architecture merges image patches in deeper layers and computes self-attention within local windows.
Open Source Model Profile · microsoft
swin-base-patch4-window12-384 is a microsoft Swin Transformer model for image classification. The card describes ImageNet-1k training at 384x384 resolution under Apache-2.0.
The model is published by microsoft as a base-sized Swin Transformer for image classification. The captured configuration identifies SwinForImageClassification with a swin model type. According to the model card, it was trained on ImageNet-1k at 384x384 resolution and introduced in the Swin Transformer paper on hierarchical vision transformers with shifted windows.
According to the model card, the architecture merges image patches in deeper layers and computes self-attention within local windows.
The card describes a base-sized Swin model trained on ImageNet-1k at 384x384 resolution.
The card positions the raw model for image classification and documents a Transformers classification example.
The hub record carries an Apache-2.0 license with transformers, PyTorch, and endpoints-compatible tags.
Source: microsoft/swin-base-patch4-window12-384
Captured: Unknown. Processed: 2026-09-07T19:34:51.836483+00:00.
Swin Transformer (base-sized model) Swin Transformer model trained on ImageNet-1k at resolution 384x384. It was introduced in the paper Swin Transformer: Hierarchical Vision Transformer using Shifted Windows by Liu et al. and first released in this repository . Disclaimer: The team releasing Swin Transformer did not write a model card for this model so this model card has been written by the Hugging Face team. Model description The Swin Transformer is a type of Vision Transformer. It builds hierarchical feature maps by merging image patches (shown in gray) in deeper layers and has linear computation complexity to input image size du…
F001F002F003F004F005F006F007F008F009F010F011F012F013F014