Skip to content

EthenEthenEthen

Open Source Model Profile · microsoft

swin-base-patch4-window12-384

swin-base-patch4-window12-384 is a microsoft Swin Transformer model for image classification. The card describes ImageNet-1k training at 384x384 resolution under Apache-2.0.

Publisher
microsoft
Task
image-classification
Model type
swin
License
apache-2.0
Library
transformers
Publication status
Accepted · not indexed

Model overview

The model is published by microsoft as a base-sized Swin Transformer for image classification. The captured configuration identifies SwinForImageClassification with a swin model type. According to the model card, it was trained on ImageNet-1k at 384x384 resolution and introduced in the Swin Transformer paper on hierarchical vision transformers with shifted windows.

Recorded capabilities

Swin hierarchical windows

According to the model card, the architecture merges image patches in deeper layers and computes self-attention within local windows.

ImageNet-1k at 384 resolution

The card describes a base-sized Swin model trained on ImageNet-1k at 384x384 resolution.

Classification usage path

The card positions the raw model for image classification and documents a Transformers classification example.

Apache-2.0 record

The hub record carries an Apache-2.0 license with transformers, PyTorch, and endpoints-compatible tags.

Use cases in the source record

  • Raw-model image classification into the 1,000 ImageNet classes using the card's documented Transformers workflow.
  • Vision-backbone research building on the card's description of hierarchical feature maps for classification and dense recognition tasks.

Limitations and unknowns

  • No parameter count or evaluation metrics were extracted from this record.
  • No context-window, hardware requirement, or quantization detail applies to this vision record beyond the extracted fields.
  • Training and architecture narrative comes from the publisher-supplied card text and was not independently verified by Ethen.
  • Provider state is historical snapshot data, not independently refreshed current availability.

Source and provenance

Source: microsoft/swin-base-patch4-window12-384

Captured: Unknown. Processed: 2026-09-07T19:34:51.836483+00:00.

Swin Transformer (base-sized model) Swin Transformer model trained on ImageNet-1k at resolution 384x384. It was introduced in the paper Swin Transformer: Hierarchical Vision Transformer using Shifted Windows by Liu et al. and first released in this repository . Disclaimer: The team releasing Swin Transformer did not write a model card for this model so this model card has been written by the Hugging Face team. Model description The Swin Transformer is a type of Vision Transformer. It builds hierarchical feature maps by merging image patches (shown in gray) in deeper layers and has linear computation complexity to input image size du…

F001F002F003F004F005F006F007F008F009F010F011F012F013F014