Skip to content

EthenEthenEthen

Open Source Model Profile · microsoft

swinv2-base-patch4-window8-256

swinv2-base-patch4-window8-256 is a Swin V2 base image classifier from microsoft. According to the model card, it was pre-trained on ImageNet-1k at 256x256 resolution.

Publisher
microsoft
Task
image-classification
Model type
swinv2
License
apache-2.0
Library
transformers
Publication status
Accepted · not indexed

Model overview

swinv2-base-patch4-window8-256 is published by microsoft as an image-classification model. The captured configuration identifies Swinv2ForImageClassification with model type swinv2. According to the model card, it is a base-sized Swin Transformer v2 model pre-trained on ImageNet-1k at 256x256.

Recorded capabilities

Swin V2 base ImageNet-1k training

According to the model card, this base-sized model was pre-trained on ImageNet-1k at 256x256 resolution.

Hierarchical vision-transformer design

According to the model card, Swin builds hierarchical feature maps by merging patches and computes self-attention within local windows for linear complexity to image size.

Three documented v2 improvements

According to the model card, v2 introduces residual-post-norm with cosine attention for stability, log-spaced position bias for resolution transfer, and SimMIM to reduce labeled-data needs.

Documented Transformers classification workflow

According to the model card, the model classifies COCO 2017 images into 1,000 ImageNet classes with AutoImageProcessor and AutoModelForImageClassification.

Use cases in the source record

  • Raw image classification into 1,000 ImageNet classes using the documented AutoImageProcessor and classification-model workflow.
  • Backbone experiments for classification and dense-recognition setups where a hierarchical window-attention transformer is relevant, as described by the card.

Limitations and unknowns

  • No parameter count was extracted from this record.
  • No evaluation results were extracted from this record.
  • The card states it was written by the Hugging Face team because the releasing team did not provide one, so its descriptions are secondary documentation rather than publisher-authored claims.
  • Provider state is historical snapshot data, not independently refreshed current availability.

Source and provenance

Source: microsoft/swinv2-base-patch4-window8-256

Captured: Unknown. Processed: 2026-09-07T19:34:51.912900+00:00.

Swin Transformer v2 (base-sized model) Swin Transformer v2 model pre-trained on ImageNet-1k at resolution 256x256. It was introduced in the paper Swin Transformer V2: Scaling Up Capacity and Resolution by Liu et al. and first released in this repository . Disclaimer: The team releasing Swin Transformer v2 did not write a model card for this model so this model card has been written by the Hugging Face team. Model description The Swin Transformer is a type of Vision Transformer. It builds hierarchical feature maps by merging image patches (shown in gray) in deeper layers and has linear computation complexity to input image size due t…

F001F002F003F004F005F006F009F010F011F012F013F014