Skip to content

EthenEthenEthen

Open Source Model Profile · microsoft

dit-base-finetuned-rvlcdip

dit-base-finetuned-rvlcdip is a BEiT-architecture image-classification model from Microsoft. According to the captured model card, it fine-tunes a Document Image Transformer on RVL-CDIP document images.

Publisher
microsoft
Task
image-classification
Model type
beit
License
Unknown
Library
transformers
Publication status
Accepted · not indexed

Model overview

dit-base-finetuned-rvlcdip is published by Microsoft as a beit image-classification model. The captured configuration identifies BeitForImageClassification. According to the captured model card, it is a base-sized Document Image Transformer pre-trained on IIT-CDIP and fine-tuned on RVL-CDIP, a 400,000-image grayscale set in 16 classes.

Recorded capabilities

RVL-CDIP document lineage

According to the captured model card, the model was pre-trained on IIT-CDIP with 42 million document images and fine-tuned on RVL-CDIP.

Self-supervised encoder design

According to the captured card, DiT is a BERT-like transformer encoder pre-trained to predict visual tokens from a discrete VAE based on masked patches.

Patch-based representation

According to the captured card, images are encoded as 16x16 patches with position embeddings, learning representations reusable with a linear classifier.

Use cases in the source record

  • Document-image classification workflows consistent with the captured image-classification tag and publisher-described RVL-CDIP fine-tune.

Limitations and unknowns

  • The captured card includes a disclaimer that the releasing team did not write the card and it was written by the Hugging Face team.
  • No license, parameter count, or evaluation results were extracted from this record.
  • No pricing, hardware-requirement, or context-window values were extracted.
  • Provider state is historical snapshot data, not independently refreshed current availability.

Source and provenance

Source: microsoft/dit-base-finetuned-rvlcdip

Captured: Unknown. Processed: 2026-09-07T19:34:51.917568+00:00.

Document Image Transformer (base-sized model) Document Image Transformer (DiT) model pre-trained on IIT-CDIP (Lewis et al., 2006), a dataset that includes 42 million document images and fine-tuned on RVL-CDIP , a dataset consisting of 400,000 grayscale images in 16 classes, with 25,000 images per class. It was introduced in the paper DiT: Self-supervised Pre-training for Document Image Transformer by Li et al. and first released in this repository . Note that DiT is identical to the architecture of BEiT . Disclaimer: The team releasing DiT did not write a model card for this model so this model card has been written by the Hugging F…

F001F002F003F004F005F007F008F009F010F011F012F013