RVL-CDIP document lineage
According to the captured model card, the model was pre-trained on IIT-CDIP with 42 million document images and fine-tuned on RVL-CDIP.
Open Source Model Profile · microsoft
dit-base-finetuned-rvlcdip is a BEiT-architecture image-classification model from Microsoft. According to the captured model card, it fine-tunes a Document Image Transformer on RVL-CDIP document images.
dit-base-finetuned-rvlcdip is published by Microsoft as a beit image-classification model. The captured configuration identifies BeitForImageClassification. According to the captured model card, it is a base-sized Document Image Transformer pre-trained on IIT-CDIP and fine-tuned on RVL-CDIP, a 400,000-image grayscale set in 16 classes.
According to the captured model card, the model was pre-trained on IIT-CDIP with 42 million document images and fine-tuned on RVL-CDIP.
According to the captured card, DiT is a BERT-like transformer encoder pre-trained to predict visual tokens from a discrete VAE based on masked patches.
According to the captured card, images are encoded as 16x16 patches with position embeddings, learning representations reusable with a linear classifier.
Source: microsoft/dit-base-finetuned-rvlcdip
Captured: Unknown. Processed: 2026-09-07T19:34:51.917568+00:00.
Document Image Transformer (base-sized model) Document Image Transformer (DiT) model pre-trained on IIT-CDIP (Lewis et al., 2006), a dataset that includes 42 million document images and fine-tuned on RVL-CDIP , a dataset consisting of 400,000 grayscale images in 16 classes, with 25,000 images per class. It was introduced in the paper DiT: Self-supervised Pre-training for Document Image Transformer by Li et al. and first released in this repository . Note that DiT is identical to the architecture of BEiT . Disclaimer: The team releasing DiT did not write a model card for this model so this model card has been written by the Hugging F…
F001F002F003F004F005F007F008F009F010F011F012F013