Distilled BERT architecture
The captured configuration identifies BertModel with model type bert and Transformers support.
Open Source Model Profile · microsoft
xtremedistil-l6-h384-uncased is a distilled BERT-family text-classification release from microsoft. According to the model card, the l6-h384 checkpoint reports 22 million parameters.
xtremedistil-l6-h384-uncased is published by microsoft as a text-classification release. The captured configuration identifies BertModel with model type bert. According to the model card, it is the l6-h384 XtremeDistilTransformers checkpoint with 6 layers, 384 hidden size, and a publisher-stated 22 million parameters.
The captured configuration identifies BertModel with model type bert and Transformers support.
According to the model card, the model uses task transfer with multi-task distillation techniques from the XtremeDistil and MiniLM lines of work.
According to the model card, this checkpoint has 6 layers, 384 hidden size, 12 attention heads, 22 million parameters, and 5.3x speedup over BERT-base.
According to the model card, results are presented on the GLUE dev set and SQuAD-v2, with two sibling checkpoints named for comparison.
Source: microsoft/xtremedistil-l6-h384-uncased
Captured: Unknown. Processed: 2026-09-07T19:34:51.734811+00:00.
XtremeDistilTransformers for Distilling Massive Neural Networks XtremeDistilTransformers is a distilled task-agnostic transformer model that leverages task transfer for learning a small universal model that can be applied to arbitrary tasks and languages as outlined in the paper XtremeDistilTransformers: Task Transfer for Task-agnostic Distillation . We leverage task transfer combined with multi-task distillation techniques from the papers XtremeDistil: Multi-stage Distillation for Massive Multilingual Models and MiniLM: Deep Self-Attention Distillation for Task-Agnostic Compression of Pre-Trained Transformers with the following Git…
F001F002F003F004F005F006F007F008F009F010F011F012F013