Skip to content

EthenEthenEthen

Open Source Model Profile · microsoft

xtremedistil-l6-h384-uncased

xtremedistil-l6-h384-uncased is a distilled BERT-family text-classification release from microsoft. According to the model card, the l6-h384 checkpoint reports 22 million parameters.

Publisher
microsoft
Task
text-classification
Model type
bert
License
mit
Library
transformers
Publication status
Accepted · not indexed

Model overview

xtremedistil-l6-h384-uncased is published by microsoft as a text-classification release. The captured configuration identifies BertModel with model type bert. According to the model card, it is the l6-h384 XtremeDistilTransformers checkpoint with 6 layers, 384 hidden size, and a publisher-stated 22 million parameters.

Recorded capabilities

Distilled BERT architecture

The captured configuration identifies BertModel with model type bert and Transformers support.

Task-agnostic distillation

According to the model card, the model uses task transfer with multi-task distillation techniques from the XtremeDistil and MiniLM lines of work.

L6-H384 size and speed claim

According to the model card, this checkpoint has 6 layers, 384 hidden size, 12 attention heads, 22 million parameters, and 5.3x speedup over BERT-base.

GLUE and SQuAD-v2 reporting context

According to the model card, results are presented on the GLUE dev set and SQuAD-v2, with two sibling checkpoints named for comparison.

Use cases in the source record

  • Lightweight text-classification and feature-extraction experiments consistent with the captured pipeline tag and task-agnostic distillation design.
  • Efficiency-sensitive deployment research comparing the l6-h384 checkpoint against the publisher-named sibling configurations.

Limitations and unknowns

  • No parameter count was extracted in structured form; the 22-million figure is a publisher-card claim.
  • No evaluation values were extracted in structured form, although the card mentions GLUE and SQuAD-v2 tables.
  • No context-window value was extracted from this record.
  • Provider state is historical snapshot data and should be refreshed before being presented as current.

Source and provenance

Source: microsoft/xtremedistil-l6-h384-uncased

Captured: Unknown. Processed: 2026-09-07T19:34:51.734811+00:00.

XtremeDistilTransformers for Distilling Massive Neural Networks XtremeDistilTransformers is a distilled task-agnostic transformer model that leverages task transfer for learning a small universal model that can be applied to arbitrary tasks and languages as outlined in the paper XtremeDistilTransformers: Task Transfer for Task-agnostic Distillation . We leverage task transfer combined with multi-task distillation techniques from the papers XtremeDistil: Multi-stage Distillation for Massive Multilingual Models and MiniLM: Deep Self-Attention Distillation for Task-Agnostic Compression of Pre-Trained Transformers with the following Git…

F001F002F003F004F005F006F007F008F009F010F011F012F013