Distilled BERT design
According to the model card, this is a distilled version of the BERT base model, described as smaller and faster while pretrained on the same corpus with BERT as teacher.
Open Source Model Profile · distilbert
distilbert-base-uncased is a 67M-parameter DistilBERT fill-mask model from distilbert. According to the model card, it is an uncased distilled BERT base model for masked-language workflows.
distilbert-base-uncased is published by distilbert as a distilbert fill-mask model. The captured configuration identifies DistilBertForMaskedLM and Safetensors metadata reports 66,985,530 parameters. According to the model card, it distills the BERT base model through self-supervised pretraining and treats uppercase and lowercase English identically.
According to the model card, this is a distilled version of the BERT base model, described as smaller and faster while pretrained on the same corpus with BERT as teacher.
According to the model card, pretraining combined masked language modeling with 15% masking and a cosine embedding loss aligning hidden states with the teacher.
According to the model card, the model is aimed at fine-tuning for whole-sentence decisions such as sequence classification, token classification, or question answering.
According to the model card, pretraining used BookCorpus with 11,038 books and English Wikipedia, lowercased and tokenized with WordPiece and a 30,000 vocabulary.
Source: distilbert/distilbert-base-uncased
Captured: Unknown. Processed: 2026-09-07T19:34:43.566230+00:00.
DistilBERT base model (uncased) This model is a distilled version of the BERT base model . It was introduced in this paper . The code for the distillation process can be found here . This model is uncased: it does not make a difference between english and English. Model description DistilBERT is a transformers model, smaller and faster than BERT, which was pretrained on the same corpus in a self-supervised fashion, using the BERT base model as a teacher. This means it was pretrained on the raw texts only, with no humans labelling them in any way (which is why it can use lots of publicly available data) with an automatic process to g…
F001F002F003F004F005F006F007F009F010F011F012F013F014F016F021F031F032F033