Whole-word masked pretraining
According to the model card, the model masks all tokens for a word at once, with each masked WordPiece token predicted independently.
Open Source Model Profile · google-bert
bert-large-cased-whole-word-masking is a 335M-parameter cased BERT fill-mask model from google-bert. Its model card documents whole-word masking on English pretraining data.
bert-large-cased-whole-word-masking is published by google-bert as a bert fill-mask model. The captured configuration identifies BertForMaskedLM and Safetensors metadata reports 334661958 parameters. According to the model card, it is a cased English model trained with whole-word masking, where all tokens for a word are masked together.
According to the model card, the model masks all tokens for a word at once, with each masked WordPiece token predicted independently.
Captured Safetensors metadata reports 334661958 parameters; the card lists 24 layers, 1024 hidden size, 16 heads, and 336M nominal parameters.
According to the model card, pretraining used BookCorpus with 11,038 books and English Wikipedia, trained for one million steps on 4 cloud TPUs in Pod configuration.
According to the model card, it is aimed at whole-sentence tasks such as sequence classification, token classification, and question answering, not text generation.
Card data records apache-2.0, and the card reports fine-tuned SQuAD 1.1 92.9/86.7 and MultiNLI 86.46.
Source: google-bert/bert-large-cased-whole-word-masking
Captured: Unknown. Processed: 2026-09-07T19:34:46.014070+00:00.
BERT large model (cased) whole word masking Pretrained model on English language using a masked language modeling (MLM) objective. It was introduced in this paper and first released in this repository . This model is cased: it makes a difference between english and English. Differently to other BERT models, this model was trained with a new technique: Whole Word Masking. In this case, all of the tokens corresponding to a word are masked at once. The overall masking rate remains the same. The training is identical -- each masked WordPiece token is predicted independently. Disclaimer: The team releasing BERT did not write a model card…
F001F002F003F004F005F006F007F009F010F011F012F013F014F015F016F017F018F022F032F033F036F037