Skip to content

EthenEthenEthen

Open Source Model Profile · google-bert

bert-large-cased-whole-word-masking

bert-large-cased-whole-word-masking is a 335M-parameter cased BERT fill-mask model from google-bert. Its model card documents whole-word masking on English pretraining data.

Publisher
google-bert
Task
fill-mask
Model type
bert
License
apache-2.0
Library
transformers
Publication status
Accepted · not indexed

Model overview

bert-large-cased-whole-word-masking is published by google-bert as a bert fill-mask model. The captured configuration identifies BertForMaskedLM and Safetensors metadata reports 334661958 parameters. According to the model card, it is a cased English model trained with whole-word masking, where all tokens for a word are masked together.

Recorded capabilities

Whole-word masked pretraining

According to the model card, the model masks all tokens for a word at once, with each masked WordPiece token predicted independently.

BERT-large 335M configuration

Captured Safetensors metadata reports 334661958 parameters; the card lists 24 layers, 1024 hidden size, 16 heads, and 336M nominal parameters.

BookCorpus and Wikipedia pretraining

According to the model card, pretraining used BookCorpus with 11,038 books and English Wikipedia, trained for one million steps on 4 cloud TPUs in Pod configuration.

Fine-tune-first design

According to the model card, it is aimed at whole-sentence tasks such as sequence classification, token classification, and question answering, not text generation.

Apache-2.0 with reported evals

Card data records apache-2.0, and the card reports fine-tuned SQuAD 1.1 92.9/86.7 and MultiNLI 86.46.

Use cases in the source record

  • Fine-tuning for whole-sentence decisions such as sequence classification, token classification, or question answering.
  • Fill-mask inference and feature extraction with Transformers in PyTorch or TensorFlow as shown in the card's usage examples.

Limitations and unknowns

  • According to the model card, the model can produce biased predictions and this bias affects fine-tuned versions.
  • No context-window value was extracted from this record.
  • Provider state is historical snapshot data, not independently refreshed current availability.
  • According to the card disclaimer, this model card was written by the Hugging Face team, not the original BERT team.

Source and provenance

Source: google-bert/bert-large-cased-whole-word-masking

Captured: Unknown. Processed: 2026-09-07T19:34:46.014070+00:00.

BERT large model (cased) whole word masking Pretrained model on English language using a masked language modeling (MLM) objective. It was introduced in this paper and first released in this repository . This model is cased: it makes a difference between english and English. Differently to other BERT models, this model was trained with a new technique: Whole Word Masking. In this case, all of the tokens corresponding to a word are masked at once. The overall masking rate remains the same. The training is identical -- each masked WordPiece token is predicted independently. Disclaimer: The team releasing BERT did not write a model card…

F001F002F003F004F005F006F007F009F010F011F012F013F014F015F016F017F018F022F032F033F036F037