Skip to content

EthenEthenEthen

Open Source Model Profile · indiejoseph

bert-base-cantonese

bert-base-cantonese is a 102.67M-parameter BERT fill-mask model from indiejoseph. According to the model card, it continues pre-training on Cantonese Common Crawl data.

Publisher
indiejoseph
Task
fill-mask
Model type
bert
License
cc-by-4.0
Library
transformers
Publication status
Accepted · not indexed

Model overview

bert-base-cantonese is published by indiejoseph as a BERT fill-mask model. The captured configuration identifies BertForMaskedLM and Safetensors metadata reports 102,674,812 parameters. According to the model card, it is a continued pre-train of bert-base-chinese on a Cantonese Common Crawl dataset with 198M tokens, consistent with hub tags listing google-bert/bert-base-chinese as base model.

Recorded capabilities

Cantonese continued pre-training

According to the model card, it is a continued pre-train of bert-base-chinese on a Cantonese Common Crawl dataset with 198M tokens.

Extended Cantonese characters

According to the model card, the publisher extended 500 additional Chinese characters common in Cantonese.

BERT masked-LM with Transformers

The captured configuration reports BertForMaskedLM with model type bert and 102,674,812 parameters, with Transformers library support.

CC-BY-4.0 licensing

The captured card data lists cc-by-4.0 licensing.

Use cases in the source record

  • Cantonese fill-mask experiments using a bert-base-chinese continuation with extended Cantonese character coverage.
  • Transformers-based masked-language-modeling workflows, including the snapshot-listed hf-inference fill-mask route subject to refresh.

Limitations and unknowns

  • No evaluation results were extracted from this record.
  • According to the model card, intended uses, limitations, and training and evaluation data are recorded as needing more information.
  • No context-window value was extracted.
  • Provider state is historical snapshot data, not independently refreshed current availability.

Source and provenance

Source: indiejoseph/bert-base-cantonese

Captured: Unknown. Processed: 2026-09-07T19:34:47.083334+00:00.

bert-base-cantonese This model is a continue pre-train version of bert-base-chinese on Cantonese Common Crawl dataset with 198m tokens. Model description This model has extended 500 more Chinese characters which very common in Cantonese, such as 冧, 噉, 麪, 笪, 冚, 乸 etc. Intended uses & limitations More information needed Training and evaluation data More information needed Training procedure Training hyperparameters The following hyperparameters were used during training: learning_rate: 0.0001 train_batch_size: 24 eval_batch_size: 8 seed: 42 gradient_accumulation_steps: 8 total_train_batch_size: 192 optimizer: Adam with betas=(0.9,0.99…

F001F002F003F004F005F006F007F008F009F010F011F012F013