Cantonese continued pre-training
According to the model card, it is a continued pre-train of bert-base-chinese on a Cantonese Common Crawl dataset with 198M tokens.
Open Source Model Profile · indiejoseph
bert-base-cantonese is a 102.67M-parameter BERT fill-mask model from indiejoseph. According to the model card, it continues pre-training on Cantonese Common Crawl data.
bert-base-cantonese is published by indiejoseph as a BERT fill-mask model. The captured configuration identifies BertForMaskedLM and Safetensors metadata reports 102,674,812 parameters. According to the model card, it is a continued pre-train of bert-base-chinese on a Cantonese Common Crawl dataset with 198M tokens, consistent with hub tags listing google-bert/bert-base-chinese as base model.
According to the model card, it is a continued pre-train of bert-base-chinese on a Cantonese Common Crawl dataset with 198M tokens.
According to the model card, the publisher extended 500 additional Chinese characters common in Cantonese.
The captured configuration reports BertForMaskedLM with model type bert and 102,674,812 parameters, with Transformers library support.
The captured card data lists cc-by-4.0 licensing.
Source: indiejoseph/bert-base-cantonese
Captured: Unknown. Processed: 2026-09-07T19:34:47.083334+00:00.
bert-base-cantonese This model is a continue pre-train version of bert-base-chinese on Cantonese Common Crawl dataset with 198m tokens. Model description This model has extended 500 more Chinese characters which very common in Cantonese, such as 冧, 噉, 麪, 笪, 冚, 乸 etc. Intended uses & limitations More information needed Training and evaluation data More information needed Training procedure Training hyperparameters The following hyperparameters were used during training: learning_rate: 0.0001 train_batch_size: 24 eval_batch_size: 8 seed: 42 gradient_accumulation_steps: 8 total_train_batch_size: 192 optimizer: Adam with betas=(0.9,0.99…
F001F002F003F004F005F006F007F008F009F010F011F012F013