Massively multilingual coverage
According to the model card, the model covers more than 1800 languages with a progressive inclusion strategy and a 3T+ token dataset.
Open Source Model Profile · jhu-clsp
mmBERT-small is a ModernBERT multilingual fill-mask encoder from jhu-clsp. According to the model card, the small variant has 140M parameters and was trained on 3T+ tokens across 1800+ languages.
mmBERT-small is published by jhu-clsp as a fill-mask encoder. The captured configuration identifies ModernBertForMaskedLM with model type modernbert and a recorded mit license. According to the model card, it is a modern multilingual encoder trained on more than 3T tokens across 1800+ languages, with the small variant documented at 140M parameters.
According to the model card, the model covers more than 1800 languages with a progressive inclusion strategy and a 3T+ token dataset.
According to the model card, it is built on a ModernBERT foundation with Flash Attention 2 and unpadding, masked language modeling, and bidirectional attention.
According to the model card, training adds languages progressively and reduces the mask ratio from 30% to 15% to 5% across phases.
According to the model card, mmBERT-small has 22 layers, 384 hidden size, 1152 intermediate size, 6 attention heads, 140M total parameters, 8192 max sequence length, and a 256,000 Gemma 2 vocabulary.
Source: jhu-clsp/mmBERT-small
Captured: Unknown. Processed: 2026-09-07T19:35:57.756076+00:00.
mmBERT: A Modern Multilingual Encoder TL;DR: A state-of-the-art multilingual encoder trained on 3T+ tokens across 1800+ languages, introducing novel techniques for learning low-resource languages during the decay phase. mmBERT is a modern multilingual encoder that significantly outperforms previous generation models like XLM-R on classification, embedding, and retrieval tasks. Built on the ModernBERT architecture with novel multilingual training innovations, mmBERT demonstrates that low-resource languages can be effectively learned during the decay phase of training. It is also significantly faster than any previous multilingual enc…
F001F002F003F004F005F006F009F010F011F014F015F016F017F019F021F022F023F024F025F026F027