Skip to content

EthenEthenEthen

Open Source Model Profile · jhu-clsp

mmBERT-small

mmBERT-small is a ModernBERT multilingual fill-mask encoder from jhu-clsp. According to the model card, the small variant has 140M parameters and was trained on 3T+ tokens across 1800+ languages.

Publisher
jhu-clsp
Task
fill-mask
Model type
modernbert
License
mit
Library
transformers
Publication status
Accepted · not indexed

Model overview

mmBERT-small is published by jhu-clsp as a fill-mask encoder. The captured configuration identifies ModernBertForMaskedLM with model type modernbert and a recorded mit license. According to the model card, it is a modern multilingual encoder trained on more than 3T tokens across 1800+ languages, with the small variant documented at 140M parameters.

Recorded capabilities

Massively multilingual coverage

According to the model card, the model covers more than 1800 languages with a progressive inclusion strategy and a 3T+ token dataset.

ModernBERT foundation

According to the model card, it is built on a ModernBERT foundation with Flash Attention 2 and unpadding, masked language modeling, and bidirectional attention.

Annealed training recipe

According to the model card, training adds languages progressively and reduces the mask ratio from 30% to 15% to 5% across phases.

Small encoder configuration

According to the model card, mmBERT-small has 22 layers, 384 hidden size, 1152 intermediate size, 6 attention heads, 140M total parameters, 8192 max sequence length, and a 256,000 Gemma 2 vocabulary.

Use cases in the source record

  • Multilingual masked language modeling with the documented Transformers masked-LM workflow, including cross-language mask prediction examples.
  • Fine-tuning for retrieval, classification, XNLI, and reranking workflows covered by the card's documented training and fine-tuning examples.

Limitations and unknowns

  • No Safetensors parameter count was extracted from this record; the 140M figure is a publisher model-card claim.
  • Performance and superiority claims in the model card, including comparisons with XLM-R and larger models, are publisher claims and were not independently verified.
  • No benchmark scores were extracted as structured evaluation evidence.
  • Provider state is historical snapshot data and should be refreshed before being presented as current.

Source and provenance

Source: jhu-clsp/mmBERT-small

Captured: Unknown. Processed: 2026-09-07T19:35:57.756076+00:00.

mmBERT: A Modern Multilingual Encoder TL;DR: A state-of-the-art multilingual encoder trained on 3T+ tokens across 1800+ languages, introducing novel techniques for learning low-resource languages during the decay phase. mmBERT is a modern multilingual encoder that significantly outperforms previous generation models like XLM-R on classification, embedding, and retrieval tasks. Built on the ModernBERT architecture with novel multilingual training innovations, mmBERT demonstrates that low-resource languages can be effectively learned during the decay phase of training. It is also significantly faster than any previous multilingual enc…

F001F002F003F004F005F006F009F010F011F014F015F016F017F019F021F022F023F024F025F026F027