Skip to content

EthenEthenEthen

Open Source Model Profile · BSC-LT

MrBERT-es

MrBERT-es is a 150M-parameter bilingual Spanish-English ModernBERT model from BSC-LT. According to the model card, it adapts MrBERT and continues pretraining on 615B balanced tokens.

Publisher
BSC-LT
Task
fill-mask
Model type
modernbert
License
apache-2.0
Library
transformers
Publication status
Accepted · not indexed

Model overview

MrBERT-es is published by BSC-LT as a ModernBERT fill-mask model for Spanish and English. The captured configuration identifies ModernBertForMaskedLM and Safetensors metadata reports 150,295,040 parameters. According to the model card, it adapts MrBERT vocabulary and was continually pretrained on 615 billion balanced English-Spanish tokens, with card data recording apache-2.0.

Recorded capabilities

Bilingual Spanish-English ModernBERT

According to the model card, MrBERT-es is a bilingual Spanish-English foundation model built on ModernBERT through vocabulary adaptation from MrBERT.

615B-token continual pretraining

According to the model card, the model was continually pretrained on 615 billion tokens evenly balanced between English and Spanish, with a masked language modeling objective.

Documented masked-LM usage

According to the model card, the publisher documents loading with AutoModelForMaskedLM and AutoTokenizer alongside fill-mask inference examples.

Compact Transformers record

The captured configuration identifies ModernBertForMaskedLM with model type modernbert and about 150M Safetensors parameters with Transformers support.

Use cases in the source record

  • Spanish and English fill-mask inference using the publisher-documented AutoModelForMaskedLM and AutoTokenizer pattern.
  • Bilingual masked language modeling research on the documented 615B-token English-Spanish continual-pretraining setup.
  • Transformers-based masked-LM experiments using the captured ModernBERT configuration and Safetensors weights.

Limitations and unknowns

  • No evaluation results were extracted from this record.
  • Provider state is historical snapshot data and should be refreshed before being presented as current.
  • No pricing, VRAM, or deployment hardware figures were extracted.

Source and provenance

Source: BSC-LT/MrBERT-es

Captured: Unknown. Processed: 2026-09-07T19:35:34.118394+00:00.

MrBERT-es Model Card MrBERT-es is a new foundational bilingual language model for Spanish and English built on the ModernBERT architecture. It uses vocabulary adaptation from MrBERT , a method that initializes all weights from MrBERT while applying a specialized treatment to the embedding matrix. This treatment carefully handles the differences between the two tokenizers. Following initialization, the model is continually pretrained on a bilingual corpus of 615 billion tokens, evenly balanced between English and Spanish. Technical Description Technical details of the MrBERT-es model. Description Value Model Parameters 150M Tokenizer…

F001F002F003F004F005F006F007F008F009F010F011F012F014F024