Skip to content

EthenEthenEthen

Open Source Model Profile · lightonai

modernbert-embed-large

modernbert-embed-large is a 394.78M-parameter LightOnAI embedding model trained from ModernBERT-large. Its model card documents Nomic Embed fine-tuning and Matryoshka dimension support.

Publisher
lightonai
Task
sentence-similarity
Model type
modernbert
License
apache-2.0
Library
sentence-transformers
Publication status
Approved for indexing

Model overview

modernbert-embed-large is published by LightOnAI as a sentence-similarity embedding model. The captured configuration identifies ModernBertModel with model type modernbert, and Safetensors metadata reports 394,781,696 parameters. According to the model card, it is an embedding model trained from ModernBERT-large, and hub tags record that base.

Recorded capabilities

ModernBERT-large embedding tune

According to the model card, this fine-tune brings ModernBERT advances to embeddings, since the base masked-language model cannot do retrieval without further tuning.

Nomic Embed training recipe

The model card says the model was fine-tuned on Nomic Embed weakly-supervised and supervised datasets through a multi-stage contrastive pipeline.

Matryoshka truncation

According to the model card, 256-dimension Matryoshka representations reduce memory with minimal loss, via truncate_dim or pre-normalization slicing.

Prefix-required inputs

The model card says inputs require search_query and search_document prefixes, following the Nomic Embed instruction pattern.

Use cases in the source record

  • Sentence-similarity and retrieval work with prefixed search queries and documents through sentence-transformers or transformers.
  • Memory-efficient embedding with 256-dimension Matryoshka truncation for large-scale similarity indexes.

Limitations and unknowns

  • No context-window value was extracted from this record.
  • Provider state is historical snapshot data and should be refreshed before being presented as current.
  • Evaluation-table figures are publisher-reported model-card values and were not independently verified by Ethen.

Source and provenance

Source: lightonai/modernbert-embed-large

Captured: Unknown. Processed: 2026-09-07T19:35:25.960244+00:00.

ModernBERT-embed-large ModernBERT-embed-large is an embedding model trained from ModernBERT-large , bringing the new advances of ModernBERT to embeddings! Indeed, ModernBERT is a base model trained for Masked Language Modeling and can not directly be used to perform tasks such as retrieval without further fine-tuning. ModernBERT-embed-large is fine-tuned on the Nomic Embed weakly-supervised and supervised datasets and also supports Matryoshka Representation Learning dimensions of 256 to reduce memory with minimal performance loss. Performance Model Dimensions Average (56) Classification (12) Clustering (11) Pair Classification (3) R…

F001F002F003F004F005F006F007F009F010F011F012F013F014F015F016F017F018