Skip to content

EthenEthenEthen

Open Source Model Profile · ibm-granite

granite-embedding-311m-multilingual-r2

granite-embedding-311m-multilingual-r2 is a 312M-parameter ModernBERT multilingual embedding model from IBM Granite with 768-dim vectors and 32k context.

Publisher
ibm-granite
Task
feature-extraction
Model type
modernbert
License
apache-2.0
Library
sentence-transformers
Publication status
Accepted · not indexed

Model overview

granite-embedding-311m-multilingual-r2 is published by ibm-granite as a ModernBERT feature-extraction model for multilingual embeddings. The captured configuration identifies ModernBertModel with model type modernbert, and Safetensors metadata reports 311,664,384 parameters. According to the model card, it upgrades the prior generation with a ModernBERT encoder, 768-dimensional vectors, and 32,768-token context.

Recorded capabilities

312M-parameter multilingual embedder

Captured configuration records ModernBertModel with model type modernbert and about 312M parameters for multilingual text embeddings.

768-dim vectors with 32k context

According to the model card, the model emits 768-dimensional vectors with up to 32,768 tokens of context and CLS pooling.

Publisher-reported retrieval gains

According to the model card, Multilingual MTEB Retrieval reaches 65.2 with a 56.3 average across retrieval benchmarks.

Production-ready deployment record

Hub data records sentence-transformers support while the card documents ONNX and OpenVINO releases with optional Flash Attention 2.

Use cases in the source record

  • Multilingual passage and document retrieval with cosine similarity over 768-dimensional query and passage embeddings.
  • Cross-lingual code retrieval across documented programming languages using the card's retrieval benchmarks as reference.
  • Long-document and multi-passage search within the documented 32,768-token context using ONNX or OpenVINO backends.

Limitations and unknowns

  • According to the model card, quality varies by language, low-resource languages rely on cross-lingual transfer, and synthetic data may add distributional bias; texts beyond 32,768 tokens are truncated.
  • Reported benchmark scores are publisher claims and were not independently verified by Ethen.
  • Provider state is historical snapshot data and should be refreshed before being presented as current.

Source and provenance

Source: ibm-granite/granite-embedding-311m-multilingual-r2

Captured: Unknown. Processed: 2026-09-07T19:34:50.404593+00:00.

Granite-Embedding-311M-Multilingual-R2 Model Summary: Granite-Embedding-311M-Multilingual-R2 is a 311M parameter dense embedding model from the Granite Embeddings collection for high-quality multilingual text embeddings. It produces 768-dimensional vectors with a context length of up to 32,768 tokens. The model supports 200+ languages (based on the multilingual pretraining corpus of the underlying encoder), with enhanced support for 52 languages and programming code that receive explicit retrieval-pair and cross-lingual training. All training data uses permissive, enterprise-friendly licenses, plus IBM-collected and IBM-generated da…

F001F002F003F004F005F006F007F008F009F010F011F012F013F014F017F021F022F024F025F026F027F032F033