Skip to content

EthenEthenEthen

Open Source Model Profile · BAAI

bge-multilingual-gemma2

Bge-multilingual-gemma2 is a 9.24B-parameter Gemma2 multilingual embedding model from BAAI. According to the model card, it is based on google/gemma-2-9b and covers retrieval, classification, and clustering tasks.

Publisher
BAAI
Task
feature-extraction
Model type
gemma2
License
gemma
Library
sentence-transformers
Publication status
Accepted · not indexed

Model overview

Bge-multilingual-gemma2 is published by BAAI as a feature-extraction embedding model. Captured configuration identifies Gemma2Model with model type gemma2, and Safetensors metadata reports 9,241,713,152 parameters. According to the model card, it is an LLM-based multilingual embedding model based on google/gemma-2-9b, trained across languages including English, Chinese, Japanese, Korean, and French.

Recorded capabilities

Multilingual embedding use

The record is a feature-extraction checkpoint with the sentence-transformers library, and the publisher describes multilingual training across retrieval, classification, and clustering tasks.

Gemma2 architecture at 9.24B scale

Captured configuration identifies Gemma2Model with model type gemma2, and Safetensors metadata reports 9,241,713,152 parameters.

Documented FlagEmbedding and Sentence-Transformers workflows

According to the model card, the model loads through FlagLLMModel with a retrieval instruction and through SentenceTransformer with optional float16 precision.

Publisher-reported multilingual benchmark results

According to the model card, the model shows publisher-described leading results on MIRACL, MTEB-pl, and MTEB-fr, plus strong results on MTEB, C-MTEB, and AIR-Bench; no independent scores were extracted.

Use cases in the source record

  • Multilingual passage retrieval with an instruction such as retrieving relevant passages for a web search query, using the documented FlagEmbedding or Sentence-Transformers workflows.
  • Multilingual classification and clustering workflows using text embeddings, which the publisher lists among the model's trained task types.

Limitations and unknowns

  • No evaluation results were extracted from this record beyond publisher-described benchmark claims.
  • No embedding dimension, context length, or quantization value was extracted.
  • Provider state is historical snapshot data, not independently refreshed current availability.

Source and provenance

Source: BAAI/bge-multilingual-gemma2

Captured: Unknown. Processed: 2026-09-07T19:34:30.247510+00:00.

FlagEmbedding For more details please refer to our Github: FlagEmbedding . BGE-Multilingual-Gemma2 is a LLM-based multilingual embedding model. It is trained on a diverse range of languages and tasks based on google/gemma-2-9b . BGE-Multilingual-Gemma2 primarily demonstrates the following advancements: Diverse training data: The model's training data spans a broad range of languages, including English, Chinese, Japanese, Korean, French, and more.Additionally, the data covers a variety of task types, such as retrieval, classification, and clustering. Outstanding performance: The model exhibits state-of-the-art (SOTA) results on multi…

F001F002F003F004F005F006F007F008F009F010F012F013F016F017