Skip to content

EthenEthenEthen

Open Source Model Profile · BAAI

bge-base-en-v1.5

bge-base-en-v1.5 is a 109M-parameter BERT-family English embedding model from BAAI. According to the model card, it uses a documented retrieval instruction for relevant-passage search.

Publisher
BAAI
Task
feature-extraction
Model type
bert
License
mit
Library
sentence-transformers
Publication status
Approved for indexing

Model overview

bge-base-en-v1.5 is published by BAAI as a feature-extraction embedding model. The captured configuration identifies BertModel with model type bert, and Safetensors metadata reports 109,482,752 parameters. According to the model card, it is the English base-scale BGE release, with queries framed by a documented retrieval instruction.

Recorded capabilities

English retrieval instruction

According to the model card, English queries use a documented instruction for generating representations for relevant-passage search.

Documented multi-framework usage

According to the model card, usage covers FlagEmbedding, Sentence-Transformers, LangChain, and Transformers, with CLS pooling and embedding normalization.

Reranker pairing guidance

According to the model card, BGE rerankers can re-rank top-k documents returned by embedding models, with hard negatives needed for reranker fine-tuning.

Use cases in the source record

  • English passage retrieval in which queries carry the documented retrieval instruction and passages are encoded for similarity search.
  • Retrieve-then-rerank pipelines where a BGE embedding model returns candidates for cross-encoder reranking.

Limitations and unknowns

  • No context-window value was extracted from this record.
  • Rank and benchmark figures in the card are publisher claims and were not independently verified.
  • Provider state is historical snapshot data, not independently refreshed current availability.

Source and provenance

Source: BAAI/bge-base-en-v1.5

Captured: Unknown. Processed: 2026-09-07T19:34:29.205086+00:00.

FlagEmbedding Model List | FAQ | Usage | Evaluation | Train | Contact | Citation | License For more details please refer to our Github: FlagEmbedding . If you are looking for a model that supports more languages, longer texts, and other retrieval methods, you can try using bge-m3 . English | 中文 FlagEmbedding focuses on retrieval-augmented LLMs, consisting of the following projects currently: Long-Context LLM : Activation Beacon Fine-tuning of LM : LM-Cocktail Dense Retrieval : BGE-M3 , LLM Embedder , BGE Embedding Reranker Model : BGE Reranker Benchmark : C-MTEB News 1/30/2024: Release BGE-M3 , a new member to BGE model series! M3 s…

F001F002F003F004F005F006F007F008F010F016F017F019F020F026F028F029F030F031F032