Skip to content

EthenEthenEthen

Open Source Model Profile · BAAI

bge-large-en-v1.5

bge-large-en-v1.5 is a 335M-parameter BERT-based English embedding model from BAAI. According to the model card, it belongs to the BGE (BAAI General Embedding) family and is usable for passage retrieval with sentence-transformers.

Publisher
BAAI
Task
feature-extraction
Model type
bert
License
mit
Library
sentence-transformers
Publication status
Accepted · not indexed

Model overview

bge-large-en-v1.5 is published by BAAI as a feature-extraction embedding model. The captured configuration identifies BertModel with model type bert, and Safetensors metadata reports 335,142,400 parameters. The hub lists sentence-transformers support, and the model card describes BGE large models as general embeddings for retrieval, paired with cross-encoder rerankers for top-k re-ranking.

Recorded capabilities

BERT embedding core at 335M

The captured configuration identifies BertModel with model type bert, and Safetensors metadata reports 335,142,400 parameters.

Four documented usage paths

According to the model card, the model is usable with FlagEmbedding, Sentence-Transformers, LangChain, and Hugging Face Transformers.

CLS pooling with normalization

According to the model card, sentence embeddings take the last hidden state of the first ([CLS]) token and are L2-normalized before similarity scoring.

MIT commercial terms

Card data records MIT licensing, and the model card says the released models can be used for commercial purposes free of charge.

Use cases in the source record

  • English passage retrieval and sentence-similarity work with normalized embeddings, CLS pooling, and query instructions for short-query to long-passage tasks.
  • Retrieval pipelines that first retrieve candidate documents with a BGE embedding model and then re-rank the top results with a BGE reranker.
  • ONNX-based feature-extraction serving through optimum, which the card documents as producing output identical to the Transformers path.

Limitations and unknowns

  • No context-window value was extracted from this record.
  • Publisher ranking claims, such as MTEB and C-MTEB placements, are historical model-card statements and were not independently verified by Ethen.
  • Provider state is historical snapshot data and should be refreshed before being presented as current.

Source and provenance

Source: BAAI/bge-large-en-v1.5

Captured: Unknown. Processed: 2026-09-07T19:34:29.695204+00:00.

FlagEmbedding Model List | FAQ | Usage | Evaluation | Train | Contact | Citation | License For more details please refer to our Github: FlagEmbedding . If you are looking for a model that supports more languages, longer texts, and other retrieval methods, you can try using bge-m3 . English | 中文 FlagEmbedding focuses on retrieval-augmented LLMs, consisting of the following projects currently: Long-Context LLM : Activation Beacon Fine-tuning of LM : LM-Cocktail Dense Retrieval : BGE-M3 , LLM Embedder , BGE Embedding Reranker Model : BGE Reranker Benchmark : C-MTEB News 1/30/2024: Release BGE-M3 , a new member to BGE model series! M3 s…

F001F002F003F004F005F006F007F008F010F017F019F026F028F029F030F031F032F033F039