Skip to content

EthenEthenEthen

Open Source Model Profile · BAAI

bge-base-zh

bge-base-zh is a 102.27M-parameter BERT embedding model from BAAI. Its model card places it in the Chinese BGE family and points users to bge-base-zh-v1.5.

Publisher
BAAI
Task
feature-extraction
Model type
bert
License
mit
Library
transformers
Publication status
Approved for indexing

Model overview

bge-base-zh is published by BAAI as a feature-extraction model. The captured configuration identifies BertModel, and Safetensors metadata reports 102,268,160 parameters. According to the model card, it is a Chinese BGE base-scale embedding model with a newer v1.5 successor.

Recorded capabilities

Chinese embedding use

The record is a feature-extraction checkpoint, and the card associates BAAI/bge-base-zh with Chinese retrieval phrasing for relevant-passage search.

Documented v1.5 successor

According to the model card, BAAI/bge-base-zh-v1.5 is recommended for its more reasonable similarity distribution with the same usage method.

Training and rerank guidance

According to the model card, BGE pre-training uses retromae plus contrastive learning, with cross-encoder rerankers suggested for top-k refinement.

Use cases in the source record

  • Chinese passage-retrieval embedding workflows that encode queries and passages for similarity search.
  • Retrieve-then-rerank pipelines in which a BGE embedding model returns top candidates for cross-encoder reranking.

Limitations and unknowns

  • No context-window value was extracted from this record.
  • No evaluation results specific to bge-base-zh were extracted; benchmark tables in the card concern related BGE and reranker releases.
  • Provider state is historical snapshot data, not independently refreshed current availability.
  • Rank and performance statements are publisher claims and were not independently verified.

Source and provenance

Source: BAAI/bge-base-zh

Captured: Unknown. Processed: 2026-09-07T19:34:29.234326+00:00.

Recommend switching to newest BAAI/bge-base-zh-v1.5 , which has more reasonable similarity distribution and same method of usage. FlagEmbedding Model List | FAQ | Usage | Evaluation | Train | Contact | Citation | License More details please refer to our Github: FlagEmbedding . English | 中文 FlagEmbedding can map any text to a low-dimensional dense vector which can be used for tasks like retrieval, classification, clustering, or semantic search. And it also can be used in vector databases for LLMs. ************* 🌟 Updates 🌟 ************* 10/12/2023: Release LLM-Embedder , a unified embedding model to support diverse retrieval augmen…

F001F002F003F004F005F006F007F009F010F011F015F016F018F021F024F025F026F027F028F032F034