Skip to content

EthenEthenEthen

Open Source Model Profile · maidalun1020

bce-embedding-base_v1

bce-embedding-base_v1 is a bilingual English-Chinese embedding model from maidalun1020 built for RAG retrieval. According to the model card, it is part of BCEmbedding by NetEase Youdao and pairs with a reranker model for precision.

Publisher
maidalun1020
Task
feature-extraction
Model type
xlm-roberta
License
apache-2.0
Library
sentence-transformers
Publication status
Accepted · not indexed

Model overview

bce-embedding-base_v1 is published by maidalun1020 as a feature-extraction embedding model. The captured configuration identifies XLMRobertaModel with model type xlm-roberta and sentence-transformers support. According to the model card, it provides bilingual and crosslingual capability in English and Chinese and is adapted for RAG across domains such as education, law, finance, medical, and literature.

Recorded capabilities

Bilingual English-Chinese retrieval

According to the model card, the model offers bilingual and crosslingual capability in English and Chinese without requiring a designed instruction.

RAG recall plus reranker precision

According to the model card, the best practice is recall of top 50-100 passages with this model followed by precision reranking with bce-reranker-base_v1.

Publisher-listed 279M size

According to the model card's model list, bce-embedding-base_v1 is an EmbeddingModel for Chinese and English with 279M parameters.

CLS pooling with normalization

According to the model card's usage code, embeddings use CLS pooling over a 512-token input followed by normalization.

Use cases in the source record

  • First-stage RAG recall where the model card recommends retrieving the top 50-100 passages with this embedding model before reranking to a final top 5-10.
  • Bilingual English-Chinese semantic search and question-answering using the model's documented semantic-vector generation.
  • RAG pipelines in frameworks the model card names, including langchain and llama_index integrations.

Limitations and unknowns

  • No Safetensors parameter count was extracted; the only size figure is the model card's 279M listing.
  • No independently measured evaluation results were extracted; comparison claims are publisher-reported card statements.
  • No embedding dimensions or context-window value was extracted from this record.
  • Provider state is historical snapshot data, not independently refreshed current availability.

Source and provenance

Source: maidalun1020/bce-embedding-base_v1

Captured: Unknown. Processed: 2026-09-07T19:34:50.901746+00:00.

BCEmbedding: Bilingual and Crosslingual Embedding for RAG 最新、最详细的bce-embedding-base_v1相关信息,请移步(The latest "Updates" should be checked in): GitHub 主要特点(Key Features): 中英双语,以及中英跨语种能力(Bilingual and Crosslingual capability in English and Chinese); RAG优化,适配更多真实业务场景(RAG adaptation for more domains, including Education, Law, Finance, Medical, Literature, FAQ, Textbook, Wikipedia, etc.); 方便集成进langchain和llamaindex(Easy integrations for langchain and llamaindex in BCEmbedding )。 EmbeddingModel 不需要“精心设计”instruction,尽可能召回有用片段。 (No need for "instruction") 最佳实践(Best practice) :embedding召回top50-100片段,reranker对这50-100片段精排,最后取top5-10片段。(1. Get top 5…

F001F002F003F004F005F006F007F009F010F013F015F016