Bilingual English-Chinese retrieval
According to the model card, the model offers bilingual and crosslingual capability in English and Chinese without requiring a designed instruction.
Open Source Model Profile · maidalun1020
bce-embedding-base_v1 is a bilingual English-Chinese embedding model from maidalun1020 built for RAG retrieval. According to the model card, it is part of BCEmbedding by NetEase Youdao and pairs with a reranker model for precision.
bce-embedding-base_v1 is published by maidalun1020 as a feature-extraction embedding model. The captured configuration identifies XLMRobertaModel with model type xlm-roberta and sentence-transformers support. According to the model card, it provides bilingual and crosslingual capability in English and Chinese and is adapted for RAG across domains such as education, law, finance, medical, and literature.
According to the model card, the model offers bilingual and crosslingual capability in English and Chinese without requiring a designed instruction.
According to the model card, the best practice is recall of top 50-100 passages with this model followed by precision reranking with bce-reranker-base_v1.
According to the model card's model list, bce-embedding-base_v1 is an EmbeddingModel for Chinese and English with 279M parameters.
According to the model card's usage code, embeddings use CLS pooling over a 512-token input followed by normalization.
Source: maidalun1020/bce-embedding-base_v1
Captured: Unknown. Processed: 2026-09-07T19:34:50.901746+00:00.
BCEmbedding: Bilingual and Crosslingual Embedding for RAG 最新、最详细的bce-embedding-base_v1相关信息,请移步(The latest "Updates" should be checked in): GitHub 主要特点(Key Features): 中英双语,以及中英跨语种能力(Bilingual and Crosslingual capability in English and Chinese); RAG优化,适配更多真实业务场景(RAG adaptation for more domains, including Education, Law, Finance, Medical, Literature, FAQ, Textbook, Wikipedia, etc.); 方便集成进langchain和llamaindex(Easy integrations for langchain and llamaindex in BCEmbedding )。 EmbeddingModel 不需要“精心设计”instruction,尽可能召回有用片段。 (No need for "instruction") 最佳实践(Best practice) :embedding召回top50-100片段,reranker对这50-100片段精排,最后取top5-10片段。(1. Get top 5…
F001F002F003F004F005F006F007F009F010F013F015F016