Skip to content

EthenEthenEthen

Open Source Model Profile · qihoo360

Zhinao-ChineseModernBert-Embedding

Zhinao-ChineseModernBert-Embedding is a 227M-parameter Chinese semantic-embedding model from qihoo360. According to the model card, it supports 512-token inputs with 768-dimension vectors for retrieval and similarity workloads.

Publisher
qihoo360
Task
sentence-similarity
Model type
modernbert
License
apache-2.0
Library
sentence-transformers
Publication status
Accepted · not indexed

Model overview

Zhinao-ChineseModernBert-Embedding is published by qihoo360 as a sentence-similarity embedding model. Captured Safetensors metadata reports 227,018,496 parameters, about 227M. According to the model card, it is the specialized embedding member of the Zhinao-ChineseModernBert series, optimized for semantic retrieval, vector databases, and RAG augmentation on top of a ModernBert base.

Recorded capabilities

768-dimension CLS embeddings

According to the model card, the Sentence Transformer outputs 768 dimensions with CLS pooling, normalization, and cosine similarity over up to 512 tokens.

ModernBert plus Qwen2 tokenizer

According to the model card, the series uses the ModernBert architecture with a Qwen2Tokenizer to reduce out-of-vocabulary rates on Chinese and mixed text.

Two-stage embedding training

According to the model card, training runs retrieval contrastive pre-training followed by multi-task fine-tuning on CMTEB and MTEB task data.

Publisher-reported 1T-token base

According to the model card, the base was pre-trained from scratch on 1T high-quality Chinese and English tokens with over 65% Chinese content.

Use cases in the source record

  • Chinese semantic retrieval and vector-database indexing using the documented 768-dimension normalized embeddings.
  • RAG retrieval augmentation and text similarity, clustering, and reranking flows described in the model card.
  • Instruction-prompted encoding with the card's documented retrieval prompts, such as the Chinese search-query instruction.

Limitations and unknowns

  • No independently measured evaluation results were extracted; benchmark leadership statements are publisher-reported card claims.
  • No VRAM, hardware, or pricing detail was extracted from this record.
  • The model card's performance comparisons against larger models come from the publisher and were not independently verified.
  • Provider state is historical snapshot data, not independently refreshed current availability.

Source and provenance

Source: qihoo360/Zhinao-ChineseModernBert-Embedding

Captured: Unknown. Processed: 2026-09-07T19:36:00.702684+00:00.

Zhinao-ChineseModernBert: 面向高吞吐低内存场景的中文基座与向量嵌入模型 项目简介 Zhinao-ChineseModernBert系列是针对 高推理速度要求、严苛内存限制 的工业级场景,从头预训练的中文Base级基座模型与语义嵌入模型。本系列基于ModernBert高效架构与Qwen2Tokenizer分词器,依托超大规模中英文语料完成全流程预训练,在保持Base级参数量(除Embedding外约100M参数)轻量化优势的同时,实现了对同量级模型的全面超越,甚至性能优于更大参数量的主流模型,为中文NLP理解任务、语义检索、向量数据库、RAG检索增强等场景提供高性价比的开箱即用解决方案。 本项目包含两个核心模型: Zhinao-ChineseModernBert :通用中文理解基座,基于两阶段掩码语言建模(MLM)预训练,支持最长1536序列长度,适配各类中文NLU下游任务。 Zhinao-ChineseModernBert-Embedding :专业中文语义嵌入模型,在基座的基础上做两阶段Embedding训练,支持最长512序列长度,专为语义检索、向量表征、相似度计算等场景深度优化。 核心亮点 1. 高效架构+先进分词体系,兼顾速度与泛化性 采用 ModernBert 高效Transformer架构从头预训练,针对高吞吐推理、长文本处理场景做了深度优化,相比传统Bert架构实现显著的推理速度提升,内存占用更友好。 适配 Qwen2T…

F001F002F003F004F005F006F007F010F011F012F013F014F015F016