Skip to content

EthenEthenEthen

Open Source Model Profile · thenlper

gte-large-zh

gte-large-zh is a 325.5M-parameter BERT embedding model from thenlper. According to the model card, it is a Chinese General Text Embeddings model with 1024 dimensions and a 512-token maximum sequence length.

Publisher
thenlper
Task
sentence-similarity
Model type
bert
License
mit
Library
sentence-transformers
Publication status
Accepted · not indexed

Model overview

gte-large-zh is published by thenlper as a sentence-similarity embedding model. The captured configuration identifies BertModel with a bert model type, and Safetensors metadata reports 325522944 parameters. According to the model card, it belongs to the GTE family trained by Alibaba DAMO Academy on large-scale relevance text pairs for retrieval, similarity, and reranking.

Recorded capabilities

Chinese GTE embedding size

According to the model card, GTE-large-zh uses 1024 embedding dimensions, 512 maximum sequence length, and about 0.67GB model size.

BERT sentence-transformers base

Captured config identifies BertModel with 325522944 parameters and a sentence-transformers library tag.

Publisher-reported CMTEB table

According to the model card, gte-large-zh averages 66.72 across 35 CMTEB datasets, with 72.49 retrieval and 57.82 STS alongside per-task comparisons to Stella, BGE, and Piccolo variants.

Multi-stage contrastive training description

The publisher describes GTE models as trained with multi-stage contrastive learning over broad domains and scenarios.

Use cases in the source record

  • Chinese text retrieval, semantic similarity, and text reranking with sentence-transformers embeddings.
  • Embedding comparisons that use the publisher-reported CMTEB table for Chinese retrieval and classification tasks.

Limitations and unknowns

  • Benchmark figures are publisher-reported CMTEB values from the model card and have not been independently verified by Ethen.
  • No local hardware requirement is provided in the current evidence.
  • Provider state is historical snapshot data, not independently refreshed current availability.

Source and provenance

Source: thenlper/gte-large-zh

Captured: Unknown. Processed: 2026-09-07T19:35:00.093290+00:00.

gte-large-zh General Text Embeddings (GTE) model. Towards General Text Embeddings with Multi-stage Contrastive Learning The GTE models are trained by Alibaba DAMO Academy. They are mainly based on the BERT framework and currently offer different sizes of models for both Chinese and English Languages. The GTE models are trained on a large-scale corpus of relevance text pairs, covering a wide range of domains and scenarios. This enables the GTE models to be applied to various downstream tasks of text embeddings, including information retrieval , semantic textual similarity , text reranking , etc. Model List Models Language Max Sequenc…

F001F002F003F004F005F006F007F008F010F011F012F013F014