Skip to content

EthenEthenEthen

Open Source Model Profile · thenlper

gte-small

gte-small is a 33.36M-parameter BERT embedding model from thenlper. According to the model card, it belongs to the GTE family for retrieval and similarity work.

Publisher
thenlper
Task
sentence-similarity
Model type
bert
License
mit
Library
sentence-transformers
Publication status
Accepted · not indexed

Model overview

gte-small is published by thenlper as a bert-based sentence-similarity model. The captured configuration identifies BertModel, and Safetensors metadata reports 33360512 parameters. According to the model card, it is a General Text Embeddings model in a family trained by Alibaba DAMO Academy on relevance text pairs.

Recorded capabilities

33.36M BERT scale

Captured config identifies BertModel and Safetensors metadata reports 33360512 parameters.

384-dimension embeddings

According to the model card, the model uses 384 dimensions with a 512 sequence length and 0.07GB size.

MTEB-reported GTE family

The card places gte-small in the GTE family and reports an MTEB average of 61.36 alongside retrieval, STS, reranking, and classification figures.

Use cases in the source record

  • Information retrieval, semantic textual similarity, and text reranking experiments the publisher lists as GTE downstream tasks.

Limitations and unknowns

  • No independently measured evaluation results were extracted; MTEB figures are publisher-reported card claims.
  • Provider state is historical snapshot data, not independently refreshed current availability.

Source and provenance

Source: thenlper/gte-small

Captured: Unknown. Processed: 2026-09-07T19:35:00.253180+00:00.

gte-small General Text Embeddings (GTE) model. Towards General Text Embeddings with Multi-stage Contrastive Learning The GTE models are trained by Alibaba DAMO Academy. They are mainly based on the BERT framework and currently offer three different sizes of models, including GTE-large , GTE-base , and GTE-small . The GTE models are trained on a large-scale corpus of relevance text pairs, covering a wide range of domains and scenarios. This enables the GTE models to be applied to various downstream tasks of text embeddings, including information retrieval , semantic textual similarity , text reranking , etc. Metrics We compared the p…

F001F002F003F004F005F006F007F010F011F012F013F014