Skip to content

EthenEthenEthen

Open Source Model Profile · thenlper

gte-base

gte-base is a 109M-parameter BERT-family sentence-similarity embedding model from thenlper. The model card describes it as a General Text Embeddings (GTE) model.

Publisher
thenlper
Task
sentence-similarity
Model type
bert
License
mit
Library
sentence-transformers
Publication status
Accepted · not indexed

Model overview

gte-base is published by thenlper as a sentence-similarity embedding model. The captured configuration identifies BertModel with a bert model type and Safetensors metadata reports 109482752 parameters. According to the model card, it is a General Text Embeddings (GTE) model in a family described as trained by Alibaba DAMO Academy.

Recorded capabilities

BERT embedding configuration

Captured configuration identifies BertModel with a bert model type, about 109M parameters, and a sentence-transformers library tag.

768-dimension embeddings

According to the model card, gte-base uses 768 dimensions with 512 sequence length and a 0.22 GB model size.

GTE training description

According to the model card, the GTE models were trained by Alibaba DAMO Academy on a large-scale corpus of relevance text pairs for retrieval, semantic textual similarity, and reranking uses.

Publisher-reported MTEB table

The model card points to the MTEB leaderboard and includes a comparison table reporting 62.39 average for gte-base.

Use cases in the source record

  • Text-embedding workflows for information retrieval using sentence-similarity representations.
  • According to the model card, semantic textual similarity and text reranking tasks covered by the GTE training description.

Limitations and unknowns

  • No independent Ethen evaluation results were extracted from this record; benchmark figures come from the publisher model card.
  • No context-window value was extracted from this record.
  • Provider state is historical snapshot data, not independently refreshed current availability.
  • Training and origin information comes from the publisher model card and has not been independently verified by Ethen.

Source and provenance

Source: thenlper/gte-base

Captured: Unknown. Processed: 2026-09-07T19:34:59.703633+00:00.

gte-base General Text Embeddings (GTE) model. Towards General Text Embeddings with Multi-stage Contrastive Learning The GTE models are trained by Alibaba DAMO Academy. They are mainly based on the BERT framework and currently offer three different sizes of models, including GTE-large , GTE-base , and GTE-small . The GTE models are trained on a large-scale corpus of relevance text pairs, covering a wide range of domains and scenarios. This enables the GTE models to be applied to various downstream tasks of text embeddings, including information retrieval , semantic textual similarity , text reranking , etc. Metrics We compared the pe…

F001F002F003F004F005F006F007F008F010F011F012F013F014