Skip to content

EthenEthenEthen

Open Source Model Profile · intfloat

e5-small

e5-small is a 33.36M-parameter BERT sentence-similarity model from intfloat. Its model card documents 12 layers, 384-dim embeddings, and query-passage prefixes.

Publisher
intfloat
Task
sentence-similarity
Model type
bert
License
mit
Library
sentence-transformers
Publication status
Accepted · not indexed

Model overview

e5-small is published by intfloat as a sentence-similarity model. The captured configuration identifies BertModel with model type bert, and Safetensors metadata reports 33,360,512 parameters. According to the model card, it has 12 layers with an embedding size of 384 and was trained with weakly-supervised contrastive pre-training.

Recorded capabilities

Documented 384-dim embeddings

According to the model card, the model has 12 layers and the embedding size is 384.

Required query-passage prefixes

According to the model card, inputs should add query: and passage: prefixes, otherwise performance degrades.

Sentence-Transformers support

The model card documents SentenceTransformer loading for intfloat/e5-small with normalized embeddings, and hub tags list sentence-transformers.

MIT license

Card data records MIT licensing, and hub tags include a license:mit entry.

Use cases in the source record

  • Query and passage embedding for MS-MARCO-style retrieval experiments using the publisher's documented prefix format.
  • Sentence-Transformers similarity workflows with normalized embeddings and the documented sentence_transformers usage snippet.

Limitations and unknowns

  • No evaluation scores were extracted from this record.
  • According to the model card, users are directed to e5-small-v2 for better performance with the same usage method; Ethen has not independently compared the versions.
  • No context-window value, hardware requirement, or inference pricing was extracted.
  • Provider state is historical snapshot data and should be refreshed before being presented as current.

Source and provenance

Source: intfloat/e5-small

Captured: Unknown. Processed: 2026-09-07T19:34:47.743126+00:00.

E5-small News (May 2023): please switch to e5-small-v2 , which has better performance and same method of usage. Text Embeddings by Weakly-Supervised Contrastive Pre-training . Liang Wang, Nan Yang, Xiaolong Huang, Binxing Jiao, Linjun Yang, Daxin Jiang, Rangan Majumder, Furu Wei, arXiv 2022 This model has 12 layers and the embedding size is 384. Usage Below is an example to encode queries and passages from the MS-MARCO passage ranking dataset. import torch.nn.functional as F from torch import Tensor from transformers import AutoTokenizer, AutoModel def average_pool ( last_hidden_states: Tensor, attention_mask: Tensor ) -> Tensor: la…

F001F002F003F004F005F006F007F008F009F010F011F012F013F014F015F016F017F018