Skip to content

EthenEthenEthen

Open Source Model Profile · intfloat

e5-small-v2

e5-small-v2 is a 33.36M-parameter BERT embedding model from intfloat. According to the model card, it has 12 layers with 384-dimensional embeddings trained by weakly-supervised contrastive pre-training.

Publisher
intfloat
Task
sentence-similarity
Model type
bert
License
mit
Library
sentence-transformers
Publication status
Approved for indexing

Model overview

e5-small-v2 is published by intfloat as a sentence-similarity model. The captured configuration identifies BertModel with model type bert, and Safetensors metadata reports 33,360,512 parameters. According to the model card, it produces text embeddings for retrieval-style query and passage encoding.

Recorded capabilities

384-dimensional embeddings

According to the model card, the model has 12 layers and an embedding size of 384.

Query and passage prefixes

According to the model card, inputs should retain query and passage prefixes, or performance degrades.

Transformers and Sentence-Transformers support

According to the model card, usage covers Transformers average pooling and SentenceTransformer encoding with normalized embeddings.

Use cases in the source record

  • Passage ranking and retrieval embedding using the documented query and passage prefixes over inputs such as MS-MARCO examples.
  • Sentence-Transformers similarity workflows that encode prefixed inputs with normalized embeddings.

Limitations and unknowns

  • No evaluation results were extracted as structured data from this record.
  • No context-window value was extracted from this record.
  • Provider state is historical snapshot data, not independently refreshed current availability.

Source and provenance

Source: intfloat/e5-small-v2

Captured: Unknown. Processed: 2026-09-07T19:34:47.883555+00:00.

E5-small-v2 Text Embeddings by Weakly-Supervised Contrastive Pre-training . Liang Wang, Nan Yang, Xiaolong Huang, Binxing Jiao, Linjun Yang, Daxin Jiang, Rangan Majumder, Furu Wei, arXiv 2022 This model has 12 layers and the embedding size is 384. Usage Below is an example to encode queries and passages from the MS-MARCO passage ranking dataset. import torch.nn.functional as F from torch import Tensor from transformers import AutoTokenizer, AutoModel def average_pool ( last_hidden_states: Tensor, attention_mask: Tensor ) -> Tensor: last_hidden = last_hidden_states.masked_fill(~attention_mask[..., None ]. bool (), 0.0 ) return last_h…

F001F002F003F004F005F006F007F008F009F010F011F012F013F014F015F016F017F018