Skip to content

EthenEthenEthen

Open Source Model Profile · ibm-granite

granite-embedding-30m-english

granite-embedding-30m-english is a 30M-parameter RoBERTa-based English embedding model from IBM Granite. According to the model card, it outputs 384-dimensional vectors for retrieval and similarity.

Publisher
ibm-granite
Task
sentence-similarity
Model type
roberta
License
apache-2.0
Library
sentence-transformers
Publication status
Accepted · not indexed

Model overview

granite-embedding-30m-english is published by ibm-granite as a sentence-similarity embedding model. The captured configuration identifies RobertaModel with model type roberta and about 30M Safetensors parameters, and card data records apache-2.0. According to the model card, it is a dense bi-encoder producing 384-dimensional vectors for similarity, retrieval, and search.

Recorded capabilities

30M bi-encoder embeddings

According to the model card, this is a 30M-parameter dense bi-encoder that produces 384-dimensional text embeddings.

Retrieval and similarity use

According to the model card, it produces fixed-length vectors for text similarity, retrieval, and search, with SentenceTransformer compatibility.

RoBERTa-like encoder

According to the model card, it uses an encoder-only RoBERTa-like architecture with 6 layers, 12 heads, GeLU activation, and 512 maximum sequence length.

Documented training stack

According to the model card, development used retrieval-oriented pretraining, contrastive fine-tuning, knowledge distillation, and model merging, trained on NVIDIA A100 80GB hardware.

Use cases in the source record

  • Text similarity, retrieval, and search applications using fixed-length 384-dimensional embeddings.
  • Multi-turn conversational retrieval experiments, which the publisher describes for the r1.1 revision lineage.

Limitations and unknowns

  • Benchmark figures such as MTEB Retrieval and MT-RAG scores are publisher-reported model-card claims and were not independently verified.
  • According to the model card, the model supports only English texts with 512-token truncation.
  • Provider state is historical snapshot data, not independently refreshed current availability.

Source and provenance

Source: ibm-granite/granite-embedding-30m-english

Captured: Unknown. Processed: 2026-09-07T19:34:50.377893+00:00.

Granite-Embedding-30m-English (revision r1.1) Model Summary: Granite-Embedding-30m-English is a 30M parameter dense bi-encoder embedding model from the Granite Embeddings suite that can be used to generate high quality text embeddings. This model produces embedding vectors of size 384 and is trained using a combination of open source relevance-pair datasets with permissive, enterprise-friendly license, and IBM collected and generated datasets. While maintaining competitive scores on academic benchmarks such as BEIR, this model also performs well on many enterprise use cases. This model is developed using retrieval oriented pre-train…

F001F002F003F004F005F006F007F008F010F011F012F013F014F016F019F020