Skip to content

EthenEthenEthen

Open Source Model Profile · denaya

indoSBERT-large

indoSBERT-large is a BERT sentence-similarity model from denaya. Its model card describes 256-dimensional embeddings intended for Indonesian semantic search and clustering.

Publisher
denaya
Task
sentence-similarity
Model type
bert
License
Unknown
Library
sentence-transformers
Publication status
Approved for indexing

Model overview

denaya publishes indoSBERT-large as a sentence-similarity embedding model on the sentence-transformers stack. Captured config lists BertModel and a bert model type. According to the model card, IndoSBERT is a modification of indobenchmark/indobert-large-p1 that was fine-tuned with a siamese network scheme inspired by SBERT.

Recorded capabilities

256-dimension embeddings

According to the model card, the model maps sentences and paragraphs to a 256-dimensional dense vector space, with a Dense head projecting pooled 1024-d features to 256.

Indonesian STS fine-tune

The model card says the model was fine-tuned on the STS Dataset (2012-2016) machine-translated into Indonesian, and that it can provide semantic embeddings for Indonesian sentences.

Siamese SBERT-style training

The publisher describes IndoSBERT as a modification of indobenchmark/indobert-large-p1 fine-tuned with a siamese network scheme inspired by SBERT (Reimers et al., 2019).

Sentence-transformers usage

Hub metadata lists sentence-transformers, and the model card shows loading via SentenceTransformer with Indonesian example sentences.

Use cases in the source record

  • Semantic search over Indonesian sentences using the publisher-described dense embeddings.
  • Clustering of sentences and paragraphs in the documented vector space.
  • SentenceTransformer encode workflows that follow the card's usage snippet.

Limitations and unknowns

  • No evaluation results were extracted from this record.
  • No parameter count was extracted from this record.
  • No license value was extracted from this record.
  • Provider state is historical snapshot data, not independently refreshed current availability.
  • Embedding dimension, sequence length, training setup, and the IndoBERT-large-p1 modification claim come from the publisher model card and have not been independently verified by Ethen.

Source and provenance

Source: denaya/indoSBERT-large

Captured: Unknown. Processed: 2026-09-07T19:34:43.586570+00:00.

indoSBERT-large This is a sentence-transformers model: It maps sentences & paragraphs to a 256 dimensional dense vector space and can be used for tasks like clustering or semantic search. IndoSBERT is a modification of https://huggingface.co/indobenchmark/indobert-large-p1 that has been fine-tuned using the siamese network scheme inspired by SBERT (Reimers et al., 2019). This model was fine-tuned with the STS Dataset (2012-2016) which was machine-translated into Indonesian languange. This model can provide meaningful semantic sentence embeddings for Indonesian sentences. Usage (Sentence-Transformers) Using this model becomes easy wh…

F001F002F003F004F005F006F008F009F010F011F012F013F016