Skip to content

EthenEthenEthen

Open Source Model Profile · dragonkue

multilingual-e5-small-ko-v2

multilingual-e5-small-ko-v2 is a 0.12B-parameter BERT embedding model from dragonkue. According to the model card, it targets Korean retrieval with 384-dimensional vectors.

Publisher
dragonkue
Task
sentence-similarity
Model type
bert
License
apache-2.0
Library
sentence-transformers
Publication status
Approved for indexing

Model overview

multilingual-e5-small-ko-v2 is published by dragonkue as a sentence-similarity model. The captured configuration identifies BertModel, and Safetensors metadata reports about 0.12B parameters. According to the model card, it finetunes intfloat/multilingual-e5-small for Korean retrieval with 384-dimensional embeddings.

Recorded capabilities

Sentence-Transformers embedding base

The captured configuration identifies BertModel with sentence-transformers support for embedding workflows.

Korean retrieval focus

According to the model card, fine-tuning on Korean query-passage pairs targets Korean retrieval performance.

384-dimensional output

According to the model card, sentences and paragraphs map to a 384-dimensional dense vector space.

Model Soup construction

According to the model card, a 6:4 weighted merge combines the Korean-specialized checkpoint with the base multilingual model.

GISTEmbedLoss training

According to the model card, training uses clustered in-batch negatives with GISTEmbedLoss and margin after MNR-only loss reduced performance.

Use cases in the source record

  • Korean semantic search, similarity, paraphrase mining, classification, and clustering using 384-dimensional embeddings.
  • Retrieval evaluation work consistent with the listed Ko-StrategyQA, AutoRAG, MIRACL, and PublicHealthQA dataset descriptions.

Limitations and unknowns

  • The publisher states stronger Korean-benchmark performance than the larger e5-base model, but no extracted numeric table is stated here.
  • The card lists overlapping epoch figures in its hyperparameter blocks, so exact epoch count should be confirmed from the training configuration.
  • Provider state is historical snapshot data and should be refreshed before being presented as current.

Source and provenance

Source: dragonkue/multilingual-e5-small-ko-v2

Captured: Unknown. Processed: 2026-09-07T19:35:55.713752+00:00.

SentenceTransformer based on intfloat/multilingual-e5-small This is a sentence-transformers model finetuned from intfloat/multilingual-e5-small on datasets that include Korean query-passage pairs for improved performance on Korean retrieval tasks. It maps sentences & paragraphs to a 384-dimensional dense vector space and can be used for semantic textual similarity, semantic search, paraphrase mining, text classification, clustering, and more. This model is a lightweight Korean retriever, designed for ease of use and strong performance in practical retrieval tasks. It is ideal for running demos or lightweight applications, offering a…

F001F002F003F004F005F006F007F008F009F010F011F014F015F016F017F021F022F023F024F028F029F030F031F032F033