Skip to content

EthenEthenEthen

Open Source Model Profile · AITeamVN

Vietnamese_Embedding

Vietnamese_Embedding is a 568M-parameter Vietnamese sentence-embedding model from AITeamVN. According to the model card, it is fine-tuned from BAAI/bge-m3 to strengthen Vietnamese retrieval.

Publisher
AITeamVN
Task
sentence-similarity
Model type
xlm-roberta
License
apache-2.0
Library
sentence-transformers
Publication status
Accepted · not indexed

Model overview

Vietnamese_Embedding is published by AITeamVN as a sentence-similarity embedding model. Captured Safetensors metadata reports 567,754,752 parameters, about 568M. According to the model card, it is a Sentence Transformer fine-tuned from BAAI/bge-m3 on approximately 300,000 Vietnamese query-document triplets with a 2048-token maximum sequence length.

Recorded capabilities

BGE-M3 Vietnamese fine-tune

According to the model card, the model is fine-tuned from BAAI/bge-m3 specifically to enhance Vietnamese retrieval capability.

2048-token inputs, 1024 dimensions

According to the model card, it supports a 2048-token maximum sequence length with 1024-dimension output and dot-product similarity.

Triplet-trained retrieval

According to the model card, training used approximately 300,000 triplets of queries with positive and negative Vietnamese documents.

Publisher-reported Legal Zalo scores

According to the model card's table, the public checkpoint reports 0.7274 Accuracy@1 and 0.8181 MRR@10, above the listed BGE-M3 baseline.

Use cases in the source record

  • Vietnamese passage retrieval and semantic search using the documented 1024-dimension dot-product similarity.
  • Legal-domain retrieval evaluation on the Legal Zalo 2021 training set, which the card states was not used for training.

Limitations and unknowns

  • No independently measured evaluation results were extracted; Legal Zalo figures are publisher-reported card claims.
  • No VRAM, hardware, or pricing detail was extracted from this record.
  • The evaluation table also covers sibling reranker and v2 checkpoints rather than this checkpoint alone.
  • Provider state is historical snapshot data, not independently refreshed current availability.

Source and provenance

Source: AITeamVN/Vietnamese_Embedding

Captured: Unknown. Processed: 2026-09-07T19:35:01.894233+00:00.

Model Card: Vietnamese_Embedding Vietnamese_Embedding is an embedding model fine-tuned from the BGE-M3 model ( https://huggingface.co/BAAI/bge-m3 ) to enhance retrieval capabilities for Vietnamese. The model was trained on approximately 300,000 triplets of queries, positive documents, and negative documents for Vietnamese. The model was trained with a maximum sequence length of 2048. Model Details Model Description Model Type: Sentence Transformer Base model: BAAI/bge-m3 Maximum Sequence Length: 2048 tokens Output Dimensionality: 1024 dimensions Similarity Function: Dot product Similarity Language: Vietnamese Licence: Apache 2.0 Usa…

F001F002F003F004F005F006F007F009F010F011F012F013F014F015F016