Skip to content

EthenEthenEthen

Open Source Model Profile · vinai

phobert-base-v2

vinai/phobert-base-v2 is a RoBERTa-based Vietnamese fill-mask model. According to the model card, the base v2 release trains a 135M-parameter model on 140GB of Vietnamese text.

Publisher
vinai
Task
fill-mask
Model type
roberta
License
agpl-3.0
Library
transformers
Publication status
Approved for indexing

Model overview

phobert-base-v2 is published by VinAI as a fill-mask model for Vietnamese. The captured configuration identifies RobertaForMaskedLM with model type roberta. According to the model card, it is a 135M-parameter RoBERTa-based base model pre-trained on 20GB of Wikipedia and news texts plus 120GB of OSCAR-2301 texts.

Recorded capabilities

Vietnamese monolingual pre-training

According to the model card, PhoBERT base and large were the first public large-scale monolingual language models pre-trained for Vietnamese.

Reported downstream results

According to the model card, PhoBERT obtained new state-of-the-art results on Vietnamese part-of-speech tagging, dependency parsing, named-entity recognition, and natural language inference.

Expanded v2 training corpus

According to the model card, base-v2 keeps the 135M base size and 256 length while adding 120GB of OSCAR-2301 texts to the original 20GB corpus.

Documented segmentation workflow

According to the model card, downstream applications should apply VnCoreNLP RDRSegmenter word segmentation to raw input before feeding PhoBERT.

Use cases in the source record

  • Vietnamese masked-language modeling for part-of-speech tagging, dependency parsing, named-entity recognition, and natural language inference fine-tuning.
  • Vietnamese text workflows that segment raw input with the VnCoreNLP RDRSegmenter before PhoBERT encoding.

Limitations and unknowns

  • No Safetensors parameter count was captured; the 135M figure is a publisher model-card claim.
  • No evaluation results were extracted from this record.
  • Provider state is historical snapshot data, not independently refreshed current availability.
  • Downstream performance claims come from the publisher model card and were not independently verified.

Source and provenance

Source: vinai/phobert-base-v2

Captured: Unknown. Processed: 2026-09-07T19:35:00.560590+00:00.

Table of contents Introduction Using PhoBERT with transformers Installation Pre-trained models Example usage Using PhoBERT with fairseq Notes PhoBERT: Pre-trained language models for Vietnamese Pre-trained PhoBERT models are the state-of-the-art language models for Vietnamese ( Pho , i.e. "Phở", is a popular food in Vietnam): Two PhoBERT versions of "base" and "large" are the first public large-scale monolingual language models pre-trained for Vietnamese. PhoBERT pre-training approach is based on RoBERTa which optimizes the BERT pre-training procedure for more robust performance. PhoBERT outperforms previous monolingual and multilin…

F001F002F003F004F005F006F008F009F010F011F015F016F017