Japanese three-corpus pre-training
According to the model card, training used Japanese Wikipedia, the Japanese CC-100 portion, and the Japanese OSCAR portion, totaling 171GB after duplicating Wikipedia ten times.
Open Source Model Profile · ku-nlp
deberta-v2-base-japanese is a 137.03M-parameter Japanese DeBERTa V2 fill-mask model from ku-nlp. Its model card documents pre-training on Japanese Wikipedia, CC-100, and OSCAR.
deberta-v2-base-japanese is published by ku-nlp as a DeBERTa V2 masked-language model. The captured configuration identifies DebertaV2ForMaskedLM and Safetensors metadata reports 137,031,168 parameters. According to the model card, it was pre-trained on Japanese Wikipedia, the Japanese portion of CC-100, and the Japanese portion of OSCAR.
According to the model card, training used Japanese Wikipedia, the Japanese CC-100 portion, and the Japanese OSCAR portion, totaling 171GB after duplicating Wikipedia ten times.
According to the model card, texts were segmented with Juman++ 2.0.0-rc3 and tokenized into subwords with SentencePiece before training with Transformers.
According to the model card, the trained model reached 0.779 accuracy on the masked-language-modeling task over a sampled evaluation set.
Source: ku-nlp/deberta-v2-base-japanese
Captured: Unknown. Processed: 2026-09-07T19:34:49.098637+00:00.
Model Card for Japanese DeBERTa V2 base Model description This is a Japanese DeBERTa V2 base model pre-trained on Japanese Wikipedia, the Japanese portion of CC-100, and the Japanese portion of OSCAR. How to use You can use this model for masked language modeling as follows: from transformers import AutoTokenizer, AutoModelForMaskedLM tokenizer = AutoTokenizer.from_pretrained( 'ku-nlp/deberta-v2-base-japanese' ) model = AutoModelForMaskedLM.from_pretrained( 'ku-nlp/deberta-v2-base-japanese' ) sentence = '京都 大学 で 自然 言語 処理 を [MASK] する 。' # input should be segmented into words by Juman++ in advance encoding = tokenizer(sentence, return…
F001F002F003F004F005F006F007F009F010F012F013F014F015F016F017F018F019F021F022F023