Skip to content

EthenEthenEthen

Open Source Model Profile · tartuNLP

EstBERT

EstBERT is a 124.49M-parameter BERT fill-mask model from tartuNLP. According to the model card, it was exclusively trained on an Estonian cased corpus at 128 and 512 sequence lengths.

Publisher
tartuNLP
Task
fill-mask
Model type
bert
License
cc-by-4.0
Library
transformers
Publication status
Accepted · not indexed

Model overview

EstBERT is published by tartuNLP as a bert-based fill-mask model. The captured configuration identifies BertForMaskedLM and Safetensors metadata reports 124493392 parameters. According to the model card, it is a pretrained BERT Base model exclusively trained on an Estonian cased corpus, with card data recording cc-by-4.0.

Recorded capabilities

Estonian-only BERT Base

According to the model card, it is a pretrained BERT Base model exclusively trained on an Estonian cased corpus.

128 and 512 sequence training

According to the model card, the model was trained on both 128 and 512 sequence lengths, with separate EstBERT_128 and EstBERT_512 downloads documented.

Estonian National Corpus 2017

According to the model card, training used the Estonian National Corpus 2017, described as the largest Estonian corpus at the time, with four sub-corpora covering reference, web, and Wikipedia material.

Transformers masked-LM use

Captured config identifies BertForMaskedLM with transformers library support, and the model card describes use through AutoTokenizer and AutoModelForMaskedLM.

CC-BY-4.0 licensing

Card data records cc-by-4.0 for this repository.

Use cases in the source record

  • Estonian fill-mask prediction using the documented transformers AutoTokenizer and AutoModelForMaskedLM workflow.
  • Estonian cased-text research using a model the publisher describes as exclusively trained on Estonian corpus material.

Limitations and unknowns

  • No evaluation results were extracted from this record.
  • No context-window value was extracted from this record.
  • Provider state is historical snapshot data, not independently refreshed current availability.
  • Training and corpus details come from the publisher model card and were not independently verified.

Source and provenance

Source: tartuNLP/EstBERT

Captured: Unknown. Processed: 2026-09-07T19:34:59.253323+00:00.

EstBERT What's this? The EstBERT model is a pretrained BERT Base model exclusively trained on Estonian cased corpus on both 128 and 512 sequence length of data. How to use? You can use the model with the transformers library the following way. from transformers import AutoTokenizer, AutoModelForMaskedLM tokenizer = AutoTokenizer.from_pretrained( "tartuNLP/EstBERT" ) model = AutoModelForMaskedLM.from_pretrained( "tartuNLP/EstBERT" ) You can also download the pretrained model from here, EstBERT_128 EstBERT_512 Dataset used to train the model The EstBERT model is trained both on 128 and 512 sequence length of data. For training the Est…

F001F002F003F004F005F006F007F008F010F012F014F015