Skip to content

EthenEthenEthen

Open Source Model Profile · indolem

indobertweet-base-uncased

indobertweet-base-uncased is a BERT-family fill-mask model from indolem. According to the model card, it is a large-scale pretrained model for Indonesian Twitter with domain-specific vocabulary.

Publisher
indolem
Task
fill-mask
Model type
bert
License
apache-2.0
Library
transformers
Publication status
Approved for indexing

Model overview

indobertweet-base-uncased is published by indolem as a fill-mask model for Indonesian Twitter. The captured configuration identifies BertForMaskedLM with model type bert. According to the model card, it extends an Indonesian BERT model with additive domain-specific vocabulary and is presented in an EMNLP 2021 paper.

Recorded capabilities

Twitter-domain vocabulary

According to the model card, the model extends an Indonesian BERT model with additive domain-specific vocabulary for Twitter.

Documented preprocessing

According to the model card, usage lowercases words, converts mentions and URLs to @USER and HTTPURL, and translates emoticons with the emoji package.

Reported seven-dataset results

According to the model card table, the vocabulary-adapted variant averages 86.5 across sentiment, emotion, hate-speech, and NER sets, above the listed IndoBERT baselines.

Use cases in the source record

  • Indonesian Twitter analysis such as sentiment, emotion, hate-speech, and NER workflows covered by the card's reported seven-dataset table.
  • Fill-mask modeling with the documented preprocessing of lowercasing, @USER and HTTPURL replacement, and emoji translation.

Limitations and unknowns

  • Reported benchmark figures and baseline comparisons are publisher claims and were not independently verified.
  • No parameter count was extracted from this record.
  • No context-window value was extracted from this record.
  • Provider state is historical snapshot data, not independently refreshed current availability.

Source and provenance

Source: indolem/indobertweet-base-uncased

Captured: Unknown. Processed: 2026-09-07T19:34:47.096437+00:00.

IndoBERTweet 🐦 1. Paper Fajri Koto, Jey Han Lau, and Timothy Baldwin. IndoBERTweet: A Pretrained Language Model for Indonesian Twitter with Effective Domain-Specific Vocabulary Initialization . In Proceedings of the 2021 Conference on Empirical Methods in Natural Language Processing ( EMNLP 2021 ), Dominican Republic (virtual). 2. About IndoBERTweet is the first large-scale pretrained model for Indonesian Twitter that is trained by extending a monolingually trained Indonesian BERT model with additive domain-specific vocabulary. In this paper, we show that initializing domain-specific vocabulary with average-pooling of BERT subword…

F001F002F003F004F005F006F007F008F009F010F011F012F013F014F015F016