Skip to content

EthenEthenEthen

Open Source Model Profile · pysentimiento

robertuito-base-uncased

robertuito-base-uncased is a 108.82M-parameter RoBERTa fill-mask model from pysentimiento. According to the model card, it targets Spanish social-media text after training on 500 million tweets.

Publisher
pysentimiento
Task
fill-mask
Model type
roberta
License
Unknown
Library
transformers
Publication status
Accepted · not indexed

Model overview

robertuito-base-uncased is published by pysentimiento as a fill-mask model. The captured configuration identifies RobertaForMaskedLM with model type roberta, and Safetensors metadata reports 108,818,866 parameters. According to the model card, it follows RoBERTa guidelines and comes in cased, uncased, and uncased-plus-deaccented flavors.

Recorded capabilities

500M-tweet Spanish training

According to the model card and cited abstract, the model was trained on more than 500 million Spanish tweets for user-generated text.

Three released flavors

According to the model card, RoBERTuito comes in cased, uncased, and uncased-plus-deaccented variants.

Reported Spanish benchmark table

According to the model card table, the uncased variant scores 0.801 hate-speech, 0.707 sentiment, 0.551 emotion, and 0.736 irony, above the listed BETO and RoBERTa-BNE baselines.

Use cases in the source record

  • Spanish social-media analysis including hate-speech, sentiment, emotion, and irony tasks covered by the card's reported benchmark.
  • Spanish masked-language modeling with the card's SentencePiece spacing guidance for mask tokens.

Limitations and unknowns

  • Benchmark comparisons and superiority statements are publisher claims and were not independently verified.
  • No license value was extracted from this record.
  • No context-window value was extracted from this record.
  • Provider state is historical snapshot data, not independently refreshed current availability.

Source and provenance

Source: pysentimiento/robertuito-base-uncased

Captured: Unknown. Processed: 2026-09-07T19:34:56.373065+00:00.

robertuito-base-uncased RoBERTuito A pre-trained language model for social media text in Spanish PAPER Github Repository RoBERTuito is a pre-trained language model for user-generated content in Spanish, trained following RoBERTa guidelines on 500 million tweets. RoBERTuito comes in 3 flavors: cased, uncased, and uncased+deaccented. We tested RoBERTuito on a benchmark of tasks involving user-generated text in Spanish. It outperforms other pre-trained language models for this language such as BETO , BERTin and RoBERTa-BNE . The 4 tasks selected for evaluation were: Hate Speech Detection (using SemEval 2019 Task 5, HatEval dataset), Se…

F001F002F003F004F005F006F007F008F009F010F011F012F013F015