500M-tweet Spanish training
According to the model card and cited abstract, the model was trained on more than 500 million Spanish tweets for user-generated text.
Open Source Model Profile · pysentimiento
robertuito-base-uncased is a 108.82M-parameter RoBERTa fill-mask model from pysentimiento. According to the model card, it targets Spanish social-media text after training on 500 million tweets.
robertuito-base-uncased is published by pysentimiento as a fill-mask model. The captured configuration identifies RobertaForMaskedLM with model type roberta, and Safetensors metadata reports 108,818,866 parameters. According to the model card, it follows RoBERTa guidelines and comes in cased, uncased, and uncased-plus-deaccented flavors.
According to the model card and cited abstract, the model was trained on more than 500 million Spanish tweets for user-generated text.
According to the model card, RoBERTuito comes in cased, uncased, and uncased-plus-deaccented variants.
According to the model card table, the uncased variant scores 0.801 hate-speech, 0.707 sentiment, 0.551 emotion, and 0.736 irony, above the listed BETO and RoBERTa-BNE baselines.
Source: pysentimiento/robertuito-base-uncased
Captured: Unknown. Processed: 2026-09-07T19:34:56.373065+00:00.
robertuito-base-uncased RoBERTuito A pre-trained language model for social media text in Spanish PAPER Github Repository RoBERTuito is a pre-trained language model for user-generated content in Spanish, trained following RoBERTa guidelines on 500 million tweets. RoBERTuito comes in 3 flavors: cased, uncased, and uncased+deaccented. We tested RoBERTuito on a benchmark of tasks involving user-generated text in Spanish. It outperforms other pre-trained language models for this language such as BETO , BERTin and RoBERTa-BNE . The 4 tasks selected for evaluation were: Hate Speech Detection (using SemEval 2019 Task 5, HatEval dataset), Se…
F001F002F003F004F005F006F007F008F009F010F011F012F013F015