Skip to content

EthenEthenEthen

Open Source Model Profile · kredor

punctuate-all

punctuate-all is an XLM-RoBERTa-family token-classification fine-tune from kredor. According to the model card, it restores punctuation across twelve languages.

Publisher
kredor
Task
token-classification
Model type
xlm-roberta
License
mit
Library
transformers
Publication status
Accepted · not indexed

Model overview

punctuate-all is published by kredor as a token-classification model. The captured configuration identifies XLMRobertaForTokenClassification with model type xlm-roberta. According to the model card, it finetunes xlm-roberta-base for punctuation restoration across twelve languages.

Recorded capabilities

XLM-RoBERTa token-classification architecture

The captured configuration identifies XLMRobertaForTokenClassification with model type xlm-roberta and Transformers support.

Twelve-language punctuation restoration

According to the model card, the model restores punctuation across twelve listed European languages.

Publisher-reported validation report

According to the model card, validation reports 0.98 accuracy with strong period, comma, and question-mark F1 and weaker hyphen and colon F1.

Europarl dataset link

Hub tags record dataset wmt/europarl alongside transformers, pytorch, and xlm-roberta tagging.

Use cases in the source record

  • Multilingual punctuation restoration for raw text across the twelve publisher-listed languages.
  • Post-processing evaluation using the publisher-reported per-mark precision, recall, and confusion patterns.

Limitations and unknowns

  • No parameter count was extracted from this record.
  • According to the model card, hyphen and colon classes show lower recall and F1 than periods, commas, and question marks in the reported table.
  • Provider state is historical snapshot data and should be refreshed before being presented as current.

Source and provenance

Source: kredor/punctuate-all

Captured: Unknown. Processed: 2026-09-07T19:34:49.248842+00:00.

This is based on Oliver Guhr's work . The difference is that it is a finetuned xlm-roberta-base instead of an xlm-roberta-large and on twelve languages instead of four. The languages are: English, German, French, Spanish, Bulgarian, Italian, Polish, Dutch, Czech, Portugese, Slovak, Slovenian. ----- report ----- precision recall f1-score support 0 0.99 0.99 0.99 73317475 . 0.94 0.95 0.95 4484845 , 0.86 0.86 0.86 6100650 ? 0.88 0.85 0.86 136479 - 0.60 0.29 0.39 233630 : 0.71 0.49 0.58 152424 accuracy 0.98 84425503 macro avg 0.83 0.74 0.77 84425503 weighted avg 0.98 0.98 0.98 84425503 ----- confusion matrix ----- t/p 0 . , ? - : 0 1.0…

F001F002F003F004F005F006F007F008F009F010F011F012F013