Skip to content

EthenEthenEthen

Open Source Model Profile · oliverguhr

fullstop-punctuation-multilang-large

fullstop-punctuation-multilang-large is a 559M-parameter XLM-RoBERTa token-classification model from oliverguhr. Its model card documents punctuation restoration across four languages.

Publisher
oliverguhr
Task
token-classification
Model type
xlm-roberta
License
mit
Library
transformers
Publication status
Accepted · not indexed

Model overview

fullstop-punctuation-multilang-large is published by oliverguhr as a token-classification model. Safetensors metadata reports 558847496 parameters, and the captured configuration identifies XLMRobertaForTokenClassification. According to the model card, it restores punctuation in transcribed spoken language across English, Italian, French, and German.

Recorded capabilities

559M XLM-RoBERTa classifier

Captured config identifies XLMRobertaForTokenClassification and Safetensors metadata reports 558847496 parameters.

Four-language punctuation

According to the model card, the model predicts punctuation for English, Italian, French, and German texts.

Five documented markers

The model card says it restores period, comma, question mark, hyphen, and colon markers.

Europarl training with domain note

The card says the model was trained on the Europarl dataset and warns it may behave differently outside political-speech text.

Use cases in the source record

  • Punctuation restoration for transcribed English, Italian, French, and German speech using the five documented markers.
  • Text cleanup pipelines that process long inputs with the publisher's documented Python package.

Limitations and unknowns

  • No evaluation results were extracted from this record.
  • No context-window value was extracted from this record.
  • Provider state is historical snapshot data, not independently refreshed current availability.
  • According to the model card, the Europarl training data consists of political speeches, so behavior on other domains is uncertain.

Source and provenance

Source: oliverguhr/fullstop-punctuation-multilang-large

Captured: Unknown. Processed: 2026-09-07T19:34:54.290393+00:00.

This model predicts the punctuation of English, Italian, French and German texts. We developed it to restore the punctuation of transcribed spoken language. This multilanguage model was trained on the Europarl Dataset provided by the SEPP-NLG Shared Task . Please note that this dataset consists of political speeches. Therefore the model might perform differently on texts from other domains. The model restores the following punctuation markers: "." "," "?" "-" ":" Sample Code We provide a simple python package that allows you to process text of any length. Install To get started install the package from pypi : pip install deepmultili…

F001F002F003F004F005F006F007F009F010F011F012F013F014