Skip to content

EthenEthenEthen

Open Source Model Profile · Helsinki-NLP

opus-mt-en-dra

opus-mt-en-dra is a Helsinki-NLP translation model from English into four Dravidian languages. Its card documents Transformer inference with SentencePiece preprocessing and published Tatoeba BLEU scores.

Publisher
Helsinki-NLP
Task
translation
Model type
marian
License
apache-2.0
Library
transformers
Publication status
Accepted · not indexed

Model overview

opus-mt-en-dra is published by Helsinki-NLP as a translation model with MarianMTModel architecture and a marian model type. According to the model card, it is a Transformer that translates English into Kannada, Malayalam, Tamil, and Telugu. Captured metadata records an Apache-2.0 license, Transformers compatibility, and matching English and Dravidian language tags.

Recorded capabilities

Four Dravidian targets

The card covers English-to-Kannada, Malayalam, Tamil, and Telugu translation with a per-sentence target-language token.

Published Tatoeba BLEU table

The card reports per-language BLEU and chr-F scores with a 10.7 multi-target BLEU average.

SentencePiece preprocessing

The card documents normalization plus 32k SentencePiece vocabularies for source and target text.

Use cases in the source record

  • English-to-Dravidian translation into Kannada, Malayalam, Tamil, or Telugu with the required target-language token.
  • Benchmark comparisons against the card's published per-language BLEU and chr-F figures.

Limitations and unknowns

  • No parameter count was extracted from this record.
  • BLEU figures are publisher-reported 2020 Tatoeba Challenge scores, not independently measured Ethen evaluations.
  • Provider state is historical snapshot data, not independently refreshed current availability.

Source and provenance

Source: Helsinki-NLP/opus-mt-en-dra

Captured: Unknown. Processed: 2026-09-07T19:34:31.334261+00:00.

eng-dra source group: English target group: Dravidian languages OPUS readme: eng-dra model: transformer source language(s): eng target language(s): kan mal tam tel model: transformer pre-processing: normalization + SentencePiece (spm32k,spm32k) a sentence initial language token is required in the form of >>id<< (id = valid target language ID) download original weights: opus-2020-07-26.zip test set translations: opus-2020-07-26.test.txt test set scores: opus-2020-07-26.eval.txt Benchmarks testset BLEU chr-F Tatoeba-test.eng-kan.eng.kan 4.7 0.348 Tatoeba-test.eng-mal.eng.mal 13.1 0.515 Tatoeba-test.eng.multi 10.7 0.463 Tatoeba-test.en…

F001F002F003F004F005F006F007F008F009F010F011