Skip to content

EthenEthenEthen

Open Source Model Profile · jb2k

bert-base-multilingual-cased-language-detection

bert-base-multilingual-cased-language-detection is a BERT text-classification fine-tune from jb2k for language detection across 45 languages.

Publisher
jb2k
Task
text-classification
Model type
bert
License
Unknown
Library
transformers
Publication status
Accepted · not indexed

Model overview

bert-base-multilingual-cased-language-detection is published by jb2k as a Transformers text-classification model. The captured configuration identifies BertForSequenceClassification and bert. According to the model card, it fine-tunes bert-base-multilingual-cased on the common language dataset for detection across 45 languages.

Recorded capabilities

45-language detection scope

According to the model card, the model detects 45 languages, with the page listing languages from Arabic and Basque to Turkish, Ukrainian, and Welsh.

bert-base-multilingual-cased fine-tune

According to the model card, this model was created by fine-tuning bert-base-multilingual-cased on the common language dataset.

BERT classification architecture

Captured config reports BertForSequenceClassification and bert for text classification.

Publisher-reported 97.8% accuracy

According to the model card, evaluation on the test split of the common language dataset achieved 97.8% accuracy.

Use cases in the source record

  • Language detection over the documented 45-language set, including major European, Asian, and regional languages.
  • Text-classification inference through a Transformers and endpoints-compatible workflow for language tagging.

Limitations and unknowns

  • No parameter count was extracted from this record.
  • No license value was captured, so reuse terms are unknown from this record.
  • The 97.8% accuracy figure is a publisher-reported test-split result, not an independently verified Ethen measurement.
  • Provider state is historical snapshot data, not independently refreshed current availability.

Source and provenance

Source: jb2k/bert-base-multilingual-cased-language-detection

Captured: Unknown. Processed: 2026-09-07T19:34:47.923326+00:00.

bert-base-multilingual-cased-language-detection A model for language detection with support for 45 languages Model description This model was created by fine-tuning bert-base-multilingual-cased on the common language dataset. This dataset has support for 45 languages, which are listed below: Arabic, Basque, Breton, Catalan, Chinese_China, Chinese_Hongkong, Chinese_Taiwan, Chuvash, Czech, Dhivehi, Dutch, English, Esperanto, Estonian, French, Frisian, Georgian, German, Greek, Hakha_Chin, Indonesian, Interlingua, Italian, Japanese, Kabyle, Kinyarwanda, Kyrgyz, Latvian, Maltese, Mongolian, Persian, Polish, Portuguese, Romanian, Romansh_…

F001F002F003F004F005F006F008F009F010F011F012