Skip to content

EthenEthenEthen

Open Source Model Profile · openai

whisper-large-v3

whisper-large-v3 is a 1.54B-parameter multilingual speech model from openai for recognition and translation. According to the model card, it shares the large-v2 architecture with 128 Mel bins and a new Cantonese token.

Publisher
openai
Task
automatic-speech-recognition
Model type
whisper
License
apache-2.0
Library
transformers
Publication status
Approved for indexing

Model overview

whisper-large-v3 is published by openai as an automatic-speech-recognition model. Captured Safetensors metadata reports 1,543,490,560 parameters, about 1.54B. According to the model card, it is a Transformer encoder-decoder trained for multilingual recognition and translation, showing 10-20% error reduction over large-v2 across many languages.

Recorded capabilities

128 Mel bins plus Cantonese

According to the model card, large-v3 differs from large-v2 with 128 Mel frequency bins instead of 80 and a new Cantonese language token.

Large-scale weak supervision

According to the model card, large-v3 trained for 2.0 epochs on 1 million hours of weakly labeled plus 4 million hours of pseudo-labeled audio.

30-second chunked long-form

According to the model card, audio beyond the 30-second receptive field uses sequential or chunked long-form algorithms.

Multilingual-only 1550M scale

According to the model card's table, the large checkpoints are multilingual-only at 1550M parameters.

Use cases in the source record

  • Long-form transcription with the card's pipeline class, including sequential sliding-window or chunked segmentation for audio beyond 30 seconds.
  • Speech translation workflows such as French-to-English with sentence-level timestamps via the documented generate options.
  • Domain fine-tuning where the card notes as little as 5 hours of labelled data can improve specific languages and tasks.

Limitations and unknowns

  • According to the model card, weakly supervised training can produce hallucinations of text not present in the audio.
  • According to the model card, accuracy is uneven across low-resource languages, accents, and dialects, with possible demographic disparities.
  • According to the model card, the sequence-to-sequence design can emit repetitive text, only partly mitigated by beam search and temperature scheduling.
  • Provider state is historical snapshot data, not independently refreshed current availability.

Source and provenance

Source: openai/whisper-large-v3

Captured: Unknown. Processed: 2026-09-07T19:34:54.464607+00:00.

Whisper Whisper is a state-of-the-art model for automatic speech recognition (ASR) and speech translation, proposed in the paper Robust Speech Recognition via Large-Scale Weak Supervision by Alec Radford et al. from OpenAI. Trained on >5M hours of labeled data, Whisper demonstrates a strong ability to generalise to many datasets and domains in a zero-shot setting. Whisper large-v3 has the same architecture as the previous large and large-v2 models, except for the following minor differences: The spectrogram input uses 128 Mel frequency bins instead of 80 A new language token for Cantonese The Whisper large-v3 model was trained on 1…

F001F002F003F004F005F006F007F010F011F012F013F015F018F022F026F027F029F031F032F033