Skip to content

EthenEthenEthen

Open Source Model Profile · openai

whisper-large-v3-turbo

whisper-large-v3-turbo is an 809M-parameter automatic-speech-recognition model from OpenAI. According to the model card, it is a pruned, fine-tuned Whisper large-v3 with decoding layers reduced from 32 to 4.

Publisher
openai
Task
automatic-speech-recognition
Model type
whisper
License
mit
Library
transformers
Publication status
Accepted · not indexed

Model overview

whisper-large-v3-turbo is published by OpenAI as an automatic-speech-recognition model. The captured configuration identifies WhisperForConditionalGeneration with whisper model type, and Safetensors metadata reports 808,878,080 parameters. According to the model card, it is otherwise the same as large-v3 except for the reduced decoder depth.

Recorded capabilities

Pruned turbo design

According to the model card, decoding layers were reduced from 32 to 4, making the model faster at the cost of minor quality degradation.

809M Whisper config

Safetensors metadata reports 808,878,080 parameters; the model card tabulates large-v3-turbo at 809M.

Transformers support

According to the model card, the model is supported in Hugging Face Transformers with pipeline and model-plus-processor APIs.

Multilingual tags

Captured tags record whisper, audio, automatic-speech-recognition, and many language codes with endpoints-compatible support.

Use cases in the source record

  • Speech-recognition transcription using the publisher-documented Transformers pipeline for arbitrary-length audio.
  • Speech-translation experiments where the publisher describes multilingual recognition and translation behavior.
  • Long-form transcription using the publisher-described sequential or chunked 30-second handling.

Limitations and unknowns

  • According to the model card, predictions may hallucinate unspoken text, perform unevenly across languages and accents, and produce repetitive text; users should evaluate before deployment.
  • No Ethen-measured evaluation results were extracted; only publisher-described behavior and limitations were captured.
  • No context-window value was extracted from this record.
  • Provider state is historical snapshot data, not independently refreshed current availability.

Source and provenance

Source: openai/whisper-large-v3-turbo

Captured: Unknown. Processed: 2026-09-07T19:34:54.513580+00:00.

Whisper Whisper is a state-of-the-art model for automatic speech recognition (ASR) and speech translation, proposed in the paper Robust Speech Recognition via Large-Scale Weak Supervision by Alec Radford et al. from OpenAI. Trained on >5M hours of labeled data, Whisper demonstrates a strong ability to generalise to many datasets and domains in a zero-shot setting. Whisper large-v3-turbo is a finetuned version of a pruned Whisper large-v3 . In other words, it's the exact same model, except that the number of decoding layers have reduced from 32 to 4. As a result, the model is way faster, at the expense of a minor quality degradation.…

F001F002F003F004F005F006F007F008F009F010F011F012F013F015F022F026F027F028F029F030F031