52-language ASR coverage
According to the model card, the family supports language identification and speech recognition for 52 languages and dialects, with a table listing 30 languages and 22 Chinese dialects.
Open Source Model Profile · Qwen
Qwen3-ASR-1.7B is a 2.35B-parameter speech-recognition model from Qwen. According to the model card, it handles language identification and ASR across 52 languages and dialects.
Qwen3-ASR-1.7B is published by Qwen as an automatic-speech-recognition model. The captured configuration identifies Qwen3ASRForConditionalGeneration and Safetensors metadata reports 2349217408 parameters. According to the model card, it belongs to a two-model ASR family built on Qwen3-Omni for language identification and transcription, with card data recording apache-2.0.
According to the model card, the family supports language identification and speech recognition for 52 languages and dialects, with a table listing 30 languages and 22 Chinese dialects.
According to the model card, a single model supports streaming and offline unified inference with long-audio transcription, where streaming uses the vLLM backend.
According to the model card, Qwen3-ForcedAligner-0.6B aligns text-speech pairs and returns word or character timestamps for up to 5 minutes of speech in 11 languages.
According to the model card, the release includes a full inference framework with vLLM batch inference, async serving, and day-0 vLLM model support, plus Transformers-backend use.
According to the model card, evaluation used bfloat16 with max_new_tokens 1024, greedy decoding, and no forced language parameter, followed by detailed WER tables including singing-voice sets.
Source: Qwen/Qwen3-ASR-1.7B
Captured: Unknown. Processed: 2026-09-07T19:34:36.257567+00:00.
Qwen3-ASR Overview Introduction The Qwen3-ASR family includes Qwen3-ASR-1.7B and Qwen3-ASR-0.6B, which support language identification and ASR for 52 languages and dialects. Both leverage large-scale speech training data and the strong audio understanding capability of their foundation model, Qwen3-Omni. Experiments show that the 1.7B version achieves state-of-the-art performance among open-source ASR models and is competitive with the strongest proprietary commercial APIs. Here are the main features: All-in-one : Qwen3-ASR-1.7B and Qwen3-ASR-0.6B support language identification and speech recognition for 30 languages and 22 Chinese…
F001F002F003F004F005F006F007F009F010F011F012F013F015F019F023F026F027F028F033F034F035F037F038