Skip to content

EthenEthenEthen

Open Source Model Profile · Qwen

Qwen3-ASR-1.7B

Qwen3-ASR-1.7B is a 2.35B-parameter speech-recognition model from Qwen. According to the model card, it handles language identification and ASR across 52 languages and dialects.

Publisher
Qwen
Task
automatic-speech-recognition
Model type
qwen3_asr
License
apache-2.0
Library
Unknown
Publication status
Accepted · not indexed

Model overview

Qwen3-ASR-1.7B is published by Qwen as an automatic-speech-recognition model. The captured configuration identifies Qwen3ASRForConditionalGeneration and Safetensors metadata reports 2349217408 parameters. According to the model card, it belongs to a two-model ASR family built on Qwen3-Omni for language identification and transcription, with card data recording apache-2.0.

Recorded capabilities

52-language ASR coverage

According to the model card, the family supports language identification and speech recognition for 52 languages and dialects, with a table listing 30 languages and 22 Chinese dialects.

Streaming and offline inference

According to the model card, a single model supports streaming and offline unified inference with long-audio transcription, where streaming uses the vLLM backend.

Forced-alignment timestamps

According to the model card, Qwen3-ForcedAligner-0.6B aligns text-speech pairs and returns word or character timestamps for up to 5 minutes of speech in 11 languages.

vLLM and toolkit support

According to the model card, the release includes a full inference framework with vLLM batch inference, async serving, and day-0 vLLM model support, plus Transformers-backend use.

Reported evaluation detail

According to the model card, evaluation used bfloat16 with max_new_tokens 1024, greedy decoding, and no forced language parameter, followed by detailed WER tables including singing-voice sets.

Use cases in the source record

  • Multilingual transcription and language identification across the publisher-listed 30 languages and 22 Chinese dialects.
  • Streaming transcription using the documented vLLM backend and Flask streaming demo path for live microphone input.
  • Word- or character-level timestamping through the documented Qwen3-ForcedAligner-0.6B alignment workflow.

Limitations and unknowns

  • Performance comparisons to open-source and commercial systems are publisher-reported model-card claims and were not independently verified.
  • Provider state is historical snapshot data, not independently refreshed current availability.
  • No pricing, VRAM requirement, or latency values were extracted from this record.

Source and provenance

Source: Qwen/Qwen3-ASR-1.7B

Captured: Unknown. Processed: 2026-09-07T19:34:36.257567+00:00.

Qwen3-ASR Overview Introduction The Qwen3-ASR family includes Qwen3-ASR-1.7B and Qwen3-ASR-0.6B, which support language identification and ASR for 52 languages and dialects. Both leverage large-scale speech training data and the strong audio understanding capability of their foundation model, Qwen3-Omni. Experiments show that the 1.7B version achieves state-of-the-art performance among open-source ASR models and is competitive with the strongest proprietary commercial APIs. Here are the main features: All-in-one : Qwen3-ASR-1.7B and Qwen3-ASR-0.6B support language identification and speech recognition for 30 languages and 22 Chinese…

F001F002F003F004F005F006F007F009F010F011F012F013F015F019F023F026F027F028F033F034F035F037F038