Skip to content

EthenEthenEthen

Open Source Model Profile · Qwen

Qwen3-ASR-0.6B

Qwen3-ASR-0.6B is a 938M-parameter Qwen speech-recognition model from Qwen. The model card describes 52-language and dialect coverage with streaming and offline inference.

Publisher
Qwen
Task
automatic-speech-recognition
Model type
qwen3_asr
License
apache-2.0
Library
Unknown
Publication status
Accepted · not indexed

Model overview

Qwen3-ASR-0.6B is published by Qwen as an automatic-speech-recognition model. The captured configuration identifies Qwen3ASRForConditionalGeneration with a qwen3_asr model type and Safetensors metadata reports 938008576 parameters. According to the model card, it is the smaller member of the Qwen3-ASR family built on Qwen3-Omni.

Recorded capabilities

Multilingual ASR coverage

According to the model card, the 0.6B model supports language identification and recognition for 30 languages and 22 Chinese dialects.

Streaming and offline inference

According to the model card, both family models support streaming and offline unified inference, with streaming described as vLLM-backend only.

vLLM and toolkit support

The model card describes a full-featured inference framework with vLLM batch inference, async serving, and streaming, plus day-0 vLLM model support.

Forced-alignment companion

The model card introduces Qwen3-ForcedAligner-0.6B for word- or character-level timestamps from text-speech pairs.

Use cases in the source record

  • Multilingual transcription and language identification across the card-listed languages and Chinese dialects.
  • Batch, served, and streaming transcription workflows using the documented qwen-asr transformers and vLLM backends.

Limitations and unknowns

  • No independent Ethen evaluation results were extracted; published WER tables mainly describe Qwen3-ASR-1.7B and should not be read as 0.6B scores.
  • No context-window, VRAM, latency, or pricing values were extracted from this record.
  • Provider state is historical snapshot data, not independently refreshed current availability.
  • Training, foundation-model, and evaluation-setup details come from the publisher model card and have not been independently verified by Ethen.

Source and provenance

Source: Qwen/Qwen3-ASR-0.6B

Captured: Unknown. Processed: 2026-09-07T19:34:36.232366+00:00.

Qwen3-ASR Overview Introduction The Qwen3-ASR family includes Qwen3-ASR-1.7B and Qwen3-ASR-0.6B, which support language identification and ASR for 52 languages and dialects. Both leverage large-scale speech training data and the strong audio understanding capability of their foundation model, Qwen3-Omni. Experiments show that the 1.7B version achieves state-of-the-art performance among open-source ASR models and is competitive with the strongest proprietary commercial APIs. Here are the main features: All-in-one : Qwen3-ASR-1.7B and Qwen3-ASR-0.6B support language identification and speech recognition for 30 languages and 22 Chinese…

F001F002F003F004F005F006F007F009F010F011F012F013F015F019F023F026F027F034F037