Qwen3-0.6B lineage
Hub tags list Qwen/Qwen3-0.6B-Base as base model and fine-tune source, and the model card gives a Qwen3-0.6B overview with 0.6B parameters and pretraining plus post-training stages.
Open Source Model Profile · Qwen
Qwen3-0.6B-MLX-bf16 is a 0.6B-parameter Qwen3 text-generation release from Qwen for MLX. According to the model card, it supports switching between thinking and non-thinking modes.
Qwen3-0.6B-MLX-bf16 is published by Qwen as a Qwen3 text-generation model. The captured configuration identifies Qwen3ForCausalLM and Safetensors metadata reports 596049920 parameters. Hub tags list Qwen3-0.6B-Base as base model, the card overview lists 0.6B parameters with pretraining and post-training stages, and card data records apache-2.0.
Hub tags list Qwen/Qwen3-0.6B-Base as base model and fine-tune source, and the model card gives a Qwen3-0.6B overview with 0.6B parameters and pretraining plus post-training stages.
According to the model card, the model supports thinking and non-thinking modes with an enable_thinking switch and /think and /no_think turn-by-turn controls.
According to the model card, Qwen3-0.6B lists 32,768 context length, 28 layers, and 16-for-Q with 8-for-KV attention heads.
The record is tagged with the mlx library, and the model card documents loading this BF16 release through the MLX load workflow with chat-template support.
According to the model card, thinking-mode guidance lists Temperature 0.6, TopP 0.95, TopK 20, and MinP 0, with boxed math formatting and JSON choice formatting for specific tasks.
Source: Qwen/Qwen3-0.6B-MLX-bf16
Captured: Unknown. Processed: 2026-09-07T19:35:51.222324+00:00.
Qwen3-0.6B-MLX-bf16 Qwen3 Highlights Qwen3 is the latest generation of large language models in Qwen series, offering a comprehensive suite of dense and mixture-of-experts (MoE) models. Built upon extensive training, Qwen3 delivers groundbreaking advancements in reasoning, instruction-following, agent capabilities, and multilingual support, with the following key features: Uniquely support of seamless switching between thinking mode (for complex logical reasoning, math, and coding) and non-thinking mode (for efficient, general-purpose dialogue) within single model , ensuring optimal performance across various scenarios. Significantl…
F001F002F003F004F005F006F007F008F009F010F011F012F013F014F015F016F017F018F020F021F022F023