235B mixture-of-experts
The captured configuration identifies Qwen3MoeForCausalLM, and the model card reports 235B total parameters with 22B activated across 128 experts.
Open Source Model Profile · Qwen
Qwen3-235B-A22B-Instruct-2507 is a 235B-parameter Qwen mixture-of-experts instruct model. According to the model card, 22B parameters activate per pass with long-context support.
Qwen3-235B-A22B-Instruct-2507 is published by Qwen as a mixture-of-experts text-generation model. The captured configuration identifies Qwen3MoeForCausalLM, and Safetensors metadata reports about 235B total parameters under apache-2.0. According to the model card, it is an updated non-thinking instruct release with 22B activated parameters.
The captured configuration identifies Qwen3MoeForCausalLM, and the model card reports 235B total parameters with 22B activated across 128 experts.
According to the model card, this 2507 release improves instruction following, reasoning, comprehension, mathematics, science, coding, and tool use over the prior non-thinking mode.
According to the model card, native 262,144-token context extends toward 1M tokens with Dual Chunk Flash Attention serving setups.
According to the model card, 1M-token processing needs about 1000 GB of GPU memory with tensor-parallel vLLM or SGLang launch settings.
According to the model card, recommended sampling uses Temperature 0.7, TopP 0.8, TopK 20, and MinP 0.
Source: Qwen/Qwen3-235B-A22B-Instruct-2507
Captured: Unknown. Processed: 2026-09-07T19:34:35.996945+00:00.
Qwen3-235B-A22B-Instruct-2507 Highlights We introduce the updated version of the Qwen3-235B-A22B non-thinking mode , named Qwen3-235B-A22B-Instruct-2507 , featuring the following key enhancements: Significant improvements in general capabilities, including instruction following, logical reasoning, text comprehension, mathematics, science, coding and tool usage . Substantial gains in long-tail knowledge coverage across multiple languages . Markedly better alignment with user preferences in subjective and open-ended tasks , enabling more helpful responses and higher-quality text generation. Enhanced capabilities in 256K long-context u…
F001F002F003F004F005F006F007F009F010F011F013F014F015F018F019F021F022F023F024F028