Skip to content

EthenEthenEthen

Open Source Model Profile · Qwen

Qwen3-235B-A22B-Instruct-2507

Qwen3-235B-A22B-Instruct-2507 is a 235B-parameter Qwen mixture-of-experts instruct model. According to the model card, 22B parameters activate per pass with long-context support.

Publisher
Qwen
Task
text-generation
Model type
qwen3_moe
License
apache-2.0
Library
transformers
Publication status
Approved for indexing

Model overview

Qwen3-235B-A22B-Instruct-2507 is published by Qwen as a mixture-of-experts text-generation model. The captured configuration identifies Qwen3MoeForCausalLM, and Safetensors metadata reports about 235B total parameters under apache-2.0. According to the model card, it is an updated non-thinking instruct release with 22B activated parameters.

Recorded capabilities

235B mixture-of-experts

The captured configuration identifies Qwen3MoeForCausalLM, and the model card reports 235B total parameters with 22B activated across 128 experts.

Instruct update

According to the model card, this 2507 release improves instruction following, reasoning, comprehension, mathematics, science, coding, and tool use over the prior non-thinking mode.

Long-context support

According to the model card, native 262,144-token context extends toward 1M tokens with Dual Chunk Flash Attention serving setups.

Documented 1M-token serving path

According to the model card, 1M-token processing needs about 1000 GB of GPU memory with tensor-parallel vLLM or SGLang launch settings.

Published sampling guidance

According to the model card, recommended sampling uses Temperature 0.7, TopP 0.8, TopK 20, and MinP 0.

Use cases in the source record

  • Instruction-following text generation for reasoning, coding, mathematics, and tool-use workflows described in the card.
  • Long-context document work within the documented 256K native and extended 1M-token serving configurations.

Limitations and unknowns

  • No structured evaluation scores were extracted from this record; benchmark detail is referenced only as external blog, GitHub, and documentation links.
  • Provider state is historical snapshot data covering four providers and should be refreshed before being presented as current.
  • According to the model card, 1M-token operation carries very large memory demands, so deployment sizing needs direct review.

Source and provenance

Source: Qwen/Qwen3-235B-A22B-Instruct-2507

Captured: Unknown. Processed: 2026-09-07T19:34:35.996945+00:00.

Qwen3-235B-A22B-Instruct-2507 Highlights We introduce the updated version of the Qwen3-235B-A22B non-thinking mode , named Qwen3-235B-A22B-Instruct-2507 , featuring the following key enhancements: Significant improvements in general capabilities, including instruction following, logical reasoning, text comprehension, mathematics, science, coding and tool usage . Substantial gains in long-tail knowledge coverage across multiple languages . Markedly better alignment with user preferences in subjective and open-ended tasks , enabling more helpful responses and higher-quality text generation. Enhanced capabilities in 256K long-context u…

F001F002F003F004F005F006F007F009F010F011F013F014F015F018F019F021F022F023F024F028