Qwen2 beta positioning
According to the card, Qwen1.5 is the beta version of Qwen2, spanning dense sizes plus a 14B MoE entry.
Open Source Model Profile · Qwen
Qwen1.5-4B is a 3.95B-parameter Qwen2 text-generation base model from Qwen. The card describes Qwen1.5 as the beta version of Qwen2.
Qwen1.5-4B is published by Qwen as a text-generation base model. The captured configuration identifies Qwen2ForCausalLM with a qwen2 model type, and Safetensors metadata reports 3,950,369,280 parameters. According to the model card, Qwen1.5 is the beta version of Qwen2, and the captured license value is other.
According to the card, Qwen1.5 is the beta version of Qwen2, spanning dense sizes plus a 14B MoE entry.
The card claims stable 32K context-length support for models of all sizes.
The card describes a Transformer decoder with SwiGLU activation, QKV bias, group query attention, sliding-window plus full attention, and an improved multilingual tokenizer.
The card advises transformers 4.37.0 or newer and recommends post-training such as SFT or RLHF rather than direct base-model generation.
Source: Qwen/Qwen1.5-4B
Captured: Unknown. Processed: 2026-09-07T19:34:35.967337+00:00.
Qwen1.5-4B Introduction Qwen1.5 is the beta version of Qwen2, a transformer-based decoder-only language model pretrained on a large amount of data. In comparison with the previous released Qwen, the improvements include: 8 model sizes, including 0.5B, 1.8B, 4B, 7B, 14B, 32B and 72B dense models, and an MoE model of 14B with 2.7B activated; Significant performance improvement in Chat models; Multilingual support of both base and chat models; Stable support of 32K context length for models of all sizes No need of trust_remote_code . For more details, please refer to our blog post and GitHub repo . Model Details Qwen1.5 is a language m…
F001F002F003F004F005F006F007F009F010F011F013F014F015