72B dense base model
The model card identifies this repository as the 72B Qwen2 base language model, matching about 72.71B captured parameters.
Open Source Model Profile · Qwen
Qwen2-72B is a Qwen 72.71B base text-generation model. Its card documents a dense Transformer design and post-training guidance.
Qwen2-72B is published by Qwen as a text-generation base model. The captured configuration identifies Qwen2ForCausalLM with Safetensors metadata reporting 72,706,203,648 parameters. According to the model card, this repository holds the 72B Qwen2 base language model in a series ranging from 0.5 to 72 billion parameters.
The model card identifies this repository as the 72B Qwen2 base language model, matching about 72.71B captured parameters.
According to the model card, the series uses SwiGLU activation, attention QKV bias, group query attention, and an improved tokenizer for multiple natural languages and code.
The card advises against direct text generation with the base model and recommends SFT, RLHF, or continued pretraining.
According to the model card, documented suites span language understanding, coding, mathematics, Chinese, and multilingual tasks with a multi-model comparison table.
Source: Qwen/Qwen2-72B
Captured: Unknown. Processed: 2026-09-07T19:34:35.754543+00:00.
Qwen2-72B Introduction Qwen2 is the new series of Qwen large language models. For Qwen2, we release a number of base language models and instruction-tuned language models ranging from 0.5 to 72 billion parameters, including a Mixture-of-Experts model. This repo contains the 72B Qwen2 base language model. Compared with the state-of-the-art opensource language models, including the previous released Qwen1.5, Qwen2 has generally surpassed most opensource models and demonstrated competitiveness against proprietary models across a series of benchmarks targeting for language understanding, language generation, multilingual capability, cod…
F001F002F003F004F005F006F007F008F009F010F011F012F014F015F016