Skip to content

EthenEthenEthen

Open Source Model Profile · Qwen

Qwen2.5-3B

Qwen2.5-3B is a 3.09B-parameter Qwen causal base model for pretraining-stage text generation. Its model card documents a 32-layer GQA design with 32K-token context.

Publisher
Qwen
Task
text-generation
Model type
qwen2
License
qwen-research
Library
Unknown
Publication status
Accepted · not indexed

Model overview

Qwen2.5-3B is published by Qwen as a text-generation base model. The captured configuration identifies Qwen2ForCausalLM with model type qwen2, and Safetensors metadata reports 3,085,938,688 parameters. According to the model card, this repository contains the base 3B model at the pretraining stage, and card data records other for license.

Recorded capabilities

3B Qwen2.5 base release

According to the model card, this is the base 3B Qwen2.5 causal language model at the pretraining stage, within a series spanning 0.5B to 72B parameters.

Documented GQA architecture

The model card documents RoPE, SwiGLU, RMSNorm, QKV bias, tied embeddings, 36 layers, and grouped-query attention with 16 Q heads and 2 KV heads.

Long-context and structured output

According to the model card, the series improves instruction following, long-text generation over 8K tokens, structured-data understanding, and JSON outputs, with up to 128K-token support.

Post-training starting point

The model card positions the base weights for SFT, RLHF, or continued pretraining rather than direct conversational use.

Use cases in the source record

  • Post-training research such as SFT, RLHF, or continued pretraining on the base weights, per the publisher's recommended path.
  • Long-context and structured-output experiments within the documented 128K-token support and 8K-token generation budget.

Limitations and unknowns

  • No VRAM or hardware requirement was extracted from this record.
  • Provider state is historical snapshot data and should be refreshed before being presented as current.
  • No evaluation scores were extracted from this record.

Source and provenance

Source: Qwen/Qwen2.5-3B

Captured: Unknown. Processed: 2026-09-07T19:34:36.013534+00:00.

Qwen2.5-3B Introduction Qwen2.5 is the latest series of Qwen large language models. For Qwen2.5, we release a number of base language models and instruction-tuned language models ranging from 0.5 to 72 billion parameters. Qwen2.5 brings the following improvements upon Qwen2: Significantly more knowledge and has greatly improved capabilities in coding and mathematics , thanks to our specialized expert models in these domains. Significant improvements in instruction following , generating long texts (over 8K tokens), understanding structured data (e.g, tables), and generating structured outputs especially JSON. More resilient to the d…

F001F002F003F004F005F006F007F009F010F011F013F014F015