Skip to content

EthenEthenEthen

Open Source Model Profile · Qwen

Qwen2.5-1.5B

Qwen2.5-1.5B is a 1.54B-parameter Qwen2-family text-generation base model from Qwen. According to the model card, it is the pretraining-stage 1.5B model with 32,768-token context.

Publisher
Qwen
Task
text-generation
Model type
qwen2
License
apache-2.0
Library
transformers
Publication status
Accepted · not indexed

Model overview

Qwen2.5-1.5B is published by Qwen as a text-generation base model. The captured configuration identifies Qwen2ForCausalLM with model type qwen2, and Safetensors metadata reports about 1.54B parameters. According to the model card, it is the base 1.5B Qwen2.5 model with RoPE, SwiGLU, RMSNorm, attention QKV bias, and tied word embeddings.

Recorded capabilities

Qwen2 causal-LM architecture

The captured configuration identifies Qwen2ForCausalLM with model type qwen2 and Transformers support.

1.54B base-model scale

Safetensors metadata reports 1,543,714,304 parameters; the model card states 1.54B total, 1.31B non-embedding, 28 layers, and grouped-query attention.

Long-context documentation

According to the model card, the model has 32,768-token full context length with support up to 128K tokens and generation up to 8K tokens.

Instruction and structured-output focus

According to the model card, the publisher describes improved instruction following, long-text generation, structured-data understanding, JSON output, and system-prompt resilience.

Post-training-oriented base model

According to the model card, this is a pretraining-stage base model; the publisher recommends SFT, RLHF, or continued pretraining rather than direct conversational use.

Use cases in the source record

  • Base-model fine-tuning for instruction following, long-text generation, and JSON structured-output tasks described by the publisher.
  • Coding- and mathematics-oriented post-training building on the publisher-described Qwen2.5 improvements.

Limitations and unknowns

  • No Ethen-measured benchmark results were extracted; series-level improvement claims are publisher descriptions.
  • According to the model card, GPU-memory and throughput details are referenced externally rather than stated in the captured text.
  • Provider state is historical snapshot data and should be refreshed before being presented as current.

Source and provenance

Source: Qwen/Qwen2.5-1.5B

Captured: Unknown. Processed: 2026-09-07T19:34:35.914220+00:00.

Qwen2.5-1.5B Introduction Qwen2.5 is the latest series of Qwen large language models. For Qwen2.5, we release a number of base language models and instruction-tuned language models ranging from 0.5 to 72 billion parameters. Qwen2.5 brings the following improvements upon Qwen2: Significantly more knowledge and has greatly improved capabilities in coding and mathematics , thanks to our specialized expert models in these domains. Significant improvements in instruction following , generating long texts (over 8K tokens), understanding structured data (e.g, tables), and generating structured outputs especially JSON. More resilient to the…

F001F002F003F004F005F006F007F008F010F011F012F013F014F015F016F017