Skip to content

EthenEthenEthen

Open Source Model Profile · Tongyi-Zhiwen

QwenLong-L1-32B

QwenLong-L1-32B is a long-context reasoning fine-tune from Tongyi-Zhiwen. Safetensors metadata reports about 32.76B Qwen2 parameters with documented RL training.

Publisher
Tongyi-Zhiwen
Task
text-generation
Model type
qwen2
License
apache-2.0
Library
transformers
Publication status
Accepted · not indexed

Model overview

QwenLong-L1-32B is published by Tongyi-Zhiwen as a Qwen2-family text-generation fine-tune. Safetensors metadata reports 32,763,876,352 parameters. According to the model card, it applies a reinforcement-learning framework for long-context reasoning on the DeepSeek-R1-Distill-Qwen-32B base.

Recorded capabilities

Qwen2 32.76B fine-tune

Safetensors metadata reports 32,763,876,352 parameters with a Qwen2ForCausalLM configuration on the documented DeepSeek-R1-Distill-Qwen-32B base.

DocQA-RL-1.6K training set

According to the model card, the publisher built a 1.6K-problem DocQA dataset covering mathematical, logical, and multi-hop reasoning.

YaRN long-document path

The model card documents YaRN RoPE scaling across transformers, vLLM, SGLang, and llama.cpp for inputs up to 131,072 tokens.

Use cases in the source record

  • Long-document question-answering experiments across mathematical, logical, and multi-hop reasoning problems described in the card.
  • Deployment trials with YaRN-scaled serving stacks such as vLLM, SGLang, transformers, or llama.cpp for long inputs.

Limitations and unknowns

  • No evaluation results were extracted as structured data from this record.
  • No verified current context-window value was extracted; long-context figures above come from publisher documentation.
  • Provider state is historical snapshot data, not independently refreshed current availability.
  • Publisher comparisons against other reasoning models are source claims and have not been independently verified by Ethen.

Source and provenance

Source: Tongyi-Zhiwen/QwenLong-L1-32B

Captured: Unknown. Processed: 2026-09-07T19:35:52.619164+00:00.

QwenLong-L1: Towards Long-Context Large Reasoning Models with Reinforcement Learning Fanqi Wan, Weizhou Shen, Shengyi Liao, Yingcheng Shi, Chenliang Li, Ziyi Yang, Ji Zhang, Fei Huang, Jingren Zhou, Ming Yan Tongyi Lab, Alibaba Group 🎉 News May 28, 2025: 🔥 We release 🤗 QwenLong-L1-32B-AWQ , which has undergone AWQ int4 quantization using the ms-swift framework. May 26, 2025: 🔥 We release 🤗 QwenLong-L1-32B , which is the first long-context LRM trained with reinforcement learning for long-context reasoning. Experiments on seven long-context DocQA benchmarks demonstrate that QwenLong-L1-32B outperforms flagship LRMs like OpenAI-o3…

F001F002F003F004F005F006F007F008F009F010F011F012F013F014F016F017F018F019F020F021F023F028F031