Qwen2 32.76B fine-tune
Safetensors metadata reports 32,763,876,352 parameters with a Qwen2ForCausalLM configuration on the documented DeepSeek-R1-Distill-Qwen-32B base.
Open Source Model Profile · Tongyi-Zhiwen
QwenLong-L1-32B is a long-context reasoning fine-tune from Tongyi-Zhiwen. Safetensors metadata reports about 32.76B Qwen2 parameters with documented RL training.
QwenLong-L1-32B is published by Tongyi-Zhiwen as a Qwen2-family text-generation fine-tune. Safetensors metadata reports 32,763,876,352 parameters. According to the model card, it applies a reinforcement-learning framework for long-context reasoning on the DeepSeek-R1-Distill-Qwen-32B base.
Safetensors metadata reports 32,763,876,352 parameters with a Qwen2ForCausalLM configuration on the documented DeepSeek-R1-Distill-Qwen-32B base.
According to the model card, the publisher built a 1.6K-problem DocQA dataset covering mathematical, logical, and multi-hop reasoning.
The model card documents YaRN RoPE scaling across transformers, vLLM, SGLang, and llama.cpp for inputs up to 131,072 tokens.
Source: Tongyi-Zhiwen/QwenLong-L1-32B
Captured: Unknown. Processed: 2026-09-07T19:35:52.619164+00:00.
QwenLong-L1: Towards Long-Context Large Reasoning Models with Reinforcement Learning Fanqi Wan, Weizhou Shen, Shengyi Liao, Yingcheng Shi, Chenliang Li, Ziyi Yang, Ji Zhang, Fei Huang, Jingren Zhou, Ming Yan Tongyi Lab, Alibaba Group 🎉 News May 28, 2025: 🔥 We release 🤗 QwenLong-L1-32B-AWQ , which has undergone AWQ int4 quantization using the ms-swift framework. May 26, 2025: 🔥 We release 🤗 QwenLong-L1-32B , which is the first long-context LRM trained with reinforcement learning for long-context reasoning. Experiments on seven long-context DocQA benchmarks demonstrate that QwenLong-L1-32B outperforms flagship LRMs like OpenAI-o3…
F001F002F003F004F005F006F007F008F009F010F011F012F013F014F016F017F018F019F020F021F023F028F031