Qwen2 math fine-tune record
Captured configuration records Qwen2ForCausalLM with model type qwen2 and about 7.62B parameters, with hub tags marking a Qwen2.5-Math-7B finetune.
Open Source Model Profile · Yuqian-Fu
SRFT-Qwen2.5-Math-7B is a 7.62B-parameter Qwen2 math fine-tune from Yuqian-Fu. Its card documents the single-stage SRFT method under MIT licensing.
SRFT-Qwen2.5-Math-7B is published by Yuqian-Fu as a Qwen2 text-generation model. The captured configuration identifies Qwen2ForCausalLM with model type qwen2, and Safetensors metadata reports 7,615,616,512 parameters. According to the model card, it applies Supervised Reinforcement Fine-Tuning, while hub tags record a math dataset and a Qwen2.5-Math-7B finetune marker.
Captured configuration records Qwen2ForCausalLM with model type qwen2 and about 7.62B parameters, with hub tags marking a Qwen2.5-Math-7B finetune.
According to the model card, SRFT is a single-stage method unifying fine-tuning paradigms through entropy-aware weighting.
Card data records MIT licensing with Transformers and safetensors support and conversational text-generation markers.
Source: Yuqian-Fu/SRFT-Qwen2.5-Math-7B
Captured: Unknown. Processed: 2026-09-07T19:35:53.208015+00:00.
📄 Introduction Supervised Reinforcement Fine-Tuning (SRFT) is a single-stage method that unifies both fine-tuning paradigms through entropy-aware weighting mechanisms. Paper: arXiv Project Website: SRFT
F001F002F003F004F005F006F007F008F009F010F011