7.62B Qwen2 scale
Captured config identifies Qwen2ForCausalLM and Safetensors metadata reports 7615616512 parameters.
Open Source Model Profile · reasonrag
Qwen2.5-7B-Instruct-ReasonRAG is a 7.62B-parameter Qwen2 text-generation fine-tune from reasonrag. According to the model card, it fine-tunes Qwen2.5-7B-Instruct on a RAG-related DPO dataset.
Qwen2.5-7B-Instruct-ReasonRAG is published by reasonrag as a conversational text-generation model. The captured configuration identifies Qwen2ForCausalLM with a qwen2 model type, and Safetensors metadata reports 7615616512 parameters. The model card describes it as a fine-tuned version of Qwen/Qwen2.5-7B-Instruct on dpo_mcts_rag_v8, consistent with hub base-model tags.
Captured config identifies Qwen2ForCausalLM and Safetensors metadata reports 7615616512 parameters.
According to the model card, the model fine-tunes Qwen/Qwen2.5-7B-Instruct on dpo_mcts_rag_v8.
The model card reports an evaluation loss of 0.8564 with accompanying reward and log-probability values.
Source: reasonrag/Qwen2.5-7B-Instruct-ReasonRAG
Captured: Unknown. Processed: 2026-09-07T19:35:29.251529+00:00.
dpo_v16 This model is a fine-tuned version of Qwen/Qwen2.5-7B-Instruct on the dpo_mcts_rag_v8 dataset. It achieves the following results on the evaluation set: Loss: 0.8564 Rewards/chosen: 1.0146 Rewards/rejected: -0.5767 Rewards/accuracies: 0.6204 Rewards/margins: 1.5913 Logps/chosen: -65.8051 Logps/rejected: -74.8917 Logits/chosen: -0.5108 Logits/rejected: -0.5206 Model description More information needed Intended uses & limitations More information needed Training and evaluation data More information needed Training procedure Training hyperparameters The following hyperparameters were used during training: learning_rate: 1e-06 tr…
F001F002F003F004F005F006F007F008F009F010F011F012F013