Skip to content

EthenEthenEthen

Open Source Model Profile · reasonrag

Qwen2.5-7B-Instruct-ReasonRAG

Qwen2.5-7B-Instruct-ReasonRAG is a 7.62B-parameter Qwen2 text-generation fine-tune from reasonrag. According to the model card, it fine-tunes Qwen2.5-7B-Instruct on a RAG-related DPO dataset.

Publisher
reasonrag
Task
text-generation
Model type
qwen2
License
other
Library
transformers
Publication status
Accepted · not indexed

Model overview

Qwen2.5-7B-Instruct-ReasonRAG is published by reasonrag as a conversational text-generation model. The captured configuration identifies Qwen2ForCausalLM with a qwen2 model type, and Safetensors metadata reports 7615616512 parameters. The model card describes it as a fine-tuned version of Qwen/Qwen2.5-7B-Instruct on dpo_mcts_rag_v8, consistent with hub base-model tags.

Recorded capabilities

7.62B Qwen2 scale

Captured config identifies Qwen2ForCausalLM and Safetensors metadata reports 7615616512 parameters.

RAG DPO fine-tune

According to the model card, the model fine-tunes Qwen/Qwen2.5-7B-Instruct on dpo_mcts_rag_v8.

Reported training figures

The model card reports an evaluation loss of 0.8564 with accompanying reward and log-probability values.

Use cases in the source record

  • Conversational text generation consistent with the captured conversational and text-generation tags.
  • DPO-training analysis using the card's reported loss, reward, and log-probability figures.

Limitations and unknowns

  • According to the model card, intended uses, limitations, and training-data sections still need more information.
  • No context-window value was extracted from this record.
  • Reported loss and reward figures are publisher claims and have not been independently verified by Ethen.
  • Provider state is historical snapshot data, not independently refreshed current availability.

Source and provenance

Source: reasonrag/Qwen2.5-7B-Instruct-ReasonRAG

Captured: Unknown. Processed: 2026-09-07T19:35:29.251529+00:00.

dpo_v16 This model is a fine-tuned version of Qwen/Qwen2.5-7B-Instruct on the dpo_mcts_rag_v8 dataset. It achieves the following results on the evaluation set: Loss: 0.8564 Rewards/chosen: 1.0146 Rewards/rejected: -0.5767 Rewards/accuracies: 0.6204 Rewards/margins: 1.5913 Logps/chosen: -65.8051 Logps/rejected: -74.8917 Logits/chosen: -0.5108 Logits/rejected: -0.5206 Model description More information needed Intended uses & limitations More information needed Training and evaluation data More information needed Training procedure Training hyperparameters The following hyperparameters were used during training: learning_rate: 1e-06 tr…

F001F002F003F004F005F006F007F008F009F010F011F012F013