Helpful-assistant DPO tuning
According to the model card, it is a Llama 2 fine-tune trained on mixed public, synthetic, and human data with Direct Preference Optimization.
Open Source Model Profile · allenai
tulu-2-dpo-7b is a Llama-family text-generation assistant fine-tune from allenai. Its model card documents Llama 2 lineage and Direct Preference Optimization on mixed human and synthetic data.
tulu-2-dpo-7b is published by allenai as a Llama-based text-generation assistant. The captured configuration identifies LlamaForCausalLM with a llama model type. According to the model card, it is a fine-tuned version of Llama 2 trained with Direct Preference Optimization, building on meta-llama/Llama-2-7b-hf.
According to the model card, it is a Llama 2 fine-tune trained on mixed public, synthetic, and human data with Direct Preference Optimization.
According to the model card, it builds on meta-llama/Llama-2-7b-hf with a Zephyr Beta-derived DPO recipe and UltraFeedback-ranked completions.
According to the model card, the 7B DPO row reports MT-Bench 6.29 and AlpacaEval 85.1% within a wider Tulu 2 comparison table.
According to the model card, it was first fine-tuned on a filtered Tulu V2 mix, then aligned on 64k GPT-4-ranked UltraFeedback prompts with a JAX DPO trainer.
The card states primarily English and an AI2 ImpACT Low-risk license, matching the captured other license value.
Source: allenai/tulu-2-dpo-7b
Captured: Unknown. Processed: 2026-09-07T19:34:39.725022+00:00.
Model Card for Tulu V2 DPO 7B Tulu is a series of language models that are trained to act as helpful assistants. Tulu V2 DPO 7B is a fine-tuned version of Llama 2 that was trained on on a mix of publicly available, synthetic and human datasets using Direct Preference Optimization (DPO) . This model is a strong alternative to Llama 2 7b Chat. For more details, read the paper: Camels in a Changing Climate: Enhancing LM Adaptation with Tulu 2 . Model description Model type: A model belonging to a suite of instruction and RLHF tuned chat models on a mix of publicly available, synthetic and human-created datasets. Language(s) (NLP): Prim…
F001F002F003F004F005F006F008F009F010F011F014F015F016F017F018F019