Skip to content

EthenEthenEthen

Open Source Model Profile · allenai

tulu-2-dpo-7b

tulu-2-dpo-7b is a Llama-family text-generation assistant fine-tune from allenai. Its model card documents Llama 2 lineage and Direct Preference Optimization on mixed human and synthetic data.

Publisher
allenai
Task
text-generation
Model type
llama
License
ai2-impact-license-low-risk
Library
transformers
Publication status
Accepted · not indexed

Model overview

tulu-2-dpo-7b is published by allenai as a Llama-based text-generation assistant. The captured configuration identifies LlamaForCausalLM with a llama model type. According to the model card, it is a fine-tuned version of Llama 2 trained with Direct Preference Optimization, building on meta-llama/Llama-2-7b-hf.

Recorded capabilities

Helpful-assistant DPO tuning

According to the model card, it is a Llama 2 fine-tune trained on mixed public, synthetic, and human data with Direct Preference Optimization.

Documented alignment lineage

According to the model card, it builds on meta-llama/Llama-2-7b-hf with a Zephyr Beta-derived DPO recipe and UltraFeedback-ranked completions.

Reported MT-Bench and AlpacaEval table

According to the model card, the 7B DPO row reports MT-Bench 6.29 and AlpacaEval 85.1% within a wider Tulu 2 comparison table.

Tulu mix and UltraFeedback stages

According to the model card, it was first fine-tuned on a filtered Tulu V2 mix, then aligned on 64k GPT-4-ranked UltraFeedback prompts with a JAX DPO trainer.

English and ImpACT license note

The card states primarily English and an AI2 ImpACT Low-risk license, matching the captured other license value.

Use cases in the source record

  • Helpful-assistant dialogue workflows consistent with the card's described assistant training.
  • Preference-optimization experiments that follow the documented DPO recipe and reported evaluation format.

Limitations and unknowns

  • No parameter count was extracted from this record.
  • No context-window value was extracted from this record.
  • Benchmark figures come from the publisher model card and have not been independently verified by Ethen.
  • Provider state is historical snapshot data, not independently refreshed current availability.

Source and provenance

Source: allenai/tulu-2-dpo-7b

Captured: Unknown. Processed: 2026-09-07T19:34:39.725022+00:00.

Model Card for Tulu V2 DPO 7B Tulu is a series of language models that are trained to act as helpful assistants. Tulu V2 DPO 7B is a fine-tuned version of Llama 2 that was trained on on a mix of publicly available, synthetic and human datasets using Direct Preference Optimization (DPO) . This model is a strong alternative to Llama 2 7b Chat. For more details, read the paper: Camels in a Changing Climate: Enhancing LM Adaptation with Tulu 2 . Model description Model type: A model belonging to a suite of instruction and RLHF tuned chat models on a mix of publicly available, synthetic and human-created datasets. Language(s) (NLP): Prim…

F001F002F003F004F005F006F008F009F010F011F014F015F016F017F018F019