Skip to content

EthenEthenEthen

Open Source Model Profile · allenai

tulu-v2.5-ppo-13b-uf-mean

tulu-v2.5-ppo-13b-uf-mean is a 13.02B-parameter Llama assistant model from allenai. Its model card documents PPO training on UltraFeedback with a 13B reward model.

Publisher
allenai
Task
text-generation
Model type
llama
License
apache-2.0
Library
transformers
Publication status
Accepted · not indexed

Model overview

tulu-v2.5-ppo-13b-uf-mean is published by allenai as a Llama-based text-generation model. The captured configuration identifies LlamaForCausalLM and Safetensors metadata reports 13015864320 parameters. According to the model card, it belongs to the Tulu helpful-assistant series and was trained with PPO on the UltraFeedback dataset starting from the Tulu 2 suite.

Recorded capabilities

13.02B Llama scale

Captured config identifies LlamaForCausalLM and Safetensors metadata reports 13015864320 parameters.

Helpful-assistant purpose

According to the model card, Tulu models are trained to act as helpful assistants.

UltraFeedback PPO

According to the model card, this model was trained on the UltraFeedback dataset using PPO with per-aspect fine-grained scores deciding chosen and rejected responses.

13B reward model and Jax PPO

According to the model card, a 13B reward model trained on UltraFeedback data was used, with later alignment via a Jax PPO trainer built on EasyLM.

Documented input format

According to the model card, the model is trained to use a specific newline-sensitive input format.

Use cases in the source record

  • English helpful-assistant text generation using the card's documented newline-sensitive input format.
  • Preference-alignment research building on the documented UltraFeedback PPO recipe and hyperparameters.

Limitations and unknowns

  • No scored evaluation results were extracted from this record.
  • No context-window value was extracted from this record.
  • Provider state is historical snapshot data, not independently refreshed current availability.
  • Linked reward-model, value-model, and dataset locations are referenced but not resolved in this record, so their current state is unknown.

Source and provenance

Source: allenai/tulu-v2.5-ppo-13b-uf-mean

Captured: Unknown. Processed: 2026-09-07T19:34:39.751237+00:00.

Model Card for Tulu V2.5 PPO 13B - UltraFeedback Mean Tulu is a series of language models that are trained to act as helpful assistants. Tulu V2.5 is a series of models trained using DPO and PPO starting from the Tulu 2 suite . This model is trained on the UltraFeedback dataset (using the per-aspect/fine-grained scores for deciding chosen and rejected) using PPO. We used a 13B RM trained on the UltraFeedback data, and then re-used the same prompts during PPO training. For more details, read the paper: Unpacking DPO and PPO: Disentangling Best Practices for Learning from Preference Feedback . .Model description Model type: One model…

F001F002F003F004F005F006F007F009F010F011F012F013F014F016F017F018F019F021F022