SimPO-linked release
According to the model card, this model was released from the SimPO preprint, with hub tags referencing arXiv 2405.14734.
Open Source Model Profile · princeton-nlp
Llama-3-Instruct-8B-RRHF-v0.2 is an 8.03B-parameter Llama text-generation model from princeton-nlp. According to the model card, it is linked to the SimPO preference-optimization preprint.
Llama-3-Instruct-8B-RRHF-v0.2 is published by princeton-nlp as a Transformers text-generation model. The captured configuration identifies LlamaForCausalLM with model type llama and about 8.03B Safetensors parameters. According to the model card, it was released from the SimPO preprint.
According to the model card, this model was released from the SimPO preprint, with hub tags referencing arXiv 2405.14734.
Captured configuration identifies LlamaForCausalLM with 8030261248 Safetensors parameters.
Hub metadata marks text-generation with Transformers library support and conversational, endpoints-compatible tags.
Source: princeton-nlp/Llama-3-Instruct-8B-RRHF-v0.2
Captured: Unknown. Processed: 2026-09-07T19:34:55.787275+00:00.
This is a model released from the preprint: SimPO: Simple Preference Optimization with a Reference-Free Reward . Please refer to our repository for more details.
F001F002F003F004F005F006F007F008F009F010