Skip to content

EthenEthenEthen

Open Source Model Profile · chujiezheng

Mistral7B-PairRM-SPPO-ExPO

This chujiezheng model is a 7.24B-parameter Mistral text-generation release. The card describes ExPO extrapolation with alpha 0.3 and reports AlpacaEval 2.0 figures.

Publisher
chujiezheng
Task
text-generation
Model type
mistral
License
apache-2.0
Library
transformers
Publication status
Accepted · not indexed

Model overview

Mistral7B-PairRM-SPPO-ExPO is published by chujiezheng as a text-generation model. The captured configuration identifies MistralForCausalLM with a mistral model type, and Safetensors metadata reports 7,241,732,096 parameters. According to the model card, it is an ExPO extrapolation of UCLA-AGI/Mistral7B-PairRM-SPPO and Mistral-7B-Instruct-v0.2 with alpha of 0.3, and the captured license is Apache-2.0.

Recorded capabilities

ExPO extrapolation method

According to the card, the model extrapolates from SFT and DPO/RLHF checkpoint weights with alpha of 0.3, following the Weak-to-Strong Extrapolation paper.

Reported AlpacaEval comparison

The card reports 35.4% win rate and 31.8% length-controlled win rate on AlpacaEval 2.0, against 32.2% and 30.5% for the original Mistral7B-PairRM-SPPO.

Mistral 7.24B configuration

Captured config identifies MistralForCausalLM and mistral, with Safetensors metadata reporting 7,241,732,096 parameters.

Apache-2.0 licensing

The captured record and hub tags both list an Apache-2.0 license.

Use cases in the source record

  • Conversational text generation through a transformers-compatible Mistral workflow.
  • Alignment-extrapolation research using the card's documented ExPO setup and reported AlpacaEval comparison.

Limitations and unknowns

  • No independent Ethen evaluation results were extracted from this record.
  • No context-window value, hardware requirement, or quantization detail was extracted.
  • Provider state is historical snapshot data, not independently refreshed current availability.
  • Extrapolation and benchmark figures come from the publisher card and were not independently verified by Ethen.

Source and provenance

Source: chujiezheng/Mistral7B-PairRM-SPPO-ExPO

Captured: Unknown. Processed: 2026-09-07T19:34:42.537685+00:00.

Mistral7B-PairRM-SPPO-ExPO The extrapolated (ExPO) model based on UCLA-AGI/Mistral7B-PairRM-SPPO and mistralai/Mistral-7B-Instruct-v0.2 , as in the " Weak-to-Strong Extrapolation Expedites Alignment " paper. Specifically, we obtain this model by extrapolating (alpha = 0.3) from the weights of the SFT and DPO/RLHF checkpoints, achieving superior alignment with human preference. This extrapolated model achieves the 35.4% win rate and 31.8% LC win rate on AlpacaEval 2.0 , outperforming the original Mistral7B-PairRM-SPPO 's 32.2% and 30.5%, respectively. Evaluation Results Evaluation results on the AlpacaEval 2.0 benchmark (you can find…

F001F002F003F004F005F006F007F008F009F010F011F012F013F014F015