Skip to content

EthenEthenEthen

Open Source Model Profile · ContextualAI

LMUnit-llama3.1-70b

LMUnit-llama3.1-70b is a 70.55B-parameter Llama evaluation model from ContextualAI that scores responses against natural-language unit tests.

Publisher
ContextualAI
Task
text-generation
Model type
llama
License
Unknown
Library
transformers
Publication status
Accepted · not indexed

Model overview

LMUnit-llama3.1-70b is published by ContextualAI as a Llama text-generation model for evaluation. The captured configuration identifies LlamaForCausalLM and Safetensors metadata reports 70,553,706,496 parameters. According to the model card, it is finetuned from Llama-3.1-70B-Instruct to score prompt-response pairs against natural-language unit tests.

Recorded capabilities

Natural-language unit-test scoring

According to the model card, the model takes a prompt, a response, and a unit test and produces a continuous score between 1 and 5.

Publisher-reported evaluation results

According to the model card, the publisher reports 93.5% RewardBench accuracy, 82.1% RewardBench2 accuracy, and results on FLASK, BiGGen Bench, LFQA, and related sets.

Multi-objective and synthetic-data training

According to the model card, training combines pairwise comparisons, direct ratings, and criteria-based judgments with a synthetic-data pipeline for fine-grained criteria.

LMUnit, vLLM, and Transformers usage

According to the model card, the publisher documents a dedicated lmunit package with vLLM sampling as well as direct Transformers loading.

Use cases in the source record

  • Fine-grained response evaluation where a prompt, response, and natural-language unit test are scored on a 1-to-5 scale.
  • Preference, direct-scoring, and benchmark-style evaluation experiments using the publisher's reported FLASK, BiGGen Bench, and RewardBench setup.
  • Evaluation pipelines built with the publisher's documented lmunit package, vLLM sampling, or Transformers loading.

Limitations and unknowns

  • No license value was extracted from this record.
  • Evaluation figures are publisher-reported claims from the model card and were not independently verified by Ethen.
  • Provider state is historical snapshot data and should be refreshed before being presented as current.
  • No context-window value, pricing, or VRAM requirement was extracted.

Source and provenance

Source: ContextualAI/LMUnit-llama3.1-70b

Captured: Unknown. Processed: 2026-09-07T19:35:38.562744+00:00.

LMUnit: Fine-grained Evaluation with Natural Language Unit Tests LMUnit is a state-of-the-art language model that is optimized for evaluating natural language unit tests. It takes three inputs: a prompt, a response, and a unit test. It then produces a continuous score between 1 and 5 where higher scores indicate that the response better satisfies the unit test criteria. The LMUnit model achieves leading averaged performance across preference, direct scoring, and fine-grained unit test evaluation tasks, as measured by FLASK and BiGGen Bench, and performs on par with frontier models for coarse evaluation of long-form responses (per LF…

F001F002F003F004F005F006F009F010F011F013F014F015F016F017