Skip to content

EthenEthenEthen

Open Source Model Profile · nvidia

Llama-3.3-Nemotron-70B-Select

Llama-3.3-Nemotron-70B-Select is a 70.55B-parameter Llama text-generation fine-tune from nvidia. According to the model card, it selects the most helpful response for Feedback-Edit inference-time scaling.

Publisher
nvidia
Task
text-generation
Model type
llama
License
nvidia-open-model-license
Library
transformers
Publication status
Accepted · not indexed

Model overview

Llama-3.3-Nemotron-70B-Select is published by nvidia as a Llama text-generation model. The captured configuration identifies LlamaForCausalLM and Safetensors metadata reports 70,553,706,496 parameters. According to the model card, it builds on Meta-Llama-3.3-70B-Instruct with scaled Bradley-Terry modeling to select the most helpful generated response.

Recorded capabilities

Select-model role

According to the model card, the model selects the most helpful LLM-generated response to user queries.

70.55B Llama with Transformers

The captured configuration reports LlamaForCausalLM with model type llama and 70,553,706,496 parameters, with Transformers library support.

128k-token input, float output

According to the model card, input is text up to 128k tokens and output is a single float quality score.

HelpSteer3 lineage

According to the model card, v1.0 training and testing use nvidia/HelpSteer3, also tagged in the hub metadata.

Documented A100 test setup

According to the model card, sample code was tested with Transformers v4.45.0 on two A100 80GB GPUs, with Triton inference on H100 and A100 hardware.

Use cases in the source record

  • Response selection in Feedback-Edit inference-time-scaling workflows for general-domain, open-ended tasks, as described by the publisher.
  • Quality-scoring experiments where a prompt-response pair receives a single float score, with higher values meaning higher quality.

Limitations and unknowns

  • No Ethen-measured evaluation results were extracted from this record.
  • According to the model card, Arena Hard figures for Feedback-Edit ITS are publisher-reported as of 18 Mar 2025 and are not Ethen-measured.
  • According to the model card, the model may amplify toxic biases and generate inaccurate, incomplete, or undesirable text.
  • The captured card data lists license value other; according to the model card, use is governed by the NVIDIA Open Model License with Llama 3.3 Community terms, which should be reviewed directly.
  • Provider state is historical snapshot data, not independently refreshed current availability.

Source and provenance

Source: nvidia/Llama-3.3-Nemotron-70B-Select

Captured: Unknown. Processed: 2026-09-07T19:35:27.512859+00:00.

Model Overview Description: Llama-3.3-Nemotron-70B-Select is a large language model that leverages Meta-Llama-3.3-70B-Instruct as the foundation and is fine-tuned using scaled Bradley-Terry modeling to select the most helpful LLM generated response to user queries. This model is ready for commercial use. License/Terms of Use: GOVERNING TERMS: Use of this model is governed by the NVIDIA Open Model License . Additional Information: Llama 3.3 Community License Agreement . Built with Llama. Arena Hard LeaderBoard As of 18 Mar 2025, augmenting models with the Feedback-Edit Inference Time Scaling (ITS) approach leads to the highest perfor…

F001F002F003F004F005F006F007F008F009F010F011F013F014F015F017F018F019F020F021F023F024