Skip to content

EthenEthenEthen

Open Source Model Profile · RedHatAI

Meta-Llama-3-8B-Instruct-FP8

Meta-Llama-3-8B-Instruct-FP8 is an 8.03B-parameter Llama text-generation model from RedHatAI. According to the model card, it is an FP8-quantized chat release for vLLM inference.

Publisher
RedHatAI
Task
text-generation
Model type
llama
License
llama3
Library
transformers
Publication status
Approved for indexing

Model overview

Meta-Llama-3-8B-Instruct-FP8 is published by RedHatAI as a text-generation model. The captured configuration identifies LlamaForCausalLM, and Safetensors metadata reports about 8.03B parameters. According to the model card, it quantizes Meta-Llama-3-8B-Instruct weights and activations to FP8 for assistant-like chat.

Recorded capabilities

Llama chat architecture

The captured configuration identifies LlamaForCausalLM with model type llama and Transformers support.

FP8 weights and activations

According to the model card, both weights and activations are quantized to FP8 for vLLM inference.

Documented parent lineage

According to the model card, this is a Neural Magic quantized version of Meta-Llama-3-8B-Instruct.

Reported OpenLLM score

According to the model card, the FP8 model scored 68.22 on OpenLLM version 1 against 68.71 unquantized.

English assistant scope

According to the model card, intended use is English assistant-like chat for commercial and research work, with other languages out of scope.

Use cases in the source record

  • English assistant-like chat workflows consistent with the documented commercial and research intent.
  • Memory-efficient vLLM inference experiments using the documented FP8 weight and activation format.

Limitations and unknowns

  • No context-window value was extracted from this record.
  • The only evaluation figure is the publisher-reported OpenLLM average; no other benchmark detail was extracted.
  • Provider state is historical snapshot data and should be refreshed before being presented as current.

Source and provenance

Source: RedHatAI/Meta-Llama-3-8B-Instruct-FP8

Captured: Unknown. Processed: 2026-09-07T19:34:36.171995+00:00.

Meta-Llama-3-8B-Instruct-FP8 Model Overview Model Architecture: Meta-Llama-3 Input: Text Output: Text Model Optimizations: Weight quantization: FP8 Activation quantization: FP8 Intended Use Cases: Intended for commercial and research use in English. Similarly to Meta-Llama-3-8B-Instruct , this models is intended for assistant-like chat. Out-of-scope: Use in any manner that violates applicable laws or regulations (including trade compliance laws). Use in languages other than English. Release Date: 6/8/2024 Version: 1.0 License(s): Llama3 Model Developers: Neural Magic Quantized version of Meta-Llama-3-8B-Instruct . It achieves an ave…

F001F002F003F004F005F006F007F008F009F010F011F012F013F014F015F016