Llama chat architecture
The captured configuration identifies LlamaForCausalLM with model type llama and Transformers support.
Open Source Model Profile · RedHatAI
Meta-Llama-3-8B-Instruct-FP8 is an 8.03B-parameter Llama text-generation model from RedHatAI. According to the model card, it is an FP8-quantized chat release for vLLM inference.
Meta-Llama-3-8B-Instruct-FP8 is published by RedHatAI as a text-generation model. The captured configuration identifies LlamaForCausalLM, and Safetensors metadata reports about 8.03B parameters. According to the model card, it quantizes Meta-Llama-3-8B-Instruct weights and activations to FP8 for assistant-like chat.
The captured configuration identifies LlamaForCausalLM with model type llama and Transformers support.
According to the model card, both weights and activations are quantized to FP8 for vLLM inference.
According to the model card, this is a Neural Magic quantized version of Meta-Llama-3-8B-Instruct.
According to the model card, the FP8 model scored 68.22 on OpenLLM version 1 against 68.71 unquantized.
According to the model card, intended use is English assistant-like chat for commercial and research work, with other languages out of scope.
Source: RedHatAI/Meta-Llama-3-8B-Instruct-FP8
Captured: Unknown. Processed: 2026-09-07T19:34:36.171995+00:00.
Meta-Llama-3-8B-Instruct-FP8 Model Overview Model Architecture: Meta-Llama-3 Input: Text Output: Text Model Optimizations: Weight quantization: FP8 Activation quantization: FP8 Intended Use Cases: Intended for commercial and research use in English. Similarly to Meta-Llama-3-8B-Instruct , this models is intended for assistant-like chat. Out-of-scope: Use in any manner that violates applicable laws or regulations (including trade compliance laws). Use in languages other than English. Release Date: 6/8/2024 Version: 1.0 License(s): Llama3 Model Developers: Neural Magic Quantized version of Meta-Llama-3-8B-Instruct . It achieves an ave…
F001F002F003F004F005F006F007F008F009F010F011F012F013F014F015F016