Skip to content

EthenEthenEthen

Open Source Model Profile · facebook

KernelLLM

KernelLLM is an 8.03B-parameter Llama fine-tune from facebook for translating PyTorch modules into Triton kernels. Its model card describes supervised tuning on paired PyTorch and Triton examples with KernelBench-Triton evaluation.

Publisher
facebook
Task
text-generation
Model type
llama
License
other
Library
transformers
Publication status
Accepted · not indexed

Model overview

KernelLLM is published by facebook as a text-generation model. The captured configuration identifies LlamaForCausalLM with model type llama, and Safetensors metadata reports 8,030,261,248 parameters. According to the model card, it is based on Llama 3.1 Instruct and trained specifically for Triton kernel authoring.

Recorded capabilities

PyTorch-to-Triton focus

According to the model card, KernelLLM translates PyTorch modules into Triton kernels.

Paired kernel dataset

The publisher describes about 25,000 paired PyTorch and Triton examples plus synthetic samples, with the filtered set released as KernelBook.

Documented SFT recipe

According to the model card, Llama-3.1-8B-Instruct was fine-tuned for 10 epochs at batch size 32, taking about 12 wall-clock hours on 16 GPUs.

Use cases in the source record

  • Drafting Triton kernels from PyTorch modules using the publisher's documented prompt template and calling code.
  • Kernel-generation experiments measured against the publisher's KernelBench-Triton variant.

Limitations and unknowns

  • According to the model card, the model may produce incorrect API references and syntax errors and has limited instruction-following ability.
  • KernelBench figures in the card are publisher-reported values, not Ethen-measured results.
  • No context-window value was extracted from this record.
  • Provider state is historical snapshot data, not independently refreshed current availability.

Source and provenance

Source: facebook/KernelLLM

Captured: Unknown. Processed: 2026-09-07T19:35:21.616101+00:00.

KernelLLM On KernelBench-Triton Level 1, our 8B parameter model exceeds models such as GPT-4o and DeepSeek V3 in single-shot performance. With multiple inferences, KernelLLM's performance outperforms DeepSeek R1. This is all from a model with two orders of magnitude fewer parameters than its competitors. Updates : 2025/06/25: We added an end-to-end example walkthrough , where we format a community-provided prompt for KernelLLM to function well. We have received many questions about how to format the prompts such that KernelLLM performs best. We hope this can help! 2025/06/15 We would like to thank the community for the creation of m…

F001F002F003F004F005F006F007F008F009F010F016F019F020F021F022F023F024F030