Instruct-thinking hybrid tuning
According to the model card, Unsloth tuning for 3 epochs on the TeichAI Claude high-reasoning dataset created an instruct and thinking hybrid without updating core knowledge.
Open Source Model Profile · DavidAU
Llama3.3-8B-Instruct-Thinking-Claude-4.5-Opus-High-Reasoning is an 8.03B-parameter Llama instruct-thinking hybrid from DavidAU. According to the model card, it was tuned with Unsloth for 3 epochs.
Llama3.3-8B-Instruct-Thinking-Claude-4.5-Opus-High-Reasoning is published by DavidAU as a Llama text-generation fine-tune. The captured configuration identifies LlamaForCausalLM with model type llama, and Safetensors metadata reports 8,030,261,248 parameters. According to the model card, Unsloth tuning for 3 epochs produced an instruct and thinking hybrid from allura-forge/Llama-3.3-8B-Instruct, with card data recording apache-2.0.
According to the model card, Unsloth tuning for 3 epochs on the TeichAI Claude high-reasoning dataset created an instruct and thinking hybrid without updating core knowledge.
According to the model card, no system prompt is needed because thinking tags self-generate, with hub tags recording thinking, reasoning, and instruct entries.
According to the model card, the publisher documents a 1.5 smoothing factor for KoboldCpp, text-generation-webui, and Silly Tavern chat or roleplay use.
The captured configuration identifies LlamaForCausalLM with model type llama and about 8.03B Safetensors parameters with Transformers support.
Source: DavidAU/Llama3.3-8B-Instruct-Thinking-Claude-4.5-Opus-High-Reasoning
Captured: Unknown. Processed: 2026-09-07T19:35:39.015150+00:00.
Llama3.3-8B-Instruct-Thinking-Claude-4.5-Opus-High-Reasoning What madness is this? Someone found "Llama3.3-8B" source (never publicly released) in the "wild", then it was adjusted back to 128k and then I added my own special madness: Training the model with Unsloth (3 epochs) and Claude 4.5-Opus High Reasoning dataset. This has created an Instruct/Thinking hybrid (128k context, Llama 3.3 model). Note this tuning was only to create an instruct/thinking model, not to update the model's core knowledge / root training. 1 example at bottom of the page. HERETIC / Uncensored Version: https://huggingface.co/DavidAU/Llama3.3-8B-Instruct-Thin…
F001F002F003F004F005F006F007F008F009F010F012F013F014F015F016F017F019F020F021F022