Meta-Llama-3-8B-Instruct-abliterated-v3 is an 8.03B-parameter Llama text-generation model from Greytechai. Its card documents refusal-direction orthogonalization of Meta-Llama-3-8B-Instruct.
Publisher
Greytechai
Task
text-generation
Model type
llama
License
llama3
Library
transformers
Publication status
Accepted · not indexed
Model overview
Meta-Llama-3-8B-Instruct-abliterated-v3 is published by Greytechai as a Llama text-generation model. The captured configuration identifies LlamaForCausalLM and Safetensors metadata reports 8,030,261,248 parameters, or about 8.03B. The model card describes it as Meta-Llama-3-8B-Instruct with weights edited to inhibit refusal, otherwise tuned like the original instruct model.
Recorded capabilities
Refusal-direction orthogonalization
The model card describes orthogonalized bfloat16 weights based on refusal-direction research, intended to inhibit refusal expression while leaving other behavior as close to the original model as the author could manage.
Weight-level behavior edit
According to the model card, the method contrasts a system prompt against a blank prompt and orthogonalizes the desired behavior into the weights, preserving other knowledge and training rather than broadly fine-tuning.
8.03B Llama Transformers record
Captured configuration records LlamaForCausalLM with model type llama, about 8.03B parameters, and Transformers support.
Use cases in the source record
Conversational text-generation experiments where reduced refusal behavior is the explicit research interest, consistent with the record's conversational tag.
Methodology replication and targeted behavior-edit experiments using the publisher's described ablation approach, rather than general broad capability changes.
Limitations and unknowns
No evaluation results were extracted from this record.
No context-window value was extracted from this record.
Provider state is historical snapshot data and should be refreshed before being presented as current.
The publisher does not guarantee non-refusal, comprehension, or absence of ethics and safety commentary; behavioral change is explicitly limited and uncertain.
Llama-3-8B-Instruct-abliterated-v3 Model Card My Jupyter "cookbook" to replicate the methodology can be found here, refined library coming soon This is meta-llama/Meta-Llama-3-8B-Instruct with orthogonalized bfloat16 safetensor weights, generated with a refined methodology based on that which was described in the preview paper/blog post: ' Refusal in LLMs is mediated by a single direction ' which I encourage you to read to understand more. Hang on, "abliteration"? Orthogonalization? Ablation? What is this? TL;DR: This model has had certain weights manipulated to "inhibit" the model's ability to express refusal. It is not in anyway g…