Humanish conversational tuning
According to the model card, DPO tuning is described as reducing assistant-like or overly neutral responses in favor of humanlike reactions.
Open Source Model Profile · vicgalle
Humanish-Roleplay-Llama-3.1-8B is an 8.03B-parameter Llama text-generation fine-tune from vicgalle. Its model card documents DPO tuning for humanlike and role-play behavior.
Humanish-Roleplay-Llama-3.1-8B is published by vicgalle as a text-generation model. The captured configuration identifies LlamaForCausalLM, and Safetensors metadata reports 8,030,261,248 parameters. According to the model card, it is a DPO-tuned Llama-3.1 variant for humanlike conversation and role-play.
According to the model card, DPO tuning is described as reducing assistant-like or overly neutral responses in favor of humanlike reactions.
According to the model card, the release steers output toward RP action formatting and documents chat-template and tokenizer use with an RP example.
According to the model card, the fine-tuning script and library versions are shared, with a reported sub-one-hour T4 runtime.
Source: vicgalle/Humanish-Roleplay-Llama-3.1-8B
Captured: Unknown. Processed: 2026-09-07T19:35:00.777217+00:00.
Humanish-Roleplay-Llama-3.1-8B A DPO-tuned Llama-3.1 to behave more "humanish", i.e., avoiding all the AI assistant slop. It also works for role-play (RP). To achieve this, the model was fine-tuned over a series of datasets: General conversations from Claude Opus, from Undi95/Meta-Llama-3.1-8B-Claude Undi95/Weyaxi-humanish-dpo-project-noemoji , to make the model react as a human, rejecting assistant-like or too neutral responses. ResplendentAI/NSFW_RP_Format_DPO , to steer the model towards using the *action* format in RP settings. Works best if in the first message you also use this format naturally (see example) Usage example conv…
F001F002F003F004F005F006F007F009F010F011F012F013F014F015