Korean SFT and DPO tuning
According to the model card, the model was fine-tuned on a curated high-quality dataset with Supervised Fine-Tuning and Direct Preference Optimization for human feedback.
Open Source Model Profile · rtzr
ko-gemma-2-9b-it is a 9.24B-parameter Korean Gemma conversational fine-tune from rtzr. According to the model card, it was trained with Supervised Fine-Tuning and Direct Preference Optimization.
ko-gemma-2-9b-it is published by rtzr as a Gemma text-generation model. The captured configuration identifies Gemma2ForCausalLM with a gemma2 model type, and Safetensors metadata reports 9,241,705,984 parameters. According to the model card, it is a Korean-language conversational fine-tune trained with Supervised Fine-Tuning and Direct Preference Optimization.
According to the model card, the model was fine-tuned on a curated high-quality dataset with Supervised Fine-Tuning and Direct Preference Optimization for human feedback.
According to the model card, input is a text string such as a question, prompt, or document, and output is generated Korean-language text such as an answer or summary.
According to the model card, a comparison table reports this model ahead of google/gemma-2-9b-it and two Korean baselines across math, reasoning, writing, coding, understanding, and grammar groupings.
According to the model card, inference uses the tokenizer chat template with a documented transformers pipeline example generating up to 2048 new tokens.
According to the model card, vLLM 0.5.1 fails to load the Gemma 2 model, so the publisher recommends vllm-openai:latest Docker or vllm 0.5.0.post1 with a supplied run command.
Source: rtzr/ko-gemma-2-9b-it
Captured: Unknown. Processed: 2026-09-07T19:34:57.128436+00:00.
Model Details Ko-Gemma-2-9B-IT Ko-Gemma-2-9B-IT is a Korean-language conversational model that is part of the Gemma family of models. It is a text-to-text, decoder-only large language model, available in Korean. We fine-tuned this model on a carefully curated high-quality dataset using Supervised Fine-Tuning (SFT). And we use Direct Preference Optimization training specifically for Human Feedback. The datasets include: Orca-Math dpo-mix-7k Some of these datasets were partially used and translated for training. In particular, a lot of repetition occurred during the translation process, so preprocessing was performed based on N-gram.…
F001F002F003F004F005F006F007F009F010F011F012F013F016F017F018F019F020F021