Japanese-enhanced Llama 3.1 lineage
According to the model card, Llama 3.1 Swallow was built by continual pre-training on Meta Llama 3.1 to enhance Japanese capability while retaining English capability.
Open Source Model Profile · tokyotech-llm
Llama-3.1-Swallow-8B-Instruct-v0.5 is an 8.03B-parameter Llama text-generation fine-tune from tokyotech-llm. Its model card documents Japanese-focused continual pre-training and instruction tuning.
Llama-3.1-Swallow-8B-Instruct-v0.5 is published by tokyotech-llm as a Llama-based text-generation model. The captured configuration identifies LlamaForCausalLM and Safetensors metadata reports 8,030,261,248 parameters. According to the model card, it was continually pre-trained from meta-llama/Llama-3.1-8B-Instruct and then instruction-tuned on synthetic Japanese-focused data.
According to the model card, Llama 3.1 Swallow was built by continual pre-training on Meta Llama 3.1 to enhance Japanese capability while retaining English capability.
According to the model card, the Instruct version was built by supervised fine-tuning on synthetic data specially built for Japanese, including data derived from lmsys-chat-1m.
According to the model card, v0.5 exhibits state-of-the-art performance among open-source LLMs with 8B or fewer parameters on Japanese MT-Bench, outperforming v0.3 by 1.5 points.
Source: tokyotech-llm/Llama-3.1-Swallow-8B-Instruct-v0.5
Captured: Unknown. Processed: 2026-09-07T19:34:59.940205+00:00.
Llama 3.1 Swallow - Built with Llama Llama 3.1 Swallow is a series of large language models (8B, 70B) that were built by continual pre-training on the Meta Llama 3.1 models. Llama 3.1 Swallow enhanced the Japanese language capabilities of the original Llama 3.1 while retaining the English language capabilities. We use approximately 200 billion tokens that were sampled from a large Japanese web corpus (Swallow Corpus Version 2), Japanese and English Wikipedia articles, and mathematical and coding contents, etc (see the Training Datasets section of the base model) for continual pre-training. The instruction-tuned models (Instruct) wer…
F001F002F003F004F005F006F007F009F010F011F012F014F015F017F019F023F025