Japanese-strengthened Llama 3.1
According to the model card, Swallow enhanced Japanese capability while retaining English capability.
Open Source Model Profile · tokyotech-llm
Llama-3.1-Swallow-8B-Instruct-v0.3 is an 8.03B-parameter Llama instruction model from tokyotech-llm. Its model card describes continual pre-training of Llama 3.1 for Japanese with English retained, plus Japanese-focused supervised fine-tuning.
Llama-3.1-Swallow-8B-Instruct-v0.3 is published by tokyotech-llm as a text-generation model. The captured configuration identifies LlamaForCausalLM with model type llama, and Safetensors metadata reports 8,030,261,248 parameters. According to the model card, it continues from meta-llama/Llama-3.1-8B-Instruct and was then instruction-tuned on the publisher's instruction datasets.
According to the model card, Swallow enhanced Japanese capability while retaining English capability.
The publisher describes approximately 200B tokens from Swallow Corpus Version 2, Japanese and English Wikipedia, and math and coding material.
According to the model card, the Instruct models used supervised fine-tuning on synthetic data built for Japanese.
Source: tokyotech-llm/Llama-3.1-Swallow-8B-Instruct-v0.3
Captured: Unknown. Processed: 2026-09-07T19:34:59.914443+00:00.
Llama 3.1 Swallow - Built with Llama Llama 3.1 Swallow is a series of large language models (8B, 70B) that were built by continual pre-training on the Meta Llama 3.1 models. Llama 3.1 Swallow enhanced the Japanese language capabilities of the original Llama 3.1 while retaining the English language capabilities. We use approximately 200 billion tokens that were sampled from a large Japanese web corpus (Swallow Corpus Version 2), Japanese and English Wikipedia articles, and mathematical and coding contents, etc (see the Training Datasets section of the base model) for continual pre-training. The instruction-tuned models (Instruct) wer…
F001F002F003F004F005F006F007F008F009F010F011F012F013F014F015F016F018F022F023F024