Skip to content

EthenEthenEthen

Open Source Model Profile · tokyotech-llm

Llama-3.1-Swallow-8B-Instruct-v0.5

Llama-3.1-Swallow-8B-Instruct-v0.5 is an 8.03B-parameter Llama text-generation fine-tune from tokyotech-llm. Its model card documents Japanese-focused continual pre-training and instruction tuning.

Publisher
tokyotech-llm
Task
text-generation
Model type
llama
License
Unknown
Library
transformers
Publication status
Accepted · not indexed

Model overview

Llama-3.1-Swallow-8B-Instruct-v0.5 is published by tokyotech-llm as a Llama-based text-generation model. The captured configuration identifies LlamaForCausalLM and Safetensors metadata reports 8,030,261,248 parameters. According to the model card, it was continually pre-trained from meta-llama/Llama-3.1-8B-Instruct and then instruction-tuned on synthetic Japanese-focused data.

Recorded capabilities

Japanese-enhanced Llama 3.1 lineage

According to the model card, Llama 3.1 Swallow was built by continual pre-training on Meta Llama 3.1 to enhance Japanese capability while retaining English capability.

Documented instruction tuning

According to the model card, the Instruct version was built by supervised fine-tuning on synthetic data specially built for Japanese, including data derived from lmsys-chat-1m.

Reported Japanese MT-Bench result

According to the model card, v0.5 exhibits state-of-the-art performance among open-source LLMs with 8B or fewer parameters on Japanese MT-Bench, outperforming v0.3 by 1.5 points.

Use cases in the source record

  • Japanese-English conversational text generation using the instruction-tuned Swallow chat behavior.
  • Japanese multi-turn dialogue experiments evaluated with the card's documented Japanese MT-Bench setup.

Limitations and unknowns

  • No context-window value was extracted from this record.
  • No license verification beyond the captured llama3.3 and gemma tags was performed; the card also cites Meta Llama 3.1 and Gemma terms.
  • Benchmark figures and state-of-the-art wording are publisher claims from the model card and were not independently verified.
  • Provider state is historical snapshot data, not independently refreshed current availability.

Source and provenance

Source: tokyotech-llm/Llama-3.1-Swallow-8B-Instruct-v0.5

Captured: Unknown. Processed: 2026-09-07T19:34:59.940205+00:00.

Llama 3.1 Swallow - Built with Llama Llama 3.1 Swallow is a series of large language models (8B, 70B) that were built by continual pre-training on the Meta Llama 3.1 models. Llama 3.1 Swallow enhanced the Japanese language capabilities of the original Llama 3.1 while retaining the English language capabilities. We use approximately 200 billion tokens that were sampled from a large Japanese web corpus (Swallow Corpus Version 2), Japanese and English Wikipedia articles, and mathematical and coding contents, etc (see the Training Datasets section of the base model) for continual pre-training. The instruction-tuned models (Instruct) wer…

F001F002F003F004F005F006F007F009F010F011F012F014F015F017F019F023F025