Skip to content

EthenEthenEthen

Open Source Model Profile · tokyotech-llm

Llama-3.1-Swallow-8B-Instruct-v0.3

Llama-3.1-Swallow-8B-Instruct-v0.3 is an 8.03B-parameter Llama instruction model from tokyotech-llm. Its model card describes continual pre-training of Llama 3.1 for Japanese with English retained, plus Japanese-focused supervised fine-tuning.

Publisher
tokyotech-llm
Task
text-generation
Model type
llama
License
Unknown
Library
transformers
Publication status
Accepted · not indexed

Model overview

Llama-3.1-Swallow-8B-Instruct-v0.3 is published by tokyotech-llm as a text-generation model. The captured configuration identifies LlamaForCausalLM with model type llama, and Safetensors metadata reports 8,030,261,248 parameters. According to the model card, it continues from meta-llama/Llama-3.1-8B-Instruct and was then instruction-tuned on the publisher's instruction datasets.

Recorded capabilities

Japanese-strengthened Llama 3.1

According to the model card, Swallow enhanced Japanese capability while retaining English capability.

Documented continual pre-training

The publisher describes approximately 200B tokens from Swallow Corpus Version 2, Japanese and English Wikipedia, and math and coding material.

Japanese synthetic instruction tuning

According to the model card, the Instruct models used supervised fine-tuning on synthetic data built for Japanese.

Use cases in the source record

  • Japanese and English conversational text generation using the publisher's instruction-tuned setup.
  • Research on cross-lingual adaptation that references the card's Swallow Corpus and synthetic instruction datasets.

Limitations and unknowns

  • No context-window value was extracted from this record.
  • No Ethen-measured evaluation results were extracted; MT-Bench figures are publisher-reported only.
  • Provider state is historical snapshot data, not independently refreshed current availability.

Source and provenance

Source: tokyotech-llm/Llama-3.1-Swallow-8B-Instruct-v0.3

Captured: Unknown. Processed: 2026-09-07T19:34:59.914443+00:00.

Llama 3.1 Swallow - Built with Llama Llama 3.1 Swallow is a series of large language models (8B, 70B) that were built by continual pre-training on the Meta Llama 3.1 models. Llama 3.1 Swallow enhanced the Japanese language capabilities of the original Llama 3.1 while retaining the English language capabilities. We use approximately 200 billion tokens that were sampled from a large Japanese web corpus (Swallow Corpus Version 2), Japanese and English Wikipedia articles, and mathematical and coding contents, etc (see the Training Datasets section of the base model) for continual pre-training. The instruction-tuned models (Instruct) wer…

F001F002F003F004F005F006F007F008F009F010F011F012F013F014F015F016F018F022F023F024