Skip to content

EthenEthenEthen

Open Source Model Profile · swiss-ai

Apertus-70B-Instruct-2509

Apertus-70B-Instruct-2509 is a 70.6B-parameter multilingual instruction-tuned model from swiss-ai with 15T-token pretraining and tool use.

Publisher
swiss-ai
Task
text-generation
Model type
apertus
License
apache-2.0
Library
transformers
Publication status
Accepted · not indexed

Model overview

Apertus-70B-Instruct-2509 is published by swiss-ai as an Apertus instruction-tuned text-generation model. The captured configuration identifies ApertusForCausalLM and Safetensors metadata reports 70,599,864,480 parameters. According to the model card, it is a multilingual decoder-only model pretrained on 15T tokens, and card data records apache-2.0.

Recorded capabilities

Fully open 70B model

According to the model card, Apertus pairs open weights with open data and full training recipes, including data-reconstruction scripts and intermediate checkpoints.

Massively multilingual coverage

The model card describes support for more than 1,000 languages, with 1,811 natively supported languages and consent-respecting compliant data.

15T-token decoder training

According to the model card, the decoder-only Transformer pretrained on 15T tokens in bfloat16 with xIELU activation, AdEMAMix, Megatron-LM, and 4,096 GH200 GPUs.

Tool use and broad runtimes

According to the model card, Apertus supports tool use and deployment through Transformers, vLLM, SGLang, and on-device MLX.

Use cases in the source record

  • Multilingual text-generation work across the broadly covered languages the publisher documents.
  • Instruction-following and agentic tool-use experiments using the documented chat template with temperature 0.8 and top-p 0.9 sampling.

Limitations and unknowns

  • No context-window value was extracted from this record.
  • Provider state is historical snapshot data and should be refreshed before being presented as current.
  • Benchmark figures are publisher-reported model-card values and were not independently verified by Ethen.
  • According to the model card, generated content may be inaccurate, inconsistent, or biased and should be verified rather than treated as definitive.

Source and provenance

Source: swiss-ai/Apertus-70B-Instruct-2509

Captured: Unknown. Processed: 2026-09-07T19:36:02.517665+00:00.

Apertus Table of Contents Model Summary How to use Evaluation Training Limitations Legal Aspects Model Summary Apertus is a 70B and 8B parameter language model designed to push the boundaries of fully-open multilingual and transparent models. The model supports over 1000 languages and long context, it uses only fully compliant and open training data, and achieves comparable performance to models trained behind closed doors. The model is a decoder-only transformer, pretrained on 15T tokens with a staged curriculum of web, code and math data. The model uses a new xIELU activation function and is trained from scratch with the AdEMAMix…

F001F002F003F004F005F006F007F009F010F012F013F014F015F018F019F020F021F023