Skip to content

EthenEthenEthen

Mistral · Flagships Analysis

Mistral Small 4

Intelligence, Performance & Price Analysis

Canonical slug: mistral-small-4 · Canonical model registry at build time

Intelligence

20 score

Speed

177.4 output tokens/sec

Latency

0.78s TTFT

Input Price

$0.15 / 1M tokens

Output Price

$0.60 / 1M tokens

Verbosity

53M Output tokens from Intelligence Index 2 out of 4 units for Verbosity . Compa

Executive Assessment

Routing Verdict & Tradeoffs

Ethen Routing Verdict

Strong general-purpose model

Mistral Small 4 (Reasoning) scores 20 on the Artificial Analysis Intelligence Index, placing it well above average among other open weight models of similar size (median: 9). Mistral Small 4 (Reasoning) generates output at 177.4 tokens per second (based on Mistral's API), which is well above average compared to other open weight models of similar size (median: 88.2 t/s). Mistral Small 4 (Reasoning) costs $0.15 per 1M input tokens (very competitive, median: $0.40) and $0.60 per 1M output tokens (better than average, median: $0.84), based on Mistral's API.

Suitable for most production tasks, but high-volume or repetitive work should still be compared against cheaper routes.

Task Fit Assessment

WorkloadRatingNotes
Complex reasoning & agentic workflowsoptimalIntelligence score 20 supports capable reasoning, but very hard tasks may benefit from higher-tier models.
High-volume chat & customer-facingoptimalOutput speed 177.4 tokens/sec and capable intelligence make this suitable for real-time chat at scale.
Latency-sensitive applicationsoptimalTTFT 0.78s — among the lowest latencies, suitable for interactive latency-critical use cases.
Cost-sensitive pipelinesoptimalOutput pricing at $0.60 is reasonable for moderate volume.

Cost Pressure Analysis

Low

Pricing is competitive — input $0.15, output $0.60. Suitable for sustained production use.

Technical Specifications

Architecture and Limits

SpecificationValueValidation Authority
Model typeOpen weightsinferred
ReasoningYesfaq
Input modalitiesMistral Small 4 (Reasoning) supports text and image input.faq
Output modalitiesMistral Small 4 (Reasoning) supports text output.faq
Context window260k tokensfaq
Open weights / sourceYes, Mistral Small 4 (Reasoning) is open weights. The model weights are publicly available and can be downloaded for self-hosting.faq
ParametersMistral Small 4 (Reasoning) has 119 billion parameters (6.5 billion active).faq
Active parametersMistral Small 4 (Reasoning) is a Mixture of Experts (MoE) model with 119 billion total parameters, but only 6.5 billion active parameters are used during inference.faq
LicenseMistral Small 4 (Reasoning) is released under the Apache 2.0 license. This license allows commercial use.faq
API availabilityYes, Mistral Small 4 (Reasoning) is available via API through 1 provider.faq

Evidence charts

Profile visualizations

Charts are the same canonical evidence cards previously published for this slug, contained inside the D18D page grammar.

AA-Omniscience Index

AA-Omniscience Index (higher is better) measures knowledge reliability and hallucination. It rewards correct answers, penalizes hallucinations, and has no penalty for refusing to answer. Scores range from -100 to 100, where 0 means as many correct as incorrect answers, and negative scores mean more incorrect than correct. · Evaluation results measured independently by Artificial Analysis

Retrieved 2026-07-08

Chart source and provenance are listed in Methodology & sources below.

Retrieval date:
2026-07-08
Methodology:
Independent test run by Artificial Analysis on dedicated hardware.

-40-200204060Claude Fable 5 (with fallback): 4040AClaude Fable 5 (with fal…Gemini 3.5 Flash: 2323GGemini 3.5 FlashGPT-5.5 (xhigh): 2020AIGPT-5.5 (xhigh)Grok 4.3 (high): 1818xGrok 4.3 (high)Kimi K2.6: 6.46.4KKimi K2.6Muse Spark: 4.14.1?Muse SparkGLM-5.2 (max): 44?GLM-5.2 (max)MiniMax-M3: 1.41.4?MiniMax-M3Nemotron 3 Ultra: -0.8-0.8NNemotron 3 UltraDeepSeek V4 Pro (max): -10-10DDeepSeek V4 Pro (max)Mistral Small 4: -29.9-29.9MiMistral Small 4

Artificial Analysis Intelligence Index by Open Weights / Proprietary

Artificial Analysis Intelligence Index v4.1 incorporates 9 evaluations: GDPval-AA v2, 𝜏³-Banking, Terminal-Bench v2.1, SciCode, Humanity's Last Exam, GPQA Diamond, CritPt, AA-Omniscience, AA-LCR · Evaluation results measured independently by Artificial Analysis

Retrieved 2026-07-08

Chart source and provenance are listed in Methodology & sources below.

Retrieval date:
2026-07-08
Methodology:
Independent test run by Artificial Analysis on dedicated hardware.

0204060Claude Fable 5 (with fallback): 6060AClaude Fable 5 (with fal…GPT-5.5 (xhigh): 5555AIGPT-5.5 (xhigh)GLM-5.2 (max): 5151?GLM-5.2 (max)Gemini 3.5 Flash: 5050GGemini 3.5 FlashMiniMax-M3: 4444?MiniMax-M3DeepSeek V4 Pro (max): 4444DDeepSeek V4 Pro (max)Kimi K2.6: 4444KKimi K2.6Muse Spark: 4343?Muse SparkNemotron 3 Ultra: 3838NNemotron 3 UltraGrok 4.3 (high): 3838xGrok 4.3 (high)Mistral Small 4: 2020MiMistral Small 4

Artificial Analysis Intelligence Index

Artificial Analysis Intelligence Index v4.1 incorporates 9 evaluations: GDPval-AA v2, 𝜏³-Banking, Terminal-Bench v2.1, SciCode, Humanity's Last Exam, GPQA Diamond, CritPt, AA-Omniscience, AA-LCR · Evaluation results measured independently by Artificial Analysis

Retrieved 2026-07-08

Chart source and provenance are listed in Methodology & sources below.

Retrieval date:
2026-07-08
Methodology:
Independent test run by Artificial Analysis on dedicated hardware.

0204060Claude Fable 5 (with fallback): 6060AClaude Fable 5 (with fal…GPT-5.5 (xhigh): 5555AIGPT-5.5 (xhigh)GLM-5.2 (max): 5151?GLM-5.2 (max)Gemini 3.5 Flash: 5050GGemini 3.5 FlashMiniMax-M3: 4444?MiniMax-M3DeepSeek V4 Pro (max): 4444DDeepSeek V4 Pro (max)Kimi K2.6: 4444KKimi K2.6Muse Spark: 4343?Muse SparkNemotron 3 Ultra: 3838NNemotron 3 UltraGrok 4.3 (high): 3838xGrok 4.3 (high)Mistral Small 4: 2020MiMistral Small 4

Artificial Analysis Openness Index: Score

Openness Index assesses model openness on a 0 to 100 normalized scale (higher is more open) · Evaluation results measured independently by Artificial Analysis

Retrieved 2026-07-08

Chart source and provenance are listed in Methodology & sources below.

Retrieval date:
2026-07-08
Methodology:
Independent test run by Artificial Analysis on dedicated hardware.

050100K2 Think V2: 8989?K2 Think V2Nemotron 3 Ultra: 8383NNemotron 3 UltraNVIDIA Nemotron 3 Super: 8383NNVIDIA Nemotron 3 SuperDeepSeek V4 Pro (max): 5050DDeepSeek V4 Pro (max)GLM-5.2 (max): 4444?GLM-5.2 (max)Qwen3 Next 80B A3B: 4444QQwen3 Next 80B A3BQwen3 Coder Next: 4242QQwen3 Coder NextMistral Small 4: 3939MiMistral Small 4gpt-oss-120b (high): 3939AIgpt-oss-120b (high)Qwen3.5 122B A10B: 3939QQwen3.5 122B A10B

Intelligence

Artificial Analysis Intelligence Index · Higher is better · Evaluation results measured independently by Artificial Analysis

Retrieved 2026-07-08

Chart source and provenance are listed in Methodology & sources below.

Retrieval date:
2026-07-08
Methodology:
Independent test run by Artificial Analysis on dedicated hardware.

0204060Claude Fable 5 (with fallback): 6060AClaude Fable 5 (with fal…GPT-5.5 (xhigh): 5555AIGPT-5.5 (xhigh)GLM-5.2 (max): 5151?GLM-5.2 (max)Gemini 3.5 Flash: 5050GGemini 3.5 FlashMiniMax-M3: 4444?MiniMax-M3DeepSeek V4 Pro (max): 4444DDeepSeek V4 Pro (max)Kimi K2.6: 4444KKimi K2.6Muse Spark: 4343?Muse SparkNemotron 3 Ultra: 3838NNemotron 3 UltraGrok 4.3 (high): 3838xGrok 4.3 (high)Mistral Small 4: 2020MiMistral Small 4

Output Speed

Output tokens per second · Higher is better · Evaluation results measured independently by Artificial Analysis

Retrieved 2026-07-08

Chart source and provenance are listed in Methodology & sources below.

Retrieval date:
2026-07-08
Methodology:
Independent test run by Artificial Analysis on dedicated hardware.

0100200300400NVIDIA Nemotron 3 Super: 364364NNVIDIA Nemotron 3 SuperHyperNova 60B 2605: 353353?HyperNova 60B 2605gpt-oss-120b (high): 314314AIgpt-oss-120b (high)Nemotron 3 Ultra: 250250NNemotron 3 UltraGLM-5.2 (max): 218218?GLM-5.2 (max)Gemini 3.5 Flash: 192192GGemini 3.5 FlashMistral Small 4: 177177MiMistral Small 4Ling 2.6 Flash: 177177?Ling 2.6 FlashGrok 4.3 (high): 164164xGrok 4.3 (high)Mistral Medium 3.5: 140140MiMistral Medium 3.5

Speed

Output tokens per second · Higher is better · Evaluation results measured independently by Artificial Analysis

Retrieved 2026-07-08

Chart source and provenance are listed in Methodology & sources below.

Retrieval date:
2026-07-08
Methodology:
Independent test run by Artificial Analysis on dedicated hardware.

088175263350gpt-oss-120b (high): 314314AIgpt-oss-120b (high)Nemotron 3 Ultra: 250250NNemotron 3 UltraGLM-5.2 (max): 218218?GLM-5.2 (max)Gemini 3.5 Flash: 192192GGemini 3.5 FlashMistral Small 4: 177177MiMistral Small 4Grok 4.3 (high): 164164xGrok 4.3 (high)MiniMax-M3: 9999?MiniMax-M3GPT-5.5 (xhigh): 8888AIGPT-5.5 (xhigh)Kimi K2.6: 7676KKimi K2.6DeepSeek V4 Pro (max): 7272DDeepSeek V4 Pro (max)

End-to-End Response Time

Seconds to output 500 tokens, including reasoning model 'thinking' time · Lower is better · Evaluation results measured independently by Artificial Analysis

Retrieved 2026-07-08

Chart source and provenance are listed in Methodology & sources below.

Retrieval date:
2026-07-08
Methodology:
Independent test run by Artificial Analysis on dedicated hardware.

0s2.5s5s7.5s10sDeepSeek V4 Pro (max): 7s7sDDeepSeek V4 Pro (max)Kimi K2.6: 6.6s6.6sKKimi K2.6Devstral 2: 6.6s6.6s?Devstral 2Llama 3.3 70B: 5.8s5.8sMLlama 3.3 70BGPT-5.5 (xhigh): 5.7s5.7sAIGPT-5.5 (xhigh)Qwen3 Coder Next: 5.6s5.6sQQwen3 Coder NextMiniMax-M3: 5s5s?MiniMax-M3Llama 4 Scout: 4.6s4.6sMLlama 4 ScoutQwen3 Next 80B A3B: 4.4s4.4sQQwen3 Next 80B A3BQwen3.5 122B A10B: 3.6s3.6sQQwen3.5 122B A10BMistral Small 4: 2.8s2.8sMiMistral Small 4

Latency: Time To First Answer Token

Seconds to first answer token received · Accounts for reasoning model 'thinking' time · Evaluation results measured independently by Artificial Analysis

Retrieved 2026-07-08

Chart source and provenance are listed in Methodology & sources below.

Retrieval date:
2026-07-08
Methodology:
Independent test run by Artificial Analysis on dedicated hardware.

0s25s50s75s100sDeepSeek V4 Pro (max): 61s61sDDeepSeek V4 Pro (max)Kimi K2.6: 58s58sKKimi K2.6MiniMax-M3: 20s20s?MiniMax-M3Qwen3 Next 80B A3B: 17s17sQQwen3 Next 80B A3BQwen3.5 122B A10B: 14s14sQQwen3.5 122B A10BMistral Medium 3.5: 14s14sMiMistral Medium 3.5Mistral Small 4: 11s11sMiMistral Small 4GLM-5.2 (max): 9.2s9.2s?GLM-5.2 (max)Nemotron 3 Ultra: 9.1s9.1sNNemotron 3 Ultragpt-oss-120b (high): 6.4s6.4sAIgpt-oss-120b (high)

Context Window

Context window: tokens limit · Higher is better · Evaluation results measured independently by Artificial Analysis

Retrieved 2026-07-08

Chart source and provenance are listed in Methodology & sources below.

Retrieval date:
2026-07-08
Methodology:
Independent test run by Artificial Analysis on dedicated hardware.

02.5M5M7.5M10MLlama 4 Scout: 10M10MMLlama 4 ScoutGemini 3.5 Flash: 1M1MGGemini 3.5 FlashClaude Fable 5 (with fallback): 1M1MAClaude Fable 5 (with fal…GLM-5.2 (max): 1M1M?GLM-5.2 (max)DeepSeek V4 Pro (max): 1M1MDDeepSeek V4 Pro (max)Grok 4.3 (high): 1M1MxGrok 4.3 (high)MiniMax-M3: 1M1M?MiniMax-M3NVIDIA Nemotron 3 Super: 1M1MNNVIDIA Nemotron 3 SuperGPT-5.5 (xhigh): 922k922kAIGPT-5.5 (xhigh)Nemotron 3 Ultra: 262k262kNNemotron 3 UltraMistral Small 4: 256k256kMiMistral Small 4

Cost per Intelligence Index Task

Weighted average cost (USD) per Artificial Analysis Intelligence Index task, segmented by token type. Lower is better · Evaluation results measured independently by Artificial Analysis

Retrieved 2026-07-08

Chart source and provenance are listed in Methodology & sources below.

Retrieval date:
2026-07-08
Methodology:
Independent test run by Artificial Analysis on dedicated hardware.

00.51Claude Fable 5 (with fallback): 0.40.4AClaude Fable 5 (with fal…GPT-5.5 (xhigh): 0.10.1AIGPT-5.5 (xhigh)Gemini 3.5 Flash: 0.10.1GGemini 3.5 FlashMistral Medium 3.5: 0.10.1MiMistral Medium 3.5GLM-5.2 (max): 00?GLM-5.2 (max)Kimi K2.6: 00KKimi K2.6Nemotron 3 Ultra: 00NNemotron 3 UltraQwen3.5 122B A10B: 00QQwen3.5 122B A10BQwen3 Coder Next: 00QQwen3 Coder NextMiniMax-M3: 00?MiniMax-M3

Cost per Task

Weighted average cost (USD) per Intelligence Index task · Lower is better · Evaluation results measured independently by Artificial Analysis

Retrieved 2026-07-08

Chart source and provenance are listed in Methodology & sources below.

Retrieval date:
2026-07-08
Methodology:
Independent test run by Artificial Analysis on dedicated hardware.

$0.00$1.00$2.00$3.00$4.00$5.00Claude Fable 5 (with fallback): $2.75$2.75AClaude Fable 5 (with fal…GPT-5.5 (xhigh): $0.86$0.86AIGPT-5.5 (xhigh)Gemini 3.5 Flash: $0.59$0.59GGemini 3.5 FlashGLM-5.2 (max): $0.37$0.37?GLM-5.2 (max)Kimi K2.6: $0.35$0.35KKimi K2.6Nemotron 3 Ultra: $0.24$0.24NNemotron 3 UltraGrok 4.3 (high): $0.14$0.14xGrok 4.3 (high)MiniMax-M3: $0.12$0.12?MiniMax-M3gpt-oss-120b (high): $0.06$0.06AIgpt-oss-120b (high)DeepSeek V4 Pro (max): $0.04$0.04DDeepSeek V4 Pro (max)

Cost to Run Artificial Analysis Intelligence Index

Cost (USD) to run all evaluations in the Artificial Analysis Intelligence Index · Evaluation results measured independently by Artificial Analysis

Retrieved 2026-07-08

Chart source and provenance are listed in Methodology & sources below.

Retrieval date:
2026-07-08
Methodology:
Independent test run by Artificial Analysis on dedicated hardware.

0200400600Claude Fable 5 (with fallback): 508508AClaude Fable 5 (with fal…GPT-5.5 (xhigh): 132132AIGPT-5.5 (xhigh)Gemini 3.5 Flash: 114114GGemini 3.5 FlashMistral Medium 3.5: 5454MiMistral Medium 3.5Qwen3 Coder Next: 3939QQwen3 Coder NextNemotron 3 Ultra: 3434NNemotron 3 UltraGLM-5.2 (max): 3333?GLM-5.2 (max)Qwen3.5 122B A10B: 3333QQwen3.5 122B A10BKimi K2.6: 2525KKimi K2.6Qwen3 Next 80B A3B: 1616QQwen3 Next 80B A3B

Pricing: Cache Hit, Input, and Output

Price (USD per M Tokens) · Evaluation results measured independently by Artificial Analysis

Retrieved 2026-07-08

Chart source and provenance are listed in Methodology & sources below.

Retrieval date:
2026-07-08
Methodology:
Independent test run by Artificial Analysis on dedicated hardware.

0250k500k750k1MClaude Fable 5 (with fallback): 6161AClaude Fable 5 (with fal…GPT-5.5 (xhigh): 3636AIGPT-5.5 (xhigh)Gemini 3.5 Flash: 1111GGemini 3.5 FlashMistral Medium 3.5: 9.29.2MiMistral Medium 3.5Qwen3 Next 80B A3B: 6.56.5QQwen3 Next 80B A3BGLM-5.2 (max): 6.16.1?GLM-5.2 (max)Kimi K2.6: 5.15.1KKimi K2.6Grok 4.3 (high): 44xGrok 4.3 (high)Qwen3.5 122B A10B: 3.63.6QQwen3.5 122B A10BNemotron 3 Ultra: 3.63.6NNemotron 3 UltraMistral Small 4: 0.750.75MiMistral Small 4cache Hitinputoutput

Time per Intelligence Index Task

Weighted average decode time (minutes) per task; excludes TTFT and overhead time · Lower is better · Evaluation results measured independently by Artificial Analysis

Retrieved 2026-07-08

Chart source and provenance are listed in Methodology & sources below.

Retrieval date:
2026-07-08
Methodology:
Independent test run by Artificial Analysis on dedicated hardware.

02468DeepSeek V4 Pro (max): 6.86.8DDeepSeek V4 Pro (max)Claude Fable 5 (with fallback): 4.84.8AClaude Fable 5 (with fal…MiniMax-M3: 3.93.9?MiniMax-M3GPT-5.5 (xhigh): 33AIGPT-5.5 (xhigh)Mistral Medium 3.5: 2.92.9MiMistral Medium 3.5GLM-5.2 (max): 2.82.8?GLM-5.2 (max)Qwen3 Next 80B A3B: 2.82.8QQwen3 Next 80B A3BQwen3 Coder Next: 2.62.6QQwen3 Coder NextGemini 3.5 Flash: 2.22.2GGemini 3.5 Flashgpt-oss-120b (high): 1.91.9AIgpt-oss-120b (high)Mistral Small 4: 1.71.7MiMistral Small 4

Model Size: Total and Active Parameters

Comparison between total model parameters and parameters active during inference · Evaluation results measured independently by Artificial Analysis

Retrieved 2026-07-08

Chart source and provenance are listed in Methodology & sources below.

Retrieval date:
2026-07-08
Methodology:
Independent test run by Artificial Analysis on dedicated hardware.

0250k500k750k1MDeepSeek V4 Pro (max): 1.6k1.6kDDeepSeek V4 Pro (max)Kimi K2.6: 1k1kKKimi K2.6GLM-5.2 (max): 753753?GLM-5.2 (max)Nemotron 3 Ultra: 550550NNemotron 3 UltraMiniMax-M3: 428428?MiniMax-M3Mistral Medium 3.5: 128128MiMistral Medium 3.5Qwen3.5 122B A10B: 125125QQwen3.5 122B A10BDevstral 2: 125125?Devstral 2NVIDIA Nemotron 3 Super: 121121NNVIDIA Nemotron 3 SuperMistral Small 4: 119119MiMistral Small 4passive Paramsactive Params

Output Tokens per Intelligence Index Task

Weighted average number of output tokens used to run one task in the Artificial Analysis Intelligence Index · Evaluation results measured independently by Artificial Analysis

Retrieved 2026-07-08

Chart source and provenance are listed in Methodology & sources below.

Retrieval date:
2026-07-08
Methodology:
Independent test run by Artificial Analysis on dedicated hardware.

05000100001500020000Qwen3 Coder Next: 1535515355QQwen3 Coder NextLing 2.6 Flash: 1307013070?Ling 2.6 FlashMiniMax-M3: 1162311623?MiniMax-M3Gemini 3.5 Flash: 98199819GGemini 3.5 FlashNemotron 3 Ultra: 77977797NNemotron 3 UltraClaude Fable 5 (with fallback): 76967696AClaude Fable 5 (with fal…gpt-oss-120b (high): 75717571AIgpt-oss-120b (high)NVIDIA Nemotron 3 Super: 74427442NNVIDIA Nemotron 3 SuperDeepSeek V4 Pro (max): 72437243DDeepSeek V4 Pro (max)Mistral Medium 3.5: 70497049MiMistral Medium 3.5Mistral Small 4: 63646364MiMistral Small 4

Methodology

Methodology & Provenance

This page is rendered from the normalized profile and page JSON for Mistral Small 4.

Benchmark values are preserved as normalized; only layout, disclosure ordering, and typography are adjusted for readability.

Frequently Asked Questions

Model FAQs & Technical Disclosures

Mistral Small 4 (Reasoning) was released on March 16, 2026.