Skip to content

EthenEthenEthen

OpenAI · Flagships Analysis

GPT-4.1 mini

Intelligence, Performance & Price Analysis

Canonical slug: gpt-4-1-mini · Canonical model registry at build time

Intelligence

15 score

Speed

85.6 output tokens/sec

Latency

1.02s TTFT

Input Price

$0.40 / 1M tokens

Output Price

$1.60 / 1M tokens

Verbosity

5.2M Output tokens from Intelligence Index 2 out of 4 units for Verbosity . Comp

Executive Assessment

Routing Verdict & Tradeoffs

Ethen Routing Verdict

Capable everyday model

GPT-4.1 mini scores 15 on the Artificial Analysis Intelligence Index, placing it above average among other non-reasoning models in a similar price tier (median: 11). GPT-4.1 mini generates output at 85.6 tokens per second (based on OpenAI's API), which is below average compared to other non-reasoning models in a similar price tier (median: 95.5 t/s). GPT-4.1 mini costs $0.40 per 1M input tokens (somewhat higher than average, median: $0.20) and $1.60 per 1M output tokens (at the higher end, median: $0.70), based on OpenAI's API.

Good for routine tasks; route complex reasoning and premium workloads to stronger models.

Task Fit Assessment

WorkloadRatingNotes
Complex reasoning & agentic workflowsviableIntelligence score 15 handles routine reasoning but may struggle with open-ended agentic tasks.
High-volume chat & customer-facingoptimalOutput speed 85.6 tokens/sec is adequate for chat.
Latency-sensitive applicationsoptimalTTFT 1.02s — adequate latency for most interactive use cases.
Cost-sensitive pipelinesoptimalOutput pricing at $1.60 is reasonable for moderate volume.

Cost Pressure Analysis

Medium

Pricing is moderate — input $0.40, output $1.60. Costs accumulate at volume but are manageable for valuable tasks.

Cheaper substitutes: Cheaper routes for predictable extraction, labeling, or summarization

Technical Specifications

Architecture and Limits

SpecificationValueValidation Authority
Model typeProprietaryinferred
ReasoningNofaq
Input modalitiesGPT-4.1 mini supports text and image input.faq
Output modalitiesGPT-4.1 mini supports text output.faq
Context windowGPT-4.1 mini has a context window of 1.0M tokens. This determines how much text and conversation history the model can process in a single request.faq
Open weights / sourceNo, GPT-4.1 mini is proprietary. The model weights are not publicly available.faq
ParametersGPT-4.1 mini is a proprietary model and OpenAI has not disclosed the model size or parameter count.faq
Knowledge cutoffGPT-4.1 mini has a knowledge cutoff of May 2024. The model's training data includes information up to this date.faq
API availabilityYes, GPT-4.1 mini is available via API through 2 providers.faq

Evidence charts

Profile visualizations

Charts are the same canonical evidence cards previously published for this slug, contained inside the D18D page grammar.

AA-Omniscience Index

AA-Omniscience Index (higher is better) measures knowledge reliability and hallucination. It rewards correct answers, penalizes hallucinations, and has no penalty for refusing to answer. Scores range from -100 to 100, where 0 means as many correct as incorrect answers, and negative scores mean more incorrect than correct. · Evaluation results measured independently by Artificial Analysis

Retrieved 2026-07-08

Chart source and provenance are listed in Methodology & sources below.

Retrieval date:
2026-07-08
Methodology:
Independent test run by Artificial Analysis on dedicated hardware.

-100-50050Claude Fable 5 (with fallback): 4040AClaude Fable 5 (with fal…Gemini 3.5 Flash: 2323GGemini 3.5 FlashGPT-5.5 (xhigh): 2020AIGPT-5.5 (xhigh)Grok 4.3 (high): 1818xGrok 4.3 (high)Kimi K2.6: 6.46.4KKimi K2.6Muse Spark: 4.14.1?Muse SparkGLM-5.2 (max): 44?GLM-5.2 (max)MiniMax-M3: 1.41.4?MiniMax-M3Nemotron 3 Ultra: -0.8-0.8NNemotron 3 UltraDeepSeek V4 Pro (max): -10-10DDeepSeek V4 Pro (max)GPT-4.1 mini: -50.1-50.1AIGPT-4.1 mini

Artificial Analysis Intelligence Index by Open Weights / Proprietary

Artificial Analysis Intelligence Index v4.1 incorporates 9 evaluations: GDPval-AA v2, 𝜏³-Banking, Terminal-Bench v2.1, SciCode, Humanity's Last Exam, GPQA Diamond, CritPt, AA-Omniscience, AA-LCR · Evaluation results measured independently by Artificial Analysis

Retrieved 2026-07-08

Chart source and provenance are listed in Methodology & sources below.

Retrieval date:
2026-07-08
Methodology:
Independent test run by Artificial Analysis on dedicated hardware.

0204060Claude Fable 5 (with fallback): 6060AClaude Fable 5 (with fal…GPT-5.5 (xhigh): 5555AIGPT-5.5 (xhigh)GLM-5.2 (max): 5151?GLM-5.2 (max)Gemini 3.5 Flash: 5050GGemini 3.5 FlashMiniMax-M3: 4444?MiniMax-M3DeepSeek V4 Pro (max): 4444DDeepSeek V4 Pro (max)Kimi K2.6: 4444KKimi K2.6Muse Spark: 4343?Muse SparkNemotron 3 Ultra: 3838NNemotron 3 UltraGrok 4.3 (high): 3838xGrok 4.3 (high)GPT-4.1 mini: 1515AIGPT-4.1 mini

Artificial Analysis Intelligence Index

Artificial Analysis Intelligence Index v4.1 incorporates 9 evaluations: GDPval-AA v2, 𝜏³-Banking, Terminal-Bench v2.1, SciCode, Humanity's Last Exam, GPQA Diamond, CritPt, AA-Omniscience, AA-LCR · Evaluation results measured independently by Artificial Analysis

Retrieved 2026-07-08

Chart source and provenance are listed in Methodology & sources below.

Retrieval date:
2026-07-08
Methodology:
Independent test run by Artificial Analysis on dedicated hardware.

0204060Claude Fable 5 (with fallback): 6060AClaude Fable 5 (with fal…GPT-5.5 (xhigh): 5555AIGPT-5.5 (xhigh)GLM-5.2 (max): 5151?GLM-5.2 (max)Gemini 3.5 Flash: 5050GGemini 3.5 FlashMiniMax-M3: 4444?MiniMax-M3DeepSeek V4 Pro (max): 4444DDeepSeek V4 Pro (max)Kimi K2.6: 4444KKimi K2.6Muse Spark: 4343?Muse SparkNemotron 3 Ultra: 3838NNemotron 3 UltraGrok 4.3 (high): 3838xGrok 4.3 (high)GPT-4.1 mini: 1515AIGPT-4.1 mini

Intelligence

Artificial Analysis Intelligence Index · Higher is better · Evaluation results measured independently by Artificial Analysis

Retrieved 2026-07-08

Chart source and provenance are listed in Methodology & sources below.

Retrieval date:
2026-07-08
Methodology:
Independent test run by Artificial Analysis on dedicated hardware.

0204060Claude Fable 5 (with fallback): 6060AClaude Fable 5 (with fal…GPT-5.5 (xhigh): 5555AIGPT-5.5 (xhigh)GLM-5.2 (max): 5151?GLM-5.2 (max)Gemini 3.5 Flash: 5050GGemini 3.5 FlashMiniMax-M3: 4444?MiniMax-M3DeepSeek V4 Pro (max): 4444DDeepSeek V4 Pro (max)Kimi K2.6: 4444KKimi K2.6Muse Spark: 4343?Muse SparkNemotron 3 Ultra: 3838NNemotron 3 UltraGrok 4.3 (high): 3838xGrok 4.3 (high)GPT-4.1 mini: 1515AIGPT-4.1 mini

Output Speed

Output tokens per second · Higher is better · Evaluation results measured independently by Artificial Analysis

Retrieved 2026-07-08

Chart source and provenance are listed in Methodology & sources below.

Retrieval date:
2026-07-08
Methodology:
Independent test run by Artificial Analysis on dedicated hardware.

088175263350gpt-oss-120b (high): 314314AIgpt-oss-120b (high)Nemotron 3 Ultra: 250250NNemotron 3 UltraGLM-5.2 (max): 218218?GLM-5.2 (max)Gemini 3.5 Flash: 192192GGemini 3.5 FlashGPT-4.1 nano: 188188AIGPT-4.1 nanoLing 2.6 Flash: 177177?Ling 2.6 FlashMistral Small 3.1: 171171MiMistral Small 3.1Grok 4.3 (high): 164164xGrok 4.3 (high)Mistral Small 3.2: 147147MiMistral Small 3.2Llama 4 Maverick: 118118MLlama 4 MaverickGPT-4.1 mini: 8686AIGPT-4.1 mini

Speed

Output tokens per second · Higher is better · Evaluation results measured independently by Artificial Analysis

Retrieved 2026-07-08

Chart source and provenance are listed in Methodology & sources below.

Retrieval date:
2026-07-08
Methodology:
Independent test run by Artificial Analysis on dedicated hardware.

088175263350gpt-oss-120b (high): 314314AIgpt-oss-120b (high)Nemotron 3 Ultra: 250250NNemotron 3 UltraGLM-5.2 (max): 218218?GLM-5.2 (max)Gemini 3.5 Flash: 192192GGemini 3.5 FlashGrok 4.3 (high): 164164xGrok 4.3 (high)MiniMax-M3: 9999?MiniMax-M3GPT-5.5 (xhigh): 8888AIGPT-5.5 (xhigh)GPT-4.1 mini: 8686AIGPT-4.1 miniKimi K2.6: 7676KKimi K2.6DeepSeek V4 Pro (max): 7272DDeepSeek V4 Pro (max)

End-to-End Response Time

Seconds to output 500 tokens, including reasoning model 'thinking' time · Lower is better · Evaluation results measured independently by Artificial Analysis

Retrieved 2026-07-08

Chart source and provenance are listed in Methodology & sources below.

Retrieval date:
2026-07-08
Methodology:
Independent test run by Artificial Analysis on dedicated hardware.

0s2.5s5s7.5s10sMistral Large 3: 9.4s9.4sMiMistral Large 3Kimi K2.6: 6.6s6.6sKKimi K2.6Ministral 3 14B: 6.4s6.4s?Ministral 3 14BLlama 3.3 70B: 5.8s5.8sMLlama 3.3 70BGPT-4.1 mini: 5.8s5.8sAIGPT-4.1 miniQwen3 Coder Next: 5.6s5.6sQQwen3 Coder NextMistral Medium 3.1: 5.6s5.6sMiMistral Medium 3.1Ministral 3 8B: 5.4s5.4s?Ministral 3 8BMiniMax-M3: 5s5s?MiniMax-M3Llama 4 Scout: 4.6s4.6sMLlama 4 Scout

Latency: Time To First Answer Token

Seconds to first answer token received · Accounts for reasoning model 'thinking' time · Evaluation results measured independently by Artificial Analysis

Retrieved 2026-07-08

Chart source and provenance are listed in Methodology & sources below.

Retrieval date:
2026-07-08
Methodology:
Independent test run by Artificial Analysis on dedicated hardware.

0s15s30s45s60sKimi K2.6: 58s58sKKimi K2.6MiniMax-M3: 20s20s?MiniMax-M3GLM-5.2 (max): 9.2s9.2s?GLM-5.2 (max)Nemotron 3 Ultra: 9.1s9.1sNNemotron 3 Ultragpt-oss-120b (high): 6.4s6.4sAIgpt-oss-120b (high)GPT-4.1 nano: 0s0sAIGPT-4.1 nanoMistral Small 3.2: 0s0sMiMistral Small 3.2Mistral Small 3.1: 0s0sMiMistral Small 3.1Ministral 3 8B: 0s0s?Ministral 3 8BLlama 4 Scout: 0s0sMLlama 4 ScoutGPT-4.1 mini: 0s0sAIGPT-4.1 mini

Context Window

Context window: tokens limit · Higher is better · Evaluation results measured independently by Artificial Analysis

Retrieved 2026-07-08

Chart source and provenance are listed in Methodology & sources below.

Retrieval date:
2026-07-08
Methodology:
Independent test run by Artificial Analysis on dedicated hardware.

02.5M5M7.5M10MLlama 4 Scout: 10M10MMLlama 4 ScoutGPT-4.1 mini: 1M1MAIGPT-4.1 miniGemini 3.5 Flash: 1M1MGGemini 3.5 FlashClaude Fable 5 (with fallback): 1M1MAClaude Fable 5 (with fal…GLM-5.2 (max): 1M1M?GLM-5.2 (max)DeepSeek V4 Pro (max): 1M1MDDeepSeek V4 Pro (max)Grok 4.3 (high): 1M1MxGrok 4.3 (high)MiniMax-M3: 1M1M?MiniMax-M3Llama 4 Maverick: 1M1MMLlama 4 MaverickGPT-4.1 nano: 1M1MAIGPT-4.1 nano

Cost per Intelligence Index Task

Weighted average cost (USD) per Artificial Analysis Intelligence Index task, segmented by token type. Lower is better · Evaluation results measured independently by Artificial Analysis

Retrieved 2026-07-08

Chart source and provenance are listed in Methodology & sources below.

Retrieval date:
2026-07-08
Methodology:
Independent test run by Artificial Analysis on dedicated hardware.

00.51GLM-5.2 (max): 00?GLM-5.2 (max)Kimi K2.6: 00KKimi K2.6Nemotron 3 Ultra: 00NNemotron 3 UltraQwen3 Coder Next: 00QQwen3 Coder NextMiniMax-M3: 00?MiniMax-M3Mistral Medium 3.1: 00MiMistral Medium 3.1Grok 4.3 (high): 00xGrok 4.3 (high)MiMo-V2-Flash: 00?MiMo-V2-FlashDeepSeek V4 Pro (max): 00DDeepSeek V4 Pro (max)GPT-4.1 mini: 00AIGPT-4.1 mini

Cost per Task

Weighted average cost (USD) per Intelligence Index task · Lower is better · Evaluation results measured independently by Artificial Analysis

Retrieved 2026-07-08

Chart source and provenance are listed in Methodology & sources below.

Retrieval date:
2026-07-08
Methodology:
Independent test run by Artificial Analysis on dedicated hardware.

$0.00$1.00$2.00$3.00$4.00$5.00Claude Fable 5 (with fallback): $2.75$2.75AClaude Fable 5 (with fal…GPT-5.5 (xhigh): $0.86$0.86AIGPT-5.5 (xhigh)Gemini 3.5 Flash: $0.59$0.59GGemini 3.5 FlashGLM-5.2 (max): $0.37$0.37?GLM-5.2 (max)Kimi K2.6: $0.35$0.35KKimi K2.6Nemotron 3 Ultra: $0.24$0.24NNemotron 3 UltraGrok 4.3 (high): $0.14$0.14xGrok 4.3 (high)MiniMax-M3: $0.12$0.12?MiniMax-M3gpt-oss-120b (high): $0.06$0.06AIgpt-oss-120b (high)GPT-4.1 mini: $0.06$0.06AIGPT-4.1 mini

Cost to Run Artificial Analysis Intelligence Index

Cost (USD) to run all evaluations in the Artificial Analysis Intelligence Index · Evaluation results measured independently by Artificial Analysis

Retrieved 2026-07-08

Chart source and provenance are listed in Methodology & sources below.

Retrieval date:
2026-07-08
Methodology:
Independent test run by Artificial Analysis on dedicated hardware.

010203040Qwen3 Coder Next: 3939QQwen3 Coder NextNemotron 3 Ultra: 3434NNemotron 3 UltraGLM-5.2 (max): 3333?GLM-5.2 (max)Kimi K2.6: 2525KKimi K2.6Mistral Medium 3.1: 1717MiMistral Medium 3.1MiniMax-M3: 1515?MiniMax-M3Grok 4.3 (high): 1313xGrok 4.3 (high)MiMo-V2-Flash: 1111?MiMo-V2-FlashMistral Large 3: 8.48.4MiMistral Large 3GPT-4.1 mini: 8.38.3AIGPT-4.1 mini

Pricing: Cache Hit, Input, and Output

Price (USD per M Tokens) · Evaluation results measured independently by Artificial Analysis

Retrieved 2026-07-08

Chart source and provenance are listed in Methodology & sources below.

Retrieval date:
2026-07-08
Methodology:
Independent test run by Artificial Analysis on dedicated hardware.

0250k500k750k1MGrok 4.3 (high): 44xGrok 4.3 (high)Nemotron 3 Ultra: 3.63.6NNemotron 3 UltraMistral Medium 3.1: 2.42.4MiMistral Medium 3.1GPT-4.1 mini: 2.12.1AIGPT-4.1 miniMistral Large 3: 22MiMistral Large 3Qwen3 Coder Next: 1.91.9QQwen3 Coder NextLlama 3.3 70B: 1.91.9MLlama 3.3 70BMiniMax-M3: 1.61.6?MiniMax-M3Llama 4 Maverick: 1.51.5MLlama 4 MaverickDeepSeek V4 Pro (max): 1.31.3DDeepSeek V4 Pro (max)cache Hitinputoutput

Time per Intelligence Index Task

Weighted average decode time (minutes) per task; excludes TTFT and overhead time · Lower is better · Evaluation results measured independently by Artificial Analysis

Retrieved 2026-07-08

Chart source and provenance are listed in Methodology & sources below.

Retrieval date:
2026-07-08
Methodology:
Independent test run by Artificial Analysis on dedicated hardware.

01234MiniMax-M3: 3.93.9?MiniMax-M3GPT-5.5 (xhigh): 33AIGPT-5.5 (xhigh)GLM-5.2 (max): 2.82.8?GLM-5.2 (max)Qwen3 Coder Next: 2.62.6QQwen3 Coder NextMinistral 3 8B: 2.52.5?Ministral 3 8BMinistral 3 14B: 2.22.2?Ministral 3 14BGemini 3.5 Flash: 2.22.2GGemini 3.5 Flashgpt-oss-120b (high): 1.91.9AIgpt-oss-120b (high)Grok 4.3 (high): 1.61.6xGrok 4.3 (high)Nemotron 3 Ultra: 1.41.4NNemotron 3 UltraGPT-4.1 mini: 0.60.6AIGPT-4.1 mini

Output Tokens per Intelligence Index Task

Weighted average number of output tokens used to run one task in the Artificial Analysis Intelligence Index · Evaluation results measured independently by Artificial Analysis

Retrieved 2026-07-08

Chart source and provenance are listed in Methodology & sources below.

Retrieval date:
2026-07-08
Methodology:
Independent test run by Artificial Analysis on dedicated hardware.

0100002000030000MiMo-V2-Flash: 2707527075?MiMo-V2-FlashMinistral 3 8B: 1568615686?Ministral 3 8BQwen3 Coder Next: 1535515355QQwen3 Coder NextLing 2.6 Flash: 1307013070?Ling 2.6 FlashMinistral 3 14B: 1174611746?Ministral 3 14BMiniMax-M3: 1162311623?MiniMax-M3Mistral Small 3.2: 80688068MiMistral Small 3.2Nemotron 3 Ultra: 77977797NNemotron 3 UltraGPT-4.1 nano: 63276327AIGPT-4.1 nanoMistral Medium 3.1: 50475047MiMistral Medium 3.1GPT-4.1 mini: 29572957AIGPT-4.1 mini

Methodology

Methodology & Provenance

This page is rendered from the normalized profile and page JSON for GPT-4.1 mini.

Benchmark values are preserved as normalized; only layout, disclosure ordering, and typography are adjusted for readability.

Frequently Asked Questions

Model FAQs & Technical Disclosures

GPT-4.1 mini was released on April 14, 2025.