Skip to content

EthenEthenEthen

NVIDIA · Flagships Analysis

NVIDIA Nemotron Nano 12B v2 VL

Intelligence, Performance & Price Analysis

Canonical slug: nvidia-nemotron-nano-12b-v2-vl · Canonical model registry at build time

Intelligence

5 score

Speed

241.4 output tokens/sec

Latency

1.07s TTFT

Input Price

$0.20 / 1M tokens

Output Price

$0.60 / 1M tokens

Executive Assessment

Routing Verdict & Tradeoffs

Ethen Routing Verdict

Budget-friendly / task-specific model

NVIDIA Nemotron Nano 12B v2 VL (Non-reasoning) scores 5 (estimated) on the Artificial Analysis Intelligence Index, placing it below average among other open weight non-reasoning models of similar size (median: 6). NVIDIA Nemotron Nano 12B v2 VL (Non-reasoning) generates output at 241.4 tokens per second (based on the median across providers serving the model), which is well above average compared to other open weight non-reasoning models of similar size (median: 101.8 t/s). NVIDIA Nemotron Nano 12B v2 VL (Non-reasoning) costs $0.20 per 1M input tokens (somewhat higher than average, median: $0.15) and $0.60 per 1M output tokens (somewhat higher than average, median: $0.32), based on the median across providers serving the model.

Best for high-volume, simple, or domain-specific tasks where cost or speed matters more than deep reasoning.

Task Fit Assessment

WorkloadRatingNotes
Complex reasoning & agentic workflowssuboptimalIntelligence score 5 is better suited for straightforward tasks than multi-step reasoning.
High-volume chat & customer-facingoptimalOutput speed 241.4 tokens/sec is adequate for chat.
Latency-sensitive applicationsoptimalTTFT 1.07s — adequate latency for most interactive use cases.
Cost-sensitive pipelinesoptimalOutput pricing at $0.60 is reasonable for moderate volume.

Cost Pressure Analysis

Low

Pricing is competitive — input $0.20, output $0.60. Suitable for sustained production use.

Technical Specifications

Architecture and Limits

SpecificationValueValidation Authority
Model typeOpen weightsinferred
ReasoningNofaq
Input modalitiesNVIDIA Nemotron Nano 12B v2 VL (Non-reasoning) supports text and image input.faq
Output modalitiesNVIDIA Nemotron Nano 12B v2 VL (Non-reasoning) supports text output.faq
Context window130k tokensfaq
Open weights / sourceYes, NVIDIA Nemotron Nano 12B v2 VL (Non-reasoning) is open weights. The model weights are publicly available and can be downloaded for self-hosting.faq
ParametersNVIDIA Nemotron Nano 12B v2 VL (Non-reasoning) has 13.2 billion parameters.faq
LicenseNVIDIA Nemotron Nano 12B v2 VL (Non-reasoning) is released under the Nvidia Open Model License license. This license allows commercial use.faq
API availabilityYes, NVIDIA Nemotron Nano 12B v2 VL (Non-reasoning) is available via API through 2 providers.faq

Evidence charts

Profile visualizations

Charts are the same canonical evidence cards previously published for this slug, contained inside the D18D page grammar.

AA-Omniscience Index

AA-Omniscience Index (higher is better) measures knowledge reliability and hallucination. It rewards correct answers, penalizes hallucinations, and has no penalty for refusing to answer. Scores range from -100 to 100, where 0 means as many correct as incorrect answers, and negative scores mean more incorrect than correct. · Evaluation results measured independently by Artificial Analysis

Retrieved 2026-07-08

Chart source and provenance are listed in Methodology & sources below.

Retrieval date:
2026-07-08
Methodology:
Independent test run by Artificial Analysis on dedicated hardware.

-200204060Claude Fable 5 (with fallback): 4040AClaude Fable 5 (with fal…Gemini 3.5 Flash: 2323GGemini 3.5 FlashGPT-5.5 (xhigh): 2020AIGPT-5.5 (xhigh)Grok 4.3 (high): 1818xGrok 4.3 (high)Kimi K2.6: 6.46.4KKimi K2.6Muse Spark: 4.14.1?Muse SparkGLM-5.2 (max): 44?GLM-5.2 (max)MiniMax-M3: 1.41.4?MiniMax-M3Nemotron 3 Ultra: -0.8-0.8NNemotron 3 UltraDeepSeek V4 Pro (max): -10-10DDeepSeek V4 Pro (max)

Artificial Analysis Intelligence Index by Open Weights / Proprietary

Artificial Analysis Intelligence Index v4.1 incorporates 9 evaluations: GDPval-AA v2, 𝜏³-Banking, Terminal-Bench v2.1, SciCode, Humanity's Last Exam, GPQA Diamond, CritPt, AA-Omniscience, AA-LCR · Evaluation results measured independently by Artificial Analysis

Retrieved 2026-07-08

Chart source and provenance are listed in Methodology & sources below.

Retrieval date:
2026-07-08
Methodology:
Independent test run by Artificial Analysis on dedicated hardware.

0204060Claude Fable 5 (with fallback): 6060AClaude Fable 5 (with fal…GPT-5.5 (xhigh): 5555AIGPT-5.5 (xhigh)GLM-5.2 (max): 5151?GLM-5.2 (max)Gemini 3.5 Flash: 5050GGemini 3.5 FlashMiniMax-M3: 4444?MiniMax-M3DeepSeek V4 Pro (max): 4444DDeepSeek V4 Pro (max)Kimi K2.6: 4444KKimi K2.6Muse Spark: 4343?Muse SparkNemotron 3 Ultra: 3838NNemotron 3 UltraGrok 4.3 (high): 3838xGrok 4.3 (high)

Artificial Analysis Intelligence Index

Artificial Analysis Intelligence Index v4.1 incorporates 9 evaluations: GDPval-AA v2, 𝜏³-Banking, Terminal-Bench v2.1, SciCode, Humanity's Last Exam, GPQA Diamond, CritPt, AA-Omniscience, AA-LCR · Evaluation results measured independently by Artificial Analysis

Retrieved 2026-07-08

Chart source and provenance are listed in Methodology & sources below.

Retrieval date:
2026-07-08
Methodology:
Independent test run by Artificial Analysis on dedicated hardware.

0204060Claude Fable 5 (with fallback): 6060AClaude Fable 5 (with fal…GPT-5.5 (xhigh): 5555AIGPT-5.5 (xhigh)GLM-5.2 (max): 5151?GLM-5.2 (max)Gemini 3.5 Flash: 5050GGemini 3.5 FlashMiniMax-M3: 4444?MiniMax-M3DeepSeek V4 Pro (max): 4444DDeepSeek V4 Pro (max)Kimi K2.6: 4444KKimi K2.6Muse Spark: 4343?Muse SparkNemotron 3 Ultra: 3838NNemotron 3 UltraGrok 4.3 (high): 3838xGrok 4.3 (high)

Artificial Analysis Openness Index: Score

Openness Index assesses model openness on a 0 to 100 normalized scale (higher is more open) · Evaluation results measured independently by Artificial Analysis

Retrieved 2026-07-08

Chart source and provenance are listed in Methodology & sources below.

Retrieval date:
2026-07-08
Methodology:
Independent test run by Artificial Analysis on dedicated hardware.

050100Nemotron 3 Ultra: 8383NNemotron 3 UltraNVIDIA Nemotron Nano 12B v2 VL: 7272NNVIDIA Nemotron Nano 12B…Granite 4.1 30B: 6161?Granite 4.1 30BDeepSeek V4 Pro (max): 5050DDeepSeek V4 Pro (max)Gemma 3 27B: 5050?Gemma 3 27BGemma 3 12B: 5050?Gemma 3 12BGLM-5.2 (max): 4444?GLM-5.2 (max)gpt-oss-120b (high): 3939AIgpt-oss-120b (high)Qwen3.6 27B: 3939QQwen3.6 27BMistral Small 3.2: 3939MiMistral Small 3.2

Intelligence

Artificial Analysis Intelligence Index · Higher is better · Evaluation results measured independently by Artificial Analysis

Retrieved 2026-07-08

Chart source and provenance are listed in Methodology & sources below.

Retrieval date:
2026-07-08
Methodology:
Independent test run by Artificial Analysis on dedicated hardware.

0204060Claude Fable 5 (with fallback): 6060AClaude Fable 5 (with fal…GPT-5.5 (xhigh): 5555AIGPT-5.5 (xhigh)GLM-5.2 (max): 5151?GLM-5.2 (max)Gemini 3.5 Flash: 5050GGemini 3.5 FlashMiniMax-M3: 4444?MiniMax-M3DeepSeek V4 Pro (max): 4444DDeepSeek V4 Pro (max)Kimi K2.6: 4444KKimi K2.6Muse Spark: 4343?Muse SparkNemotron 3 Ultra: 3838NNemotron 3 UltraGrok 4.3 (high): 3838xGrok 4.3 (high)NVIDIA Nemotron Nano 12B v2 VL: 4.64.6NNVIDIA Nemotron Nano 12B…

Output Speed

Output tokens per second · Higher is better · Evaluation results measured independently by Artificial Analysis

Retrieved 2026-07-08

Chart source and provenance are listed in Methodology & sources below.

Retrieval date:
2026-07-08
Methodology:
Independent test run by Artificial Analysis on dedicated hardware.

088175263350gpt-oss-120b (high): 314314AIgpt-oss-120b (high)Nemotron 3 Ultra: 250250NNemotron 3 UltraNVIDIA Nemotron Nano 12B v2 VL: 241241NNVIDIA Nemotron Nano 12B…GLM-5.2 (max): 218218?GLM-5.2 (max)Gemini 3.5 Flash: 192192GGemini 3.5 FlashMistral Small 3.1: 171171MiMistral Small 3.1Grok 4.3 (high): 164164xGrok 4.3 (high)Mistral Small 3.2: 147147MiMistral Small 3.2Llama 3.1 8B: 144144MLlama 3.1 8BMiniMax-M3: 9999?MiniMax-M3

Speed

Output tokens per second · Higher is better · Evaluation results measured independently by Artificial Analysis

Retrieved 2026-07-08

Chart source and provenance are listed in Methodology & sources below.

Retrieval date:
2026-07-08
Methodology:
Independent test run by Artificial Analysis on dedicated hardware.

088175263350gpt-oss-120b (high): 314314AIgpt-oss-120b (high)Nemotron 3 Ultra: 250250NNemotron 3 UltraNVIDIA Nemotron Nano 12B v2 VL: 241241NNVIDIA Nemotron Nano 12B…GLM-5.2 (max): 218218?GLM-5.2 (max)Gemini 3.5 Flash: 192192GGemini 3.5 FlashGrok 4.3 (high): 164164xGrok 4.3 (high)MiniMax-M3: 9999?MiniMax-M3GPT-5.5 (xhigh): 8888AIGPT-5.5 (xhigh)Kimi K2.6: 7676KKimi K2.6DeepSeek V4 Pro (max): 7272DDeepSeek V4 Pro (max)

End-to-End Response Time

Seconds to output 500 tokens, including reasoning model 'thinking' time · Lower is better · Evaluation results measured independently by Artificial Analysis

Retrieved 2026-07-08

Chart source and provenance are listed in Methodology & sources below.

Retrieval date:
2026-07-08
Methodology:
Independent test run by Artificial Analysis on dedicated hardware.

0s2.5s5s7.5s10sQwen3.6 27B: 8.3s8.3sQQwen3.6 27BDevstral Small 2: 7.4s7.4s?Devstral Small 2Claude Fable 5 (with fallback): 7.1s7.1sAClaude Fable 5 (with fal…DeepSeek V4 Pro (max): 7s7sDDeepSeek V4 Pro (max)Kimi K2.6: 6.6s6.6sKKimi K2.6Ministral 3 14B: 6.4s6.4s?Ministral 3 14BGPT-5.5 (xhigh): 5.7s5.7sAIGPT-5.5 (xhigh)Ministral 3 8B: 5.4s5.4s?Ministral 3 8BMiniMax-M3: 5s5s?MiniMax-M3Llama 3.1 8B: 3.5s3.5sMLlama 3.1 8BNVIDIA Nemotron Nano 12B v2 VL: 2.1s2.1sNNVIDIA Nemotron Nano 12B…

Latency: Time To First Answer Token

Seconds to first answer token received · Accounts for reasoning model 'thinking' time · Evaluation results measured independently by Artificial Analysis

Retrieved 2026-07-08

Chart source and provenance are listed in Methodology & sources below.

Retrieval date:
2026-07-08
Methodology:
Independent test run by Artificial Analysis on dedicated hardware.

0s25s50s75s100sDeepSeek V4 Pro (max): 61s61sDDeepSeek V4 Pro (max)Kimi K2.6: 58s58sKKimi K2.6MiniMax-M3: 20s20s?MiniMax-M3GLM-5.2 (max): 9.2s9.2s?GLM-5.2 (max)Nemotron 3 Ultra: 9.1s9.1sNNemotron 3 Ultragpt-oss-120b (high): 6.4s6.4sAIgpt-oss-120b (high)Mistral Small 3.2: 0s0sMiMistral Small 3.2Mistral Small 3.1: 0s0sMiMistral Small 3.1Ministral 3 8B: 0s0s?Ministral 3 8BMinistral 3 14B: 0s0s?Ministral 3 14BNVIDIA Nemotron Nano 12B v2 VL: 0s0sNNVIDIA Nemotron Nano 12B…

Context Window

Context window: tokens limit · Higher is better · Evaluation results measured independently by Artificial Analysis

Retrieved 2026-07-08

Chart source and provenance are listed in Methodology & sources below.

Retrieval date:
2026-07-08
Methodology:
Independent test run by Artificial Analysis on dedicated hardware.

0250k500k750k1MGemini 3.5 Flash: 1M1MGGemini 3.5 FlashClaude Fable 5 (with fallback): 1M1MAClaude Fable 5 (with fal…GLM-5.2 (max): 1M1M?GLM-5.2 (max)DeepSeek V4 Pro (max): 1M1MDDeepSeek V4 Pro (max)Grok 4.3 (high): 1M1MxGrok 4.3 (high)MiniMax-M3: 1M1M?MiniMax-M3GPT-5.5 (xhigh): 922k922kAIGPT-5.5 (xhigh)Nemotron 3 Ultra: 262k262kNNemotron 3 UltraMuse Spark: 262k262k?Muse SparkQwen3.6 27B: 262k262kQQwen3.6 27BNVIDIA Nemotron Nano 12B v2 VL: 128k128kNNVIDIA Nemotron Nano 12B…

Cost per Intelligence Index Task

Weighted average cost (USD) per Artificial Analysis Intelligence Index task, segmented by token type. Lower is better · Evaluation results measured independently by Artificial Analysis

Retrieved 2026-07-08

Chart source and provenance are listed in Methodology & sources below.

Retrieval date:
2026-07-08
Methodology:
Independent test run by Artificial Analysis on dedicated hardware.

00.51Claude Fable 5 (with fallback): 0.40.4AClaude Fable 5 (with fal…GPT-5.5 (xhigh): 0.10.1AIGPT-5.5 (xhigh)Gemini 3.5 Flash: 0.10.1GGemini 3.5 FlashQwen3.6 27B: 00QQwen3.6 27BGLM-5.2 (max): 00?GLM-5.2 (max)Kimi K2.6: 00KKimi K2.6Nemotron 3 Ultra: 00NNemotron 3 UltraMiniMax-M3: 00?MiniMax-M3Grok 4.3 (high): 00xGrok 4.3 (high)DeepSeek V4 Pro (max): 00DDeepSeek V4 Pro (max)

Cost per Task

Weighted average cost (USD) per Intelligence Index task · Lower is better · Evaluation results measured independently by Artificial Analysis

Retrieved 2026-07-08

Chart source and provenance are listed in Methodology & sources below.

Retrieval date:
2026-07-08
Methodology:
Independent test run by Artificial Analysis on dedicated hardware.

$0.00$1.00$2.00$3.00$4.00$5.00Claude Fable 5 (with fallback): $2.75$2.75AClaude Fable 5 (with fal…GPT-5.5 (xhigh): $0.86$0.86AIGPT-5.5 (xhigh)Gemini 3.5 Flash: $0.59$0.59GGemini 3.5 FlashGLM-5.2 (max): $0.37$0.37?GLM-5.2 (max)Kimi K2.6: $0.35$0.35KKimi K2.6Nemotron 3 Ultra: $0.24$0.24NNemotron 3 UltraGrok 4.3 (high): $0.14$0.14xGrok 4.3 (high)MiniMax-M3: $0.12$0.12?MiniMax-M3gpt-oss-120b (high): $0.06$0.06AIgpt-oss-120b (high)DeepSeek V4 Pro (max): $0.04$0.04DDeepSeek V4 Pro (max)

Cost to Run Artificial Analysis Intelligence Index

Cost (USD) to run all evaluations in the Artificial Analysis Intelligence Index · Evaluation results measured independently by Artificial Analysis

Retrieved 2026-07-08

Chart source and provenance are listed in Methodology & sources below.

Retrieval date:
2026-07-08
Methodology:
Independent test run by Artificial Analysis on dedicated hardware.

0200400600Claude Fable 5 (with fallback): 508508AClaude Fable 5 (with fal…GPT-5.5 (xhigh): 132132AIGPT-5.5 (xhigh)Gemini 3.5 Flash: 114114GGemini 3.5 FlashQwen3.6 27B: 8888QQwen3.6 27BNemotron 3 Ultra: 3434NNemotron 3 UltraGLM-5.2 (max): 3333?GLM-5.2 (max)Kimi K2.6: 2525KKimi K2.6MiniMax-M3: 1515?MiniMax-M3Grok 4.3 (high): 1313xGrok 4.3 (high)DeepSeek V4 Pro (max): 7.27.2DDeepSeek V4 Pro (max)

Pricing: Cache Hit, Input, and Output

Price (USD per M Tokens) · Evaluation results measured independently by Artificial Analysis

Retrieved 2026-07-08

Chart source and provenance are listed in Methodology & sources below.

Retrieval date:
2026-07-08
Methodology:
Independent test run by Artificial Analysis on dedicated hardware.

0250k500k750k1MClaude Fable 5 (with fallback): 6161AClaude Fable 5 (with fal…GPT-5.5 (xhigh): 3636AIGPT-5.5 (xhigh)Gemini 3.5 Flash: 1111GGemini 3.5 FlashGLM-5.2 (max): 6.16.1?GLM-5.2 (max)Kimi K2.6: 5.15.1KKimi K2.6Qwen3.6 27B: 4.24.2QQwen3.6 27BGrok 4.3 (high): 44xGrok 4.3 (high)Nemotron 3 Ultra: 3.63.6NNemotron 3 UltraMiniMax-M3: 1.61.6?MiniMax-M3DeepSeek V4 Pro (max): 1.31.3DDeepSeek V4 Pro (max)NVIDIA Nemotron Nano 12B v2 VL: 0.80.8NNVIDIA Nemotron Nano 12B…cache Hitinputoutput

Time per Intelligence Index Task

Weighted average decode time (minutes) per task; excludes TTFT and overhead time · Lower is better · Evaluation results measured independently by Artificial Analysis

Retrieved 2026-07-08

Chart source and provenance are listed in Methodology & sources below.

Retrieval date:
2026-07-08
Methodology:
Independent test run by Artificial Analysis on dedicated hardware.

0510Kimi K2.6: 8.28.2KKimi K2.6DeepSeek V4 Pro (max): 6.86.8DDeepSeek V4 Pro (max)Claude Fable 5 (with fallback): 4.84.8AClaude Fable 5 (with fal…MiniMax-M3: 3.93.9?MiniMax-M3GPT-5.5 (xhigh): 33AIGPT-5.5 (xhigh)Qwen3.6 27B: 33QQwen3.6 27BGLM-5.2 (max): 2.82.8?GLM-5.2 (max)Ministral 3 8B: 2.52.5?Ministral 3 8BMinistral 3 14B: 2.22.2?Ministral 3 14BGemini 3.5 Flash: 2.22.2GGemini 3.5 Flash

Model Size: Total and Active Parameters

Comparison between total model parameters and parameters active during inference · Evaluation results measured independently by Artificial Analysis

Retrieved 2026-07-08

Chart source and provenance are listed in Methodology & sources below.

Retrieval date:
2026-07-08
Methodology:
Independent test run by Artificial Analysis on dedicated hardware.

0250k500k750k1MDeepSeek V4 Pro (max): 1.6k1.6kDDeepSeek V4 Pro (max)Kimi K2.6: 1k1kKKimi K2.6GLM-5.2 (max): 753753?GLM-5.2 (max)Nemotron 3 Ultra: 550550NNemotron 3 UltraMiniMax-M3: 428428?MiniMax-M3gpt-oss-120b (high): 117117AIgpt-oss-120b (high)EXAONE 4.5 33B: 3434?EXAONE 4.5 33BGranite 4.1 30B: 3030?Granite 4.1 30BQwen3.6 27B: 2828QQwen3.6 27BGemma 3 27B: 2727?Gemma 3 27BNVIDIA Nemotron Nano 12B v2 VL: 1313NNVIDIA Nemotron Nano 12B…passive Paramsactive Params

Methodology

Methodology & Provenance

This page is rendered from the normalized profile and page JSON for NVIDIA Nemotron Nano 12B v2 VL.

Benchmark values are preserved as normalized; only layout, disclosure ordering, and typography are adjusted for readability.

Frequently Asked Questions

Model FAQs & Technical Disclosures

NVIDIA Nemotron Nano 12B v2 VL (Non-reasoning) was released on October 28, 2025.