nvidia

NVIDIA: Nemotron 3 Ultra (batch)

NVIDIA's Nemotron 3 Ultra is designed for text-based tasks with an extensive context length of up to 512,288 tokens and supports reasoning and tools integration. Its versatility lies in handling complex textual inputs while enabling structured outputs. The model demonstrates strong performance on AI Index coding but shows less prowess in the agentic domain. Given its blended benchmark score of 63.0 across three independent tests, Nemotron 3 Ultra is a solid choice for those requiring robust text processing and reasoning capabilities. However, with pricing at $0.3 per million input tokens and $1.8 per million output tokens, it may be more expensive than other options available in the market. This cost-benefit analysis suggests that Nemotron 3 Ultra is particularly suitable for applications where high precision and detailed reasoning are critical, despite its higher price point.

Quality Score
99/100
price + capability + benchmarks
Input Price
$0.30
per 1M tokens
Output Price
$1.80
per 1M tokens
Context Window
512,288
tokens

Benchmark results

Independent, published benchmarks. Blended score 63.0 across 3 benchmarks, last refreshed 2026-08-13. How scoring works →

BenchmarkMeasuresScore
AI Index broad capability composite 63.2
AI Index Coding software engineering tasks 81.3
AI Index Agentic multi-step tool-using tasks 45.4
Model ID
nvidia/nemotron-3-ultra-550b-a55b:batch
Vendor
nvidia
Released
June 2026
Tokenizer
Other
Input Modalities
text
Output Modalities
text
Max Output
default
Tool Calling
✓ supported
Structured Output
✓ supported
Reasoning Mode
✓ supported
Vision
text only
Audio
no
Moderated
no

What it costs in practice

Computed from the current $0.30/M input and $1.80/M output rates. Run your own numbers →

JobTokensCost
Summarize a 50-page report 30k in / 1.5k out $0.01
Classify 1,000 customer emails 500k in / 50k out $0.24
A month of a busy support chatbot 5M in / 2M out $5.10

Similar models

Quick answers

How much does NVIDIA: Nemotron 3 Ultra (batch) cost?
$0.30 per million input tokens and $1.80 per million output tokens.
What is NVIDIA: Nemotron 3 Ultra (batch)'s context window?
512,288 tokens, roughly 768 pages of text in a single request.
Does NVIDIA: Nemotron 3 Ultra (batch) support tool calling?
It supports tool calling, structured output, a reasoning mode.
Can NVIDIA: Nemotron 3 Ultra (batch) process images?
No, it is text-only on the input side.