NVIDIA: Nemotron 3 Ultra (batch)
NVIDIA's Nemotron 3 Ultra is designed for text-based tasks with an extensive context length of up to 512,288 tokens and supports reasoning and tools integration. Its versatility lies in handling complex textual inputs while enabling structured outputs. The model demonstrates strong performance on AI Index coding but shows less prowess in the agentic domain. Given its blended benchmark score of 63.0 across three independent tests, Nemotron 3 Ultra is a solid choice for those requiring robust text processing and reasoning capabilities. However, with pricing at $0.3 per million input tokens and $1.8 per million output tokens, it may be more expensive than other options available in the market. This cost-benefit analysis suggests that Nemotron 3 Ultra is particularly suitable for applications where high precision and detailed reasoning are critical, despite its higher price point.
Benchmark results
Independent, published benchmarks. Blended score 63.0 across 3 benchmarks, last refreshed 2026-08-13. How scoring works →
| Benchmark | Measures | Score |
|---|---|---|
| AI Index | broad capability composite | 63.2 |
| AI Index Coding | software engineering tasks | 81.3 |
| AI Index Agentic | multi-step tool-using tasks | 45.4 |
- Model ID
- nvidia/nemotron-3-ultra-550b-a55b:batch
- Vendor
- nvidia
- Released
- June 2026
- Tokenizer
- Other
- Input Modalities
- text
- Output Modalities
- text
- Max Output
- default
- Tool Calling
- ✓ supported
- Structured Output
- ✓ supported
- Reasoning Mode
- ✓ supported
- Vision
- text only
- Audio
- no
- Moderated
- no
What it costs in practice
Computed from the current $0.30/M input and $1.80/M output rates. Run your own numbers →
| Job | Tokens | Cost |
|---|---|---|
| Summarize a 50-page report | 30k in / 1.5k out | $0.01 |
| Classify 1,000 customer emails | 500k in / 50k out | $0.24 |
| A month of a busy support chatbot | 5M in / 2M out | $5.10 |
Similar models
NVIDIA: Nemotron 3.5 Lightning
NVIDIA: Nemotron 3 Super
NVIDIA: Nemotron 3 Nano 30B A3B
NVIDIA: Nemotron 3 Nano Omni (free)
NVIDIA: Nemotron 3 Ultra
NVIDIA: Nemotron 3 Super (free)
Quick answers
- How much does NVIDIA: Nemotron 3 Ultra (batch) cost?
- $0.30 per million input tokens and $1.80 per million output tokens.
- What is NVIDIA: Nemotron 3 Ultra (batch)'s context window?
- 512,288 tokens, roughly 768 pages of text in a single request.
- Does NVIDIA: Nemotron 3 Ultra (batch) support tool calling?
- It supports tool calling, structured output, a reasoning mode.
- Can NVIDIA: Nemotron 3 Ultra (batch) process images?
- No, it is text-only on the input side.