Qwen: Qwen2.5 VL 72B Instruct
Qwen: Qwen2.5 VL 72B Instruct is an AI model designed to process text and image inputs within a context length of up to 128,000 tokens; however, it does not support reasoning or structured output, nor can it interact with external tools. Given its capabilities, this model might be suitable for tasks involving complex text analysis and visual data interpretation but lacks features that could aid in more advanced logical processing. Considering the pricing at $0.25 per million input tokens and $0.75 per million output tokens, Qwen is competitive for projects requiring significant text and image handling, especially when budget is a primary concern. Until further benchmarking confirms its performance against industry standards, users should weigh these costs against the model’s current limitations before making a selection.
- Model ID
- qwen/qwen2.5-vl-72b-instruct
- Vendor
- qwen
- Released
- February 2025
- Tokenizer
- Qwen
- Input Modalities
- text, image
- Output Modalities
- text
- Max Output
- default
- Tool Calling
- not supported
- Structured Output
- ✓ supported
- Reasoning Mode
- not supported
- Vision
- ✓ accepts images
- Audio
- no
- Moderated
- no
What it costs in practice
Computed from the current $0.25/M input and $0.75/M output rates. Run your own numbers →
| Job | Tokens | Cost |
|---|---|---|
| Summarize a 50-page report | 30k in / 1.5k out | under $0.01 |
| Classify 1,000 customer emails | 500k in / 50k out | $0.16 |
| A month of a busy support chatbot | 5M in / 2M out | $2.75 |
Price & spec history
Tracked daily by PicksByModel since 2026-07-17.
| Date | Input /M | Output /M | Context |
|---|---|---|---|
| 2026-08-05 | $0.25 | $0.75 | 128,000 |
| 2026-07-22 | $0.80 | $1.00 | 128,000 |
| 2026-07-17 | $0.80 | $1.00 | 131,072 |
Similar models
Qwen: Qwen3 30B A3B Thinking 2507
Qwen: Qwen2.5 7B Instruct
Qwen2.5 72B Instruct
Qwen: Qwen3 235B A22B
Qwen: Qwen3 VL 32B Instruct
Qwen: Qwen3 30B A3B
Quick answers
- How much does Qwen: Qwen2.5 VL 72B Instruct cost?
- $0.25 per million input tokens and $0.75 per million output tokens.
- What is Qwen: Qwen2.5 VL 72B Instruct's context window?
- 128,000 tokens, roughly 192 pages of text in a single request.
- Does Qwen: Qwen2.5 VL 72B Instruct support tool calling?
- It supports structured output.
- Can Qwen: Qwen2.5 VL 72B Instruct process images?
- Yes, it accepts image input.