Z.ai: GLM 4.6V
Z.ai: GLM 4.6V is capable of processing text, images, and video inputs within a context length of 131,072 tokens. It supports reasoning and tools integration but does not offer structured output or open weights. This model stands out with its broad input modalities and the ability to handle complex tasks through reasoning. Considering a price point of $0.3 per million input tokens and $0.9 per million output tokens, Z.ai: GLM 4.6V is more expensive than average but offers robust capabilities across multiple modalities. Its blended benchmark score of 16.7 places it in the upper-middle range among models tested, making it a suitable choice for applications requiring versatile input handling and reasoning support.
Benchmark results
Independent, published benchmarks. Blended score 16.7 across 1 benchmark, last refreshed 2026-08-13. How scoring works →
| Benchmark | Measures | Score |
|---|---|---|
| AI Index | broad capability composite | 18.0 |
- Model ID
- z-ai/glm-4.6v
- Vendor
- z-ai
- Released
- December 2025
- Tokenizer
- Other
- Input Modalities
- image, text, video
- Output Modalities
- text
- Max Output
- 32,768 tokens
- Tool Calling
- ✓ supported
- Structured Output
- ✓ supported
- Reasoning Mode
- ✓ supported
- Vision
- ✓ accepts images
- Audio
- no
- Moderated
- no
What it costs in practice
Computed from the current $0.30/M input and $0.90/M output rates. Run your own numbers →
| Job | Tokens | Cost |
|---|---|---|
| Summarize a 50-page report | 30k in / 1.5k out | $0.01 |
| Classify 1,000 customer emails | 500k in / 50k out | $0.20 |
| A month of a busy support chatbot | 5M in / 2M out | $3.30 |
Similar models
Z.ai: GLM 5V Turbo
Z.ai: GLM 4.7 Flash
Z.ai: GLM 4.7
Z.ai: GLM 4.6
Z.ai: GLM 5.2 (batch)
Z.ai: GLM 5.2
Quick answers
- How much does Z.ai: GLM 4.6V cost?
- $0.30 per million input tokens and $0.90 per million output tokens.
- What is Z.ai: GLM 4.6V's context window?
- 131,072 tokens, roughly 196 pages of text in a single request.
- Does Z.ai: GLM 4.6V support tool calling?
- It supports tool calling, structured output, a reasoning mode.
- Can Z.ai: GLM 4.6V process images?
- Yes, it accepts image input.