z-ai

Z.ai: GLM 4.6V

Z.ai: GLM 4.6V is capable of processing text, images, and video inputs within a context length of 131,072 tokens. It supports reasoning and tools integration but does not offer structured output or open weights. This model stands out with its broad input modalities and the ability to handle complex tasks through reasoning. Considering a price point of $0.3 per million input tokens and $0.9 per million output tokens, Z.ai: GLM 4.6V is more expensive than average but offers robust capabilities across multiple modalities. Its blended benchmark score of 16.7 places it in the upper-middle range among models tested, making it a suitable choice for applications requiring versatile input handling and reasoning support.

Quality Score
100/100
price + capability + benchmarks
Input Price
$0.30
per 1M tokens
Output Price
$0.90
per 1M tokens
Context Window
131,072
tokens

Benchmark results

Independent, published benchmarks. Blended score 16.7 across 1 benchmark, last refreshed 2026-08-13. How scoring works →

BenchmarkMeasuresScore
AI Index broad capability composite 18.0
Model ID
z-ai/glm-4.6v
Vendor
z-ai
Released
December 2025
Tokenizer
Other
Input Modalities
image, text, video
Output Modalities
text
Max Output
32,768 tokens
Tool Calling
✓ supported
Structured Output
✓ supported
Reasoning Mode
✓ supported
Vision
✓ accepts images
Audio
no
Moderated
no

What it costs in practice

Computed from the current $0.30/M input and $0.90/M output rates. Run your own numbers →

JobTokensCost
Summarize a 50-page report 30k in / 1.5k out $0.01
Classify 1,000 customer emails 500k in / 50k out $0.20
A month of a busy support chatbot 5M in / 2M out $3.30

Similar models

Quick answers

How much does Z.ai: GLM 4.6V cost?
$0.30 per million input tokens and $0.90 per million output tokens.
What is Z.ai: GLM 4.6V's context window?
131,072 tokens, roughly 196 pages of text in a single request.
Does Z.ai: GLM 4.6V support tool calling?
It supports tool calling, structured output, a reasoning mode.
Can Z.ai: GLM 4.6V process images?
Yes, it accepts image input.