GLM-5.3 Flash

Native multimodal open model with hybrid sparse/linear attention for long context.

GLM-5.3-Flash is a native multimodal model suited for efficient coding and long-horizon agent tasks. Its hybrid sparse and linear attention architecture maintains accurate long-context behavior.

Provider
Z.ai
Type
Language model
Released
Aug 26, 2026
Context window
1M tokens
Max output
944K tokens
Price
$0.15 input / $0.50 output per 1M tokens
Input
text, image, video
Output
text
Open weights
Yes
License
MIT

Best for

Strengths

More from Z.ai

Sources