GLM-5.3 Flash
Native multimodal open model with hybrid sparse/linear attention for long context.
GLM-5.3-Flash is a native multimodal model suited for efficient coding and long-horizon agent tasks. Its hybrid sparse and linear attention architecture maintains accurate long-context behavior.
- Provider
- Z.ai
- Type
- Language model
- Released
- Aug 26, 2026
- Context window
- 1M tokens
- Max output
- 944K tokens
- Price
- $0.15 input / $0.50 output per 1M tokens
- Input
- text, image, video
- Output
- text
- Open weights
- Yes
- License
- MIT
Best for
- Coding
- Agents
- High volume / low cost
- Long context
- Vision
Strengths
- Excellent intelligence per dollar
- Open weights
More from Z.ai
- GLM-5.3 — Z.ai. Open reasoning model for complex software engineering and long-horizon agents.