glm-5.2
GLM-5.2 is Z.ai’s latest flagship model, designed for long-horizon task scenarios. Compared with GLM-5.1, it delivers significant improvements in long-horizon task performance. The 753B-parameter model supports a stable 1-million-token context window, offers stronger coding capabilities, and provides multiple thinking-effort levels, enabling a flexible balance between performance and latency. GLM-5.2 introduces the IndexShare architecture optimization, reducing per-token FLOPs by 2.9× at a 1-million-token context length. It also improves the MTP layer to support speculative decoding, increasing the maximum acceptance length by up to 20%.
Pricing snapshot
Last updated 2026-07-15. Live pricing may change after publication.
Supported API formats
- anthropic
- openai
Integrate through one API key
Review the protocol guide, then use this model ID in your request.