Z.ai's flagship model for long-horizon tasks — a solid 1M-token context, flexible coding effort levels, and MIT open-source weights.
No credit card required
Source-backed capabilities from Z.ai's official release
A stable 1M-token context engineered for sustained long-horizon work — not just a larger window, but reliable quality across messy coding-agent trajectories.
Multiple thinking effort levels let you balance performance against latency. Max effort unlocks additional compute for challenging tasks.
Every 4 sparse attention layers share a lightweight indexer, reducing per-token FLOPs by 2.9× at 1M context without quality loss.
Pure open — no regional limits, no technical borders. Weights available on HuggingFace and ModelScope for local deployment.
Rule-based + LLM-judge pipeline blocks reward hacking during agentic RL, ensuring real task-solving — not shortcut exploitation.
Top open-source model across FrontierSWE, PostTrainBench, and SWE-Marathon — trailing only the Opus 4.8 series on long-horizon benchmarks.
As published by Z.ai
| Benchmark | GLM-5.2 | Context |
|---|---|---|
| Terminal-Bench 2.1 (Terminus-2) | 81.0 | Strongest open-source model; within 4 pts of Opus 4.8 |
| SWE-bench Pro | 62.1 | Up from 58.4 in GLM-5.1; ahead of GPT-5.5 (58.6) |
| FrontierSWE Dominance | 74.4 | Trails Opus 4.8 by only 1%; top open-source model |
| PostTrainBench | 34.3 | 2nd overall, ahead of Opus 4.7 and GPT-5.5 |
| AIME 2026 | 99.2 | Near-perfect math reasoning |
| GPQA-Diamond | 91.2 | Graduate-level science QA |
GLM-5.2 vs. GLM-5.1 and frontier closed-source models
| Benchmark | GLM-5.2 | GLM-5.1 | Opus 4.8 | GPT-5.5 | Gemini 3.1 Pro |
|---|---|---|---|---|---|
| HLE | 40.5 | 31.0 | 49.8* | 41.4* | 45.0 |
| HLE w/ Tools | 54.7 | 52.3 | 57.9* | 52.2* | 51.4* |
| AIME 2026 | 99.2 | 95.3 | 95.7 | 98.3 | 98.2 |
| GPQA-Diamond | 91.2 | 86.2 | 93.6 | 93.6 | 94.3 |
| Benchmark | GLM-5.2 | GLM-5.1 | Opus 4.8 | GPT-5.5 | Gemini 3.1 Pro |
|---|---|---|---|---|---|
| SWE-bench Pro | 62.1 | 58.4 | 69.2 | 58.6 | 54.2 |
| Terminal-Bench 2.1 (Terminus-2) | 81.0 | 63.5 | 85.0 | 84.0 | 74.0 |
| FrontierSWE | 74.4 | 30.5 | 75.1 | 72.6 | 39.6 |
| PostTrainBench | 34.3 | 20.1 | 37.2 | 28.4 | 21.6 |
| SWE-Marathon | 13.0 | 1.0 | 26.0 | 12.0 | 4.0 |
| Benchmark | GLM-5.2 | GLM-5.1 | Opus 4.8 | GPT-5.5 | Gemini 3.1 Pro |
|---|---|---|---|---|---|
| MCP-Atlas (Public Set) | 76.8 | 71.8 | 77.8 | 75.3 | 69.2 |
| Tool-Decathlon | 48.2 | 40.7 | 59.9 | 55.6 | 48.8 |
* Scores from full set evaluation. Source: Z.ai — GLM-5.2: Built for Long-Horizon Tasks
Get started in minutes
Install globally with npm. Works on macOS, Linux, and Windows.
Create your account — takes less than a minute.
Use /model to pick GLM-5.2 from the model selector. Use GLM-5.2[1m] for 1M context.
Run GLM-5.2 in AdaL CLI with 1M-token context, flexible effort levels, and production-ready coding workflows.
Get Started with AdaL