Z.ai’s GLM-5.2 has become the highest-scoring open-weight model on version 4.1 of the Artificial Analysis Intelligence Index, reaching 51 points. That places it seven points above MiniMax-M3 and DeepSeek V4 Pro at their reported 44, and eight above Kimi K2.6 at 43. The ranking reflects Artificial Analysis’s evaluation suite and should not be read as a universal measure of model quality.
GLM-5.2 uses the same mixture-of-experts scale reported for GLM-5.1: 744 billion total parameters, with 40 billion active during inference. Its index score rises by 11 points from the earlier model. Artificial Analysis reports gains across most components, including 16 percentage points on both CritPt scientific reasoning and TerminalBench v2.1, 15 points on a banking task in tau3, 12 points on Humanity’s Last Exam and nine points on AA-LCR.
On GDPval-AA v2, the service’s primary test of real-world agentic performance, GLM-5.2 scored 1,524. That exceeded MiniMax-M3 at 1,418 and DeepSeek V4 Pro at 1,328, while sitting near GPT-5.5 with high reasoning at 1,514. The benchmark uses human performance as a 1,000 Elo baseline, rotates frontier-model judges and permits trajectories of as many as 250 turns. Those design choices shape what the score captures.
The improvement comes with substantial output. Artificial Analysis measured an average of 43,000 output tokens per Intelligence Index task, including about 37,000 reasoning tokens. GLM-5.1 used 26,000, MiniMax-M3 24,000 and Kimi K2.6 35,000. GLM-5.2 consequently sits outside the most attractive region of the service’s intelligence-versus-output-token chart, even while appearing on its intelligence-versus-cost Pareto frontier.
Average cost was approximately $0.46 per index task. That was above GLM-5.1 at $0.25, Kimi K2.6 at $0.31, MiniMax-M3 at $0.18 and DeepSeek V4 Pro at $0.05. Artificial Analysis nevertheless found no cheaper model matching GLM-5.2’s intelligence score. Z.ai’s first-party API pricing was listed at $1.40 per million input tokens, $4.40 per million output tokens and $0.26 per million cache-hit tokens.
The model also expands context from 200,000 tokens in GLM-5.1 to one million. The service measured its attempt rate as unchanged at 47%, alongside accuracy of 25.1%, up from 24.2%. It is available from Z.ai and several third-party providers. On the AA-Omniscience measure, it improved modestly through slightly higher accuracy and a hallucination rate of 28.1%, down from 29.4%.
The combined results present a trade-off rather than an unqualified win. GLM-5.2 leads the selected open-weight capability index and offers a much longer context window, but it produces more tokens and costs more per evaluated task than several peers. Deployment decisions will depend on whether the measured capability gains justify that additional inference burden.



