GLM-5.3
Z AI · 2 thinking configurations
Intelligence44.8
Input / 1M$1.40
Output / 1M$4.40
Context1M
Tariff: Z AI · Source: artificial-analysis · USD per million tokens
Thinking configuration
What does more thinking buy you?
Compare this model’s measured thinking configurations.
Scroll the chart ↔ · tap a point for details
Efficient frontier: best measured score at each cost2 measured configurations · 0 without both values
Selected: thinking max · Click a point to switch configuration
GLM-5.3 is a large-scale reasoning model from Z.ai, built for complex software engineering and long-horizon agent tasks. It supports text input and output with a 1M-token context window, and improves...
Intelligence IndexArtificial Analysis composite index, not a percentage. Different versions of the index are not directly comparable.
44.8 pointsDeepSWE · pass@1Software-engineering tasks on mini-swe-agent v1.1. Pass@1 across evaluated runs, with measured cost per task. The 95% intervals may overlap.
69.0 % ±3.0GPQA DiamondGraduate-level scientific questions. Measures specialist reasoning rather than web-search quality.
91.7 %Humanity’s Last ExamDifficult questions across many disciplines, compared under the same evaluation setup.
42.3 %SciCodeCode generation for scientific research problems.
59.0 %Terminal-Bench 2.1Agent tasks in a terminal environment. Results also depend on the evaluation harness.
83.9 %Omniscience accuracyFactual accuracy in Artificial Analysis's Omniscience evaluation.
33.9 %Specifications from the catalog
Inputtext
Outputtext
Tool callingYes
Structured outputYes
Speed82.7 tok/s
First token2.85 s
Maximum output943.7K
Catalog listing date2026-08-18
DeepSWE v1.1 · 113 tasks · mini-swe-agent · cost $3.99 / task · 124 mean steps · 80.4K mean output tokens. Published by the source 9/22/2026.