- Chinese startup Z.ai says its GLM-5.2 model outperformed GPT-5.5 on key reasoning and coding benchmarks, highlighting rapid advances in AI and growing competition between Chinese and Western developers.
- The model scored higher on tests like Humanity’s Last Exam and FrontierSWE, with a massive context window and focus on complex engineering tasks, signaling a shift toward autonomous AI systems.
- On Humanity’s Last Exam — a rigorous benchmark featuring thousands of expert-level questions across disciplines — the model scored 54.7, compared to GPT-5.5’s 52.2.
- The model also showed gains on FrontierSWE, a benchmark that measures an AI system’s ability to complete complex, open-ended software engineering projects.
- Here too, GLM-5.2 reportedly outperformed GPT-5.5 by a small margin, reinforcing its strength in coding and engineering workflows.
- One of GLM-5.2’s standout features is its ability to process extremely large volumes of data at once.
- The model supports a context window of up to one million tokens, enabling it to work with large codebases, long documents and extended conversations without losing track of context.
- Z.ai says the model has been trained to handle full-scale development workflows, including analysing requirements, writing code, debugging and managing complex engineering processes — all within a single continuous task.
- The launch highlights intensifying competition between Chinese and American AI firms, particularly as open-source and hybrid models begin to rival proprietary systems in performance.
- Analysts note that China’s rapid advancements in AI development are narrowing the gap with global leaders, especially in areas like coding and applied engineering.
Chinese startup Z.ai has announced that its GLM-5.2 model outperforms GPT-5.5 on several key benchmarks, including a score of 54.7 on the rigorous Humanity’s Last Exam, which features thousands of expert-level questions across various disciplines. This performance highlights the rapid advancements in AI and the intensifying competition between Chinese and American AI firms.1239
The GLM-5.2 model also excelled on the FrontierSWE benchmark, which assesses an AI's capability to handle complex software engineering tasks. Here, it reportedly outperformed GPT-5.5 by a small margin, showcasing its strengths in coding and engineering workflows.45
One of the standout features of GLM-5.2 is its ability to process extremely large volumes of data simultaneously, supporting a context window of up to one million tokens. This allows it to manage large codebases, lengthy documents, and extended conversations without losing context.7
The launch of GLM-5.2 underscores the growing competition in the AI landscape, particularly as open-source and hybrid models begin to rival proprietary systems. Analysts note that China’s rapid advancements in AI are narrowing the gap with global leaders, especially in coding and applied engineering. However, experts caution that benchmark performance alone may not fully capture real-world reliability or usability, emphasizing the need for independent validation and broader deployment to determine if GLM-5.2 can truly challenge established models.10
“Z.ai's GLM-5.2 model reportedly outperformed GPT-5.5 on key reasoning and coding benchmarks, signaling rapid advancements in AI. This development highlights the intensifying competition between Chinese and Western AI developers.”
