GLM-5.2's Million-Token Context, Tested: An 85-Page World Cup Preview in One Click

·Toolin Editorial Team

Zhipu's GLM-5.2 supports a 1 million token context window and will be open-sourced under MIT next week; in testing it produced an 85-page World Cup preview deck, with multi-agent parallelism that beat expectations

GLM-5.2's Million-Token Context, Tested: An 85-Page World Cup Preview in One Click

GLM-5.2 is Zhipu's newly released large language model, supporting a 1 million token context window and being open-sourced under the MIT license. Inside the Claude Code framework it shows outstanding long-context handling, multi-agent parallel scheduling, and autonomous fact-checking. This article uses a hands-on project — an 85-page World Cup preview deck — to show how GLM-5.2 really performs on complex engineering work.

Key Highlights

GLM-5.2's key specs:

  • Context window: 1 million tokens (1M)
  • License: MIT, no geographic restrictions
  • Open-source release: officially next week
  • Positioning: Zhipu's strongest model to date, with a marked boost in engineering capability

The Project: An 85-Page World Cup Preview Deck

Task Definition

The goal is to turn every group-stage match of the 2026 World Cup into a complete preview deck, with these requirements:

  • One page per matchup with national flags, kickoff time, venue, key players, core insights, and score predictions
  • Matches already played use real results
  • Plus an outlook page for each group and an overall cover

Final output: 1 cover + 12 group previews + 72 matches = 85 pages.

The First Hurdle: Autonomous Error Correction

Early in the task, GLM-5.2 ran into a classic trap: past World Cups had 32 teams and 48 group-stage matches, but the 2026 World Cup has been restructured to 48 teams in 12 groups — actually 72 matches.

It didn't just run with the wrong number; it stopped on its own initiative, cross-checked multiple sources including FIFA's official site and ESPN, and confirmed the correct figure of 72.

This ability to "know it might be misremembering" is a key indicator for judging a model's intelligence.

The Multi-Agent Parallel Architecture

Faced with 85 pages, GLM-5.2 didn't process them one by one in sequence; it automatically designed a five-layer pipeline:

  1. Single data source: everything starts from one data source to keep the data consistent
  2. 12 subagents in parallel: each researches one of the 12 groups and produces structured content
  3. Unified template rendering: the HTML is rendered from a single template to prevent style drift
  4. Batch aggregation: all pages are aggregated into an overview wall
  5. Visual verification: a vision model is called in to check every page for overflow, clipping, and similar issues

The key design decision: subagents only produce structured content and never write HTML directly. That's the only way all 85 pages stay locked into a single visual system.

The Output

The full 85 pages landed in about an hour, with two style options to choose from:

  • A broadcast-special style: deep charcoal background with gold scorelines
  • A magazine-editorial style: paper white with large serif type

Across 72 matches the style stays unified, the information hierarchy is clear, and the key points, predictions, and data are all there.

A World Cup preview page generated by GLM-5.2

The information hierarchy is clear, and the style stays unified across all 72 matches.

As a text-only model, GLM-5.2 handles multimodal verification by calling a vision model, autonomously checking pages for overflow and clipping.

GLM-5.2's multimodal verification workflow

Engineering Capability Tests

Beyond the World Cup project, GLM-5.2 also passed a series of engineering capability tests:

Dynamic moon-phase clock: about 925 lines of pure front-end code with zero external dependencies, producing five layers of concentric SVG, seven gears, 60 minute markers, elliptical star trails, and a moon-phase dial. After finding a bug, it voluntarily tore the whole thing down and rewrote it, refusing to pile up tech debt.

3D penalty shootout: built with Three.js + Cannon.js, featuring five rounds of attack and defense, three AI difficulty levels, and Magnus-effect curve physics. It even consulted papers to get real biomechanical parameters for goalkeeper dives.

Mini Excel: spent an hour recreating the core desktop Excel experience in a pure browser environment, including a formula engine, 30+ functions, and 60-step undo/redo.

The mechanical astronomical clock generated by GLM-5.2

After discovering the moon-phase bug, GLM-5.2 rewrote it from scratch, replacing the mask approach with a double-arc path.

Verification of the paper data sources GLM-5.2 consulted

Every cited data source was verified — no fabricated content.

Strengths and Weaknesses

Strengths

  • The 1M context genuinely delivers: long project specs and skills are still honored deep into the window — no "forgetting the beginning halfway through"
  • Autonomous error correction: it stops to verify on its own when it hits a knowledge blind spot
  • Multi-agent scheduling: large projects are split up and parallelized automatically
  • Unrestricted open source: MIT license, no geographic restrictions, and the API can't be cut off without warning

Weaknesses

  • Aesthetics need work: the visual appeal of generated interfaces has room to improve
  • Occasionally uneven pacing: on complex tasks it sometimes thinks for a long time and is slow to output runnable code
  • Occasionally seems stuck: you need to manually type a continue instruction

Recommendation

If you have a large project to complete, GLM-5.2 + the Claude Code framework is a very solid choice. In actual use, native Claude Code and Claude Code wired to GLM-5.2 have become hard to tell apart in output quality and interaction experience.

FAQ

  • When will GLM-5.2 be open-sourced? Officially next week, under the MIT license.
  • Does it support multiple languages? Yes — it performs well in both Chinese and English scenarios.
  • How do I hook it into Claude Code? Just wire GLM-5.2 in as the backend model of the Claude Code framework.

Related articles

AutoMIA: Two Images In, a 3D-Printable Mirror Illusion Out — Runs on a Single RTX 3090
AI Products

AutoMIA: Two Images In, a 3D-Printable Mirror Illusion Out — Runs on a Single RTX 3090

AutoMIA, a Tsinghua CVPR'26 Highlight open-source project, turns two input images into a 3D-printable voxel mirror-illusion art model (STL-ready) — single RTX 3090, 76 seconds per design.

Toolin Editorial Team
Macaron-V1: An Open-Source Coding Model That Compresses Personalized Experience into LoRAs with MoL
AI Products

Macaron-V1: An Open-Source Coding Model That Compresses Personalized Experience into LoRAs with MoL

Mind Lab open-sources Macaron-V1 (Venti / Coding-Venti / Tall tiers), a Mixture-of-LoRA architecture built on GLM-5.2 744B plus four 1B LoRA experts, with a free hosted API and a UI4A plugin.

Toolin Editorial Team
UniWorld-View: Single Image/Video to Any Camera Trajectory, Tops the WorldScore Leaderboard, Fully Open-Sourced
AI Products

UniWorld-View: Single Image/Video to Any Camera Trajectory, Tops the WorldScore Leaderboard, Fully Open-Sourced

UniWorld-View, open-sourced by Peking University + Rabbitpre + Pengcheng Laboratory, generates novel-view videos along any camera trajectory from a single image/video — it tops Fei-Fei Li's team's WorldScore leaderboard, with code and weights fully open under Apache-2.0 and one-click downloads.

Toolin Editorial Team
Claude Managed Agents Ships Six Updates: Skill Cap Raised to 500, With Ready-to-Run Payloads
AI Products

Claude Managed Agents Ships Six Updates: Skill Cap Raised to 500, With Ready-to-Run Payloads

Anthropic's managed agent platform CMA ships six updates at once: per-session skills up from 20 to 500, a five-level effort setting writable into per-agent config, and seeded sessions created in one step with initial_events.

Toolin Editorial Team
DojoAgents: Build a Financial Research Agent Locally in 10 Minutes
AI Tutorials

DojoAgents: Build a Financial Research Agent Locally in 10 Minutes

An open-source agent framework from Shenchong Intelligence covering A-shares, US stocks, and Hong Kong stocks — deploy a financial research agent that autonomously analyzes market themes, locally, in 10 minutes.

Toolin Editorial Team
OpenFPM's Experimental Metal Backend: Running CUDA Programs on Apple Silicon
AI Products

OpenFPM's Experimental Metal Backend: Running CUDA Programs on Apple Silicon

An experimental Metal backend PR for the open-source scientific computing framework OpenFPM runs original CUDA kernels nearly unchanged on an M3 Pro via Clang/HIP-SPIR-V-Vulkan-MoltenVK, with a measured ~10x speedup on a 3D SPH benchmark.

Toolin Editorial Team