STAR-PólyaMath: The Open-Source Reasoning Framework That Teaches LLMs to Correct Their Mistakes

·Toolin Editorial Team

A multi-agent reasoning framework open-sourced jointly by Tsinghua and Microsoft, where the Reasoner, Verifier, and Meta-Strategist roles make long-horizon reasoning verifiable and traceable — beating GPT-5.5 on the Apex benchmark by 13.5%.

STAR-PólyaMath: The Open-Source Reasoning Framework That Teaches LLMs to Correct Their Mistakes

On the hardest math competition problems, even the strongest LLMs keep spinning their wheels on the same mistake — they pile details onto a seemingly plausible route and write quite persuasive arguments, but lack the metacognition to realize "this is not a solution that needs more polishing; it is a dead end." STAR-PólyaMath, open-sourced jointly by Tsinghua University and Microsoft Research Asia, solves this problem with a multi-agent framework: a loop driven by three roles — Reasoner, Verifier, and Meta-Strategist — makes the reasoning process verifiable, traceable, and able to accumulate experience across attempts. It takes the top score on all eight elite math competition benchmarks, leading same-backbone GPT-5.5 by 13.5% on the Apex benchmark.

STAR-PólyaMath takes the top score on all eight math competition benchmarks

Project Background

  • Paper title: STAR-PólyaMath: Multi-Agent Reasoning under Persistent Meta-Strategic Supervision
  • Paper link: arxiv.org/abs/2605.19338v1
  • Open-source link: github.com/Julius-Woo/STAR-PolyaMath
  • First author: Wu Jia'ao (PhD student, T-STAR Lab, School of Artificial Intelligence, Tsinghua University)
  • Corresponding authors: Zhang Xian (Principal Researcher, Microsoft Research Asia), Dong Yinpeng (Assistant Professor, School of Artificial Intelligence, Tsinghua University)

The framework draws its inspiration from the problem-solving steps George Pólya laid out in How to Solve It — understand the problem, devise a plan, carry out the plan, look back — structured into four phases: exploration, planning and decomposition, step-by-step execution with challenge loops, and solution generation.

The Three Difficulties of Long-Horizon Reasoning

Frontier LLMs are already extremely strong at routine reasoning, but when facing long-horizon reasoning that demands multi-step exploration, hypothesis testing, and even tearing things down to start over, three systematic failure modes keep reappearing.

Hallucination accumulation and trustworthiness. Models tend to hold high confidence in their own intermediate conclusions; one seemingly tiny error (say, a missed boundary case) keeps amplifying through the subsequent derivation.

Memory loss across attempts. When a proof path fails and needs backtracking, most systems either retain so much context that the error cannot be located precisely, or lose key information from earlier attempts — and end up retrying directions that have already been falsified.

Imbalance between reasoning and tool use. Running code is a reliable verification method, but models trained on tool-use data systematically favor code over exploring mathematical structure. Conversely, pure natural-language reasoning struggles with problems that require symbolic construction.

The Core Architecture: Three Agent Roles

The whole framework is coordinated by a non-reasoning Python orchestrator (Orchestrator) managing three agent roles.

The Reasoner

It handles the actual problem-solving: explores the problem structure, proposes plans, executes each step of reasoning or computation, and defends its arguments when challenged. Its output must always pass through the verification stage.

Within a single attempt (that is, sequentially executing one plan), the Reasoner keeps full memory; but on backtracking and replanning its memory resets, to reduce contamination from erroneous reasoning.

The Verifier

It independently reviews the Reasoner's output and keeps no memory. The review has two gating mechanisms:

  • Goal Gate: checks whether the step genuinely accomplished the goal declared in the plan, preventing "semantic drift" (arguments that are correct but accomplished only a trivial solution).
  • Logic Gate: audits the correctness of the reasoning content.

After review it issues one of four verdicts: Accept, Challenge, Trace-Back, or Propose-Replan.

The Meta-Strategist

This is the framework's most critical innovation. It performs no concrete mathematical reasoning at all; instead it gives guidance at a higher level — like an experienced mentor. It maintains a single persistent session across the entire problem-solving process, accumulating all previous attempts, abandoned strategies, and long-standing failure modes.

STAR-PólyaMath's system workflow, driven by a loop of three agent roles

At critical moments, the Meta-Strategist issues concrete strategic advice, for example making the final ruling when the Verifier proposes replanning. When it detects the Reasoner stuck in successive rounds of meaningless computation, it can issue a mandatory instruction to switch to a pure-reasoning mode that forbids the use of code.

A Typical Case

MathArena Apex 2025 Problem 2 (from Turkey TST 2025 P5) is a typical example, and the correct answer is k = 1/2. GPT-5.5 at maximum thinking effort made 8 independent attempts at this problem and got just 1 right — it converged quickly on a suboptimal construction (arriving at the wrong answer 3/4) and worked hard to supply logically self-consistent arguments propping up the wrong conclusion.

On the same problem, STAR-PólyaMath's Reasoner also failed its first attempt (arriving at 3/4), but the Verifier kept challenging its proof. After three timed-out failures, the Meta-Strategist made a key judgment: "this direction is fundamentally wrong" — explicitly forbidding subsequent reasoning from re-anchoring on 3/4 and authorizing a replan. The new approach found a denser construction, pushed the result to 1/2, and completed a rigorous proof verified both by mathematical derivation and by actually running construction code.

💡 Tip: The backbone-swap experiments make it clear that the performance gains come from the structured reasoning harness framework, not the model itself. With the backbone swapped from GPT-5.5 to GPT-5.2 or Claude Opus 4.7, the framework still beats direct calls to the corresponding models on every benchmark. That means you can approach frontier results with cheaper models.

The Verification Mechanism: Layered Verification Labels

STAR-PólyaMath gives every step of long-horizon reasoning checkability through layered verification labels. Every intermediate assertion must be tagged as:

  • [verified]: validated by executed code
  • [easy-verify]: checkable by simple computation
  • [hard-verify]: requires rigorous mathematical review

This set of labels determines how hard the Verifier scrutinizes: code-verified results are taken as trustworthy directly, while pure mathematical arguments receive the strictest logical review.

The actual run statistics show how clearly adaptive this layering strategy is:

  • AIME, HMMT (computation-heavy): about 36-43% of assertions verified by code
  • IMO, Putnam (proof-heavy): over 85% of assertions are [hard-verify]

Experimental Results

STAR-PólyaMath uses GPT-5.5 (xhigh effort) as the backbone model for all three agents and takes the top score on all 8 elite math competition benchmarks.

STAR-PólyaMath reaches 93.75% on Apex 2025, leading same-backbone GPT-5.5 by 13.5%

Highlight numbers:

  • Apex 2025: 93.75%, versus just 80.21% for direct same-backbone GPT-5.5 calls — a 13.5% gap. These are the hardest problems requiring multi-step proofs and strategy switches, and the scenario where the Meta-Strategist delivers its maximum value.
  • AIME 2025/2026, Putnam 2025, HMMT 2026: perfect scores.

Compute cost correlates strongly with problem difficulty:

  • AIME level: solved in 8 minutes on average, with 100% solved directly in the exploration phase and the Meta-Strategist almost never triggered.
  • Apex 2025 and IMO 2025 level: 55+ minutes on average, with the Meta-Strategist stepping in 1.6-2.2 times per problem.

The framework does not impose unnecessary overhead on easy problems, but invests ample compute in exploration and reasoning on genuinely hard ones.

Beyond Math: A Generalizable Reasoning Paradigm

STAR-PólyaMath's design does not depend on anything special to the mathematics domain. Its core mechanisms (decomposing long-horizon tasks into verifiable sub-steps, structured checking of every step, memory across attempts, high-level supervision) inherently apply to any scenario that needs long-horizon, traceable, verifiable reasoning.

  • Code generation: a similar framework could structure the "generate-test-debug" loop as a state machine with backtracking, with the Meta-Strategist judging "the current architecture direction itself is wrong and needs a rewrite" after repeated failed patches.
  • Scientific discovery: the Reasoner proposes hypotheses and experiment designs, the Verifier reviews the experimental results, and the Meta-Strategist judges after multiple failed rounds whether "the experimental method or the underlying hypothesis should be revised."

The project has open-sourced the complete code framework, every role's prompt and skill definitions, and the run configurations, making it easy for the community to port this reasoning protocol to other domains.

💡 Usage advice: If your work involves complex reasoning that needs multi-step verification (mathematical proofs, algorithm design, research exploration), try it directly from the open-source code. Do not apply the full framework to simple tasks — AIME-level problems are solved 100% in the exploration phase, and the extra agent orchestration only adds cost.

Related articles

Baidu Open-Sources Unlimited OCR: a 500M-Active Small Model Reads 40 Pages in One Pass Without Forgetting
AI Products

Baidu Open-Sources Unlimited OCR: a 500M-Active Small Model Reads 40 Pages in One Pass Without Forgetting

Baidu open-sources Unlimited OCR, an end-to-end OCR model with 3B total / 500M active parameters that sets a new OmniDocBench SOTA and transcribes dozens of pages per inference without forgetting.

Toolin Editorial Team
Sakana Fugu: the Orchestrator That Doesn't Answer Itself, Just Directs Other Models to Do the Work
AI Products

Sakana Fugu: the Orchestrator That Doesn't Answer Itself, Just Directs Other Models to Do the Work

Sakana AI releases the Fugu family of orchestrator models, smartly dispatching GPT, Claude, and Gemini to finish tasks, with performance approaching Fable 5 and Mythos Preview.

Toolin Editorial Team
Seko Infinite Canvas Hands-On: Drop in One Idea, and the Agent Finishes a Whole AI Video for You
AI Tutorials

Seko Infinite Canvas Hands-On: Drop in One Idea, and the Agent Finishes a Whole AI Video for You

A hands-on guide to the Seko infinite canvas + Seedance 2.0: 720P cost down 50% and 1080P down 80%, letting even beginners produce multi-episode AI video epics in 10 minutes.

Toolin Editorial Team
WeChat Xiaowei Beta Hands-On: Swipe Right and Turn All of WeChat into an Agent
AI Products

WeChat Xiaowei Beta Hands-On: Swipe Right and Turn All of WeChat into an Agent

A beta look at WeChat's official AI assistant Xiaowei: chat summaries, auto-replies, invoking mini programs, money transfers, Moments browsing, and building small tools — eight capabilities in one read.

Toolin Editorial Team
The Ministry of Education's "Sunshine Volunteer" AI Assistant: Free College-Application Plans Generated in Seconds
AI Products

The Ministry of Education's "Sunshine Volunteer" AI Assistant: Free College-Application Plans Generated in Seconds

The Ministry of Education has officially upgraded its "Sunshine Volunteer" system, with the "Zhihui Xiaozhao" AI assistant online 24 hours a day, providing reach-match-safety application plans free of charge based on official data.

Toolin Editorial Team
GenShield: An All-in-One Open Source Framework for AI-Generated Image Detection and Repair
AI Products

GenShield: An All-in-One Open Source Framework for AI-Generated Image Detection and Repair

A Peking University team has open sourced GenShield, unifying AI-generated image detection and artifact correction in a single autoregressive framework with 98.8% detection accuracy

Toolin Editorial Team