Arbor Goes Open Source: Teaching Agents to Do Research Like Scientists, Not Blind Trial and Error

·Toolin Editorial Team

Renmin University's Gaoling School and Microsoft Research have open-sourced Arbor, an autonomous research framework that uses a Hypothesis-Tree to organize repeated experiments into an accumulable research state. It topped the HuggingFace daily chart, with held-out gains on long-horizon tasks averaging 2.5x those of Codex and Claude Code.

Arbor Goes Open Source: Teaching Agents to Do Research Like Scientists, Not Blind Trial and Error

Today's AI agents can write code, run experiments, and revise plans, but the vast majority cannot clear one core hurdle: they can execute, but they cannot do autonomous research. They try one direction and fail, try another and fail again, and even when they occasionally succeed they struggle to distill the mechanism behind that success. Researchers from Renmin University of China's Gaoling School of Artificial Intelligence and Microsoft Research have open-sourced Arbor — a general-purpose, practical autonomous research framework and toolkit that uses a "Hypothesis-Tree" to organize a series of short experiments into a research state that accumulates over the long term and can be audited. The paper has reached #1 on the HuggingFace Daily Papers chart.

This piece breaks down what problem Arbor solves, its core mechanism, real-world results, and how to get started.

What problem it solves

The paper formalizes autonomous research as Autonomous Optimization (AO): given an initial artifact (model training code, an agent harness, a data generation pipeline) plus a research objective and an executable evaluator, the agent improves the artifact through multiple rounds of experiments with no step-by-step human supervision, looking only at the dev set, and ultimately improves results on the test set.

Existing agents get stuck on three things:

  • Conversation history is long and scattered, and cannot carry structured research judgments
  • The working directory only records code changes, without explaining the hypotheses behind them
  • Logs preserve the results, but cannot tell the agent why something succeeded or failed

The result: the agent can run many trials, but those trials never add up to real research progress.

The figure above shows a hypothesis tree from a Math-Reasoning Data Synthesis run, along with the corresponding development score curve.

Core mechanism: the Hypothesis-Tree plus insight propagation

Arbor's key is not making the agent try more times, but changing how exploration is organized.

The Hypothesis-Tree

The research space is structured as a tree in which every node is a hypothesis to be tested. The agent runs isolated executions around hypotheses, compares evidence, and updates the search tree, rather than grinding away blindly along a single trajectory.

The insight propagation mechanism

This is the most crucial component. Failures are not discarded negative samples — they are research evidence that gets attributed, abstracted, and propagated; successes are not isolated score bumps but reusable local findings. These insights propagate up the tree and continuously shape the subsequent search distribution — the agent repeats failed paths less often and refines around effective mechanisms more easily.

Arbor framework overview

Arbor emphasizes two design points:

  • Generality: it is not tied to any specific benchmark or task format. As long as there is an artifact to optimize, a clear objective, and an executable feedback signal, models / harnesses / data can all be optimized.
  • Practicality: a standalone CLI and Agent SDK are open-sourced and can be plugged in directly.

Measured results

Across six real AO tasks (covering math reasoning data synthesis, SWE, Terminal-Bench, and more) plus MLE-Bench Lite:

  • On the six tasks, it delivered improvements averaging 2.5x the relative held-out gains of Codex and Claude Code
  • Terminal-Bench 2.0: held-out pass rate improved from an initial 69.81 to 77.36
  • MLE-Bench Lite: Arbor with GPT-5.5 reached 86.36% Any Medal, the current SOTA

Ablation studies further confirm the importance of insight propagation (on MLE-Bench Lite):

ConfigurationAny Medal
Full Arbor81.82%
Remove the hypothesis tree63.64%
Keep the tree, also remove upward insight propagation54.54%

One counterintuitive finding: removing only insight propagation hurts more than removing the whole tree outright. A tree that does not propagate experience can do no more than line its experiments up in a row, while offering none of the semantic memory that downstream decisions actually need.

💡 Tip: Arbor consumes tokens in the same order of magnitude as baselines like Claude Code, yet achieves larger held-out gains. That suggests the bottleneck is not how much compute you spend, but how that compute is organized.

How to get started

Arbor provides a standalone CLI and an Agent SDK. As long as you have an artifact to optimize, an executable evaluator, and a clear objective, you can plug in and start autonomous optimization.

Who it is for

  • AI researchers: use Arbor as the execution backbone for long-horizon autonomous experimentation, letting rounds of trial and error settle into reusable research memory.
  • Model / agent teams: use it to optimize training code, agent harnesses, and data pipelines, getting steadier held-out improvements than a bare coding agent run.
  • Long-horizon task developers: any scenario that calls for "continuously optimizing one research object" falls within Arbor's definition of AO.

Current limitations

The authors are candid: Arbor does not mean agents have reached human-researcher-level creativity. The quality of ideas agents generate still has plenty of room to improve; on hard tasks they may struggle to propose genuinely novel mechanisms, or give up on promising directions too early. How to generate higher-quality hypotheses, distinguish real improvements from chance overfitting, and maintain reliable memory over longer horizons all remain open questions.

References