Arbor Goes Open Source: Teaching Agents to Do Research Like Scientists, Not Blind Trial and Error
Renmin University's Gaoling School and Microsoft Research have open-sourced Arbor, an autonomous research framework that uses a Hypothesis-Tree to organize repeated experiments into an accumulable research state. It topped the HuggingFace daily chart, with held-out gains on long-horizon tasks averaging 2.5x those of Codex and Claude Code.


Arbor Goes Open Source: Teaching Agents to Do Research Like Scientists, Not Blind Trial and Error
Renmin University's Gaoling School and Microsoft Research have open-sourced Arbor, an autonomous research framework that uses a Hypothesis-Tree to organize repeated experiments into an accumulable research state. It topped the HuggingFace daily chart, with held-out gains on long-horizon tasks averaging 2.5x those of Codex and Claude Code.
Today's AI agents can write code, run experiments, and revise plans, but the vast majority cannot clear one core hurdle: they can execute, but they cannot do autonomous research. They try one direction and fail, try another and fail again, and even when they occasionally succeed they struggle to distill the mechanism behind that success. Researchers from Renmin University of China's Gaoling School of Artificial Intelligence and Microsoft Research have open-sourced Arbor — a general-purpose, practical autonomous research framework and toolkit that uses a "Hypothesis-Tree" to organize a series of short experiments into a research state that accumulates over the long term and can be audited. The paper has reached #1 on the HuggingFace Daily Papers chart.
This piece breaks down what problem Arbor solves, its core mechanism, real-world results, and how to get started.
What problem it solves
The paper formalizes autonomous research as Autonomous Optimization (AO): given an initial artifact (model training code, an agent harness, a data generation pipeline) plus a research objective and an executable evaluator, the agent improves the artifact through multiple rounds of experiments with no step-by-step human supervision, looking only at the dev set, and ultimately improves results on the test set.
Existing agents get stuck on three things:
- Conversation history is long and scattered, and cannot carry structured research judgments
- The working directory only records code changes, without explaining the hypotheses behind them
- Logs preserve the results, but cannot tell the agent why something succeeded or failed
The result: the agent can run many trials, but those trials never add up to real research progress.
The figure above shows a hypothesis tree from a Math-Reasoning Data Synthesis run, along with the corresponding development score curve.
Core mechanism: the Hypothesis-Tree plus insight propagation
Arbor's key is not making the agent try more times, but changing how exploration is organized.
The Hypothesis-Tree
The research space is structured as a tree in which every node is a hypothesis to be tested. The agent runs isolated executions around hypotheses, compares evidence, and updates the search tree, rather than grinding away blindly along a single trajectory.
The insight propagation mechanism
This is the most crucial component. Failures are not discarded negative samples — they are research evidence that gets attributed, abstracted, and propagated; successes are not isolated score bumps but reusable local findings. These insights propagate up the tree and continuously shape the subsequent search distribution — the agent repeats failed paths less often and refines around effective mechanisms more easily.

Arbor emphasizes two design points:
- Generality: it is not tied to any specific benchmark or task format. As long as there is an artifact to optimize, a clear objective, and an executable feedback signal, models / harnesses / data can all be optimized.
- Practicality: a standalone CLI and Agent SDK are open-sourced and can be plugged in directly.
Measured results
Across six real AO tasks (covering math reasoning data synthesis, SWE, Terminal-Bench, and more) plus MLE-Bench Lite:
- On the six tasks, it delivered improvements averaging 2.5x the relative held-out gains of Codex and Claude Code
- Terminal-Bench 2.0: held-out pass rate improved from an initial 69.81 to 77.36
- MLE-Bench Lite: Arbor with GPT-5.5 reached 86.36% Any Medal, the current SOTA
Ablation studies further confirm the importance of insight propagation (on MLE-Bench Lite):
| Configuration | Any Medal |
|---|---|
| Full Arbor | 81.82% |
| Remove the hypothesis tree | 63.64% |
| Keep the tree, also remove upward insight propagation | 54.54% |
One counterintuitive finding: removing only insight propagation hurts more than removing the whole tree outright. A tree that does not propagate experience can do no more than line its experiments up in a row, while offering none of the semantic memory that downstream decisions actually need.
💡 Tip: Arbor consumes tokens in the same order of magnitude as baselines like Claude Code, yet achieves larger held-out gains. That suggests the bottleneck is not how much compute you spend, but how that compute is organized.
How to get started
- Paper: https://arxiv.org/pdf/2606.11926
- Code repository: https://github.com/RUC-NLPIR/Arbor
- Project homepage: https://ruc-nlpir.github.io/Arbor/
Arbor provides a standalone CLI and an Agent SDK. As long as you have an artifact to optimize, an executable evaluator, and a clear objective, you can plug in and start autonomous optimization.
Who it is for
- AI researchers: use Arbor as the execution backbone for long-horizon autonomous experimentation, letting rounds of trial and error settle into reusable research memory.
- Model / agent teams: use it to optimize training code, agent harnesses, and data pipelines, getting steadier held-out improvements than a bare coding agent run.
- Long-horizon task developers: any scenario that calls for "continuously optimizing one research object" falls within Arbor's definition of AO.
Current limitations
The authors are candid: Arbor does not mean agents have reached human-researcher-level creativity. The quality of ideas agents generate still has plenty of room to improve; on hard tasks they may struggle to propose genuinely novel mechanisms, or give up on promising directions too early. How to generate higher-quality hypotheses, distinguish real improvements from chance overfitting, and maintain reliable memory over longer horizons all remain open questions.
References
- Paper: Toward Generalist Autonomous Research via Hypothesis-Tree Refinement (arXiv:2606.11926)
- GitHub: https://github.com/RUC-NLPIR/Arbor