DeNovoSWE: The First Long-Horizon Doc2Repo Training Set, Teaching Code Agents to Build Repositories
Renmin University's Gaoling School has released the DeNovoSWE dataset — 4818 real task instances that train Code Agents to generate complete repositories from documentation, lifting Qwen3-30B from 5.8% to 47.2% on BeyondSWE-Doc2Repo.


DeNovoSWE: The First Long-Horizon Doc2Repo Training Set, Teaching Code Agents to Build Repositories
Renmin University's Gaoling School has released the DeNovoSWE dataset — 4818 real task instances that train Code Agents to generate complete repositories from documentation, lifting Qwen3-30B from 5.8% to 47.2% on BeyondSWE-Doc2Repo.
DeNovoSWE is a long-horizon software engineering dataset released by the Gaoling School of Artificial Intelligence at Renmin University of China, built specifically to train Code Agents to start from a single document and produce a complete, executable, verifiable software repository. It contains 4818 real task instances and is open-sourced for both training and evaluation.
This is one 2026 open-source resource worth every Code Agent developer's attention. If you are training or fine-tuning code agents and want to push their abilities from "fixing bugs" to "building repositories," DeNovoSWE offers a reusable path.
What problem it solves
Over the past year, Code Agents have improved rapidly on real software engineering tasks like SWE-bench. But as models get better and better at "resolving an issue, patching a few lines," a key question has come into focus: do agents actually possess long-horizon software engineering capability?
Real-world software development is rarely fixing one function or adding one conditional. It looks more like: understand requirements → plan the architecture → create files → design APIs → handle dependencies → wire up modules → get the entire repository passing tests. In other words, the hard part is long-horizon repository-level generation. Judging by how frontier models perform on BeyondSWE-Doc2Repo and NL2RepoBench, the results are far from ideal.

Core mechanism: Divide & Conquer + Critic & Repair
DeNovoSWE does not hand-write its documents. Instead, a sandboxed multi-agent workflow automatically constructs high-quality instances, in a process that boils down to two steps.
Divide stage: the system analyzes the target repository and decomposes it into multiple repository capabilities (such as authentication and connections, data read/write, batch processing, and export flows). At the same time it runs the original unit tests to collect execution traces, distinguishing interfaces invoked directly by tests (must be documented in detail), core indirect components that affect observable behavior (need coverage), and non-core internal implementations (left for the agent to improvise).
Conquer stage: a Draft-Critic-Repair mechanism generates documentation capability by capability — a Draft agent writes the first version, a Critic agent checks for missing key APIs or behavioral contracts, and a Repair agent fixes issues based on the feedback. The loop iterates until every capability section is clear, complete, and aligned with the evaluation.

💡 Tip: The core idea behind DeNovoSWE is documents that are readable, implementable, and verifiable all at once — describing the key behaviors the evaluation depends on (import paths, public APIs, inputs and outputs, default parameters, exception behavior, configuration options, pattern strings, return fields, and so on) without turning into a copy of the implementation code.
Leakage prevention and difficulty filtering
To make agents genuinely rely on the document rather than reproducing code from "memory," DeNovoSWE performs strict cleanup in the task environment: original source code and tests are removed, git history is reset, and potential leakage channels — caches, leftover site-packages, pip wheels, temporary build artifacts — are all wiped.
It also introduces difficulty-aware trajectory filtering: easy tasks demand higher pass rates, while hard tasks are not discarded wholesale just for failing to hit a perfect score. This matters especially for long-horizon tasks — the more complex the repository, the harder it is to pass every test in one shot, yet trajectories on hard repos with low scores and partial success still carry valuable long-horizon planning and implementation ability.
Experimental results
DeNovoSWE ultimately contains 4818 high-quality document-to-repository task instances. Experiments show it delivers significant gains in models' long-horizon repository generation capability:
| Model / data | BeyondSWE-Doc2Repo | NL2RepoBench |
|---|---|---|
| Qwen3-30B-A3B-Instruct (original) | 5.8% | 4.3% |
| + Scale-SWE-Agent (issue-level data) | 29.2% | 18.3% |
| + DeNovoSWE | 47.2% | 23.0% |
On the stronger Qwen3.5-35B-A3B backbone, DeNovoSWE likewise delivers consistent gains: BeyondSWE-Doc2Repo rises from 43.8% to 50.0%, and NL2RepoBench from 23.5% to 27.1%. That indicates the benefit comes not from accidental alignment with one particular model, but from the quality of the long-horizon data itself.
💡 Tip: Data aimed at "fixing bugs" cannot fully substitute for long-horizon data aimed at "generating complete repositories." If you want agents to truly learn repository-level engineering, you need training environments purpose-built for long-horizon tasks.
Resources
- Paper: https://arxiv.org/pdf/2606.10728
- Code repository: https://github.com/AweAI-Team/DeNovoSWE
- Dataset: https://huggingface.co/collections/AweAI-Team/denovoswe
Who it is for
- Researchers and engineers currently training or fine-tuning Code Agents
- Teams focused on long-horizon benchmarks
- Practitioners looking to upgrade their code agents from "repository maintainers" to "architects"
Related articles

StepFun's Step Edge On-Device Family: 0.1-Second Local Agent Tool Calls
StepFun releases the Step Edge on-device model family — base/Audio/GUI/Gen — aimed at phones and cars, with local tool calls as fast as 0.1 seconds and sensitive data never leaving the device.

Google Open-Sources Gemma 4: Four Sizes, 256K Context, Runs Locally
Google releases the Gemma 4 open model family (E2B/E4B/26B MoE/31B Dense) under Apache 2.0, with 256K context, native multimodality, and local deployment channels.

Grok 4.5: xAI's Coding Model Co-Trained with Cursor, Starting at $2
xAI releases Grok 4.5, co-trained with Cursor and built for real engineering tasks. API pricing is $2/$6 per million tokens, and it's live on all Cursor plans.

Meta Muse Spark 1.1: An Agent Model with a Million-Token Context, Now in API Public Beta
Meta has released Muse Spark 1.1, a multimodal reasoning model built for agent tasks with a 1 million-token context window, alongside a public preview of the Meta Model API.

GPT-5.6 Is Here: Three New Model Tiers + ChatGPT Work, Codex Merged into the Desktop App
OpenAI has launched the GPT-5.6 family (Sol/Terra/Luna) and the ChatGPT Work agent, and merged Codex into the ChatGPT desktop app — with full pricing and availability details.

ChatGPT's Big Voice Upgrade: GPT-Live Full-Duplex Mode Goes Live
OpenAI has released the GPT-Live full-duplex voice model family. ChatGPT can finally talk while listening, chime in naturally, and know when to stay silent, replacing the original Advanced Voice Mode.