DeNovoSWE: The First Long-Horizon Doc2Repo Training Set That Teaches Code Agents to Build Repositories

·Toolin Editorial Team

Renmin University of China releases an open training set of 4818 real task instances targeting repository-level code generation, lifting Qwen3-30B's pass rate from 5.8% to 47.2%.

DeNovoSWE: The First Long-Horizon Doc2Repo Training Set That Teaches Code Agents to Build Repositories

Over the past year, benchmarks like SWE-bench have taught Code Agents to "fix bugs" and "patch conditional logic," but real software development is usually about understanding requirements, planning architecture, creating files, designing APIs, handling dependencies, wiring modules together, and finally getting an entire repository to pass its tests. On this kind of long-horizon repository-level generation task, frontier models still perform poorly on BeyondSWE-Doc2Repo and NL2RepoBench. DeNovoSWE, a dataset recently released by the Gaoling School of Artificial Intelligence at Renmin University of China, targets exactly this gap — it is the first open, high-quality training set for long-horizon Doc2Repo tasks, containing 4818 real task instances and lifting Qwen3-30B-A3B-Instruct's pass rate on BeyondSWE-Doc2Repo from 5.8% to 47.2%. This article breaks down its design philosophy, method, and usage, as a resource reference for teams training their own Code Agents.

What Is DeNovoSWE

One-line definition: an open training set that teaches Code Agents to generate complete, runnable repositories from documentation.

  • Publisher: Gaoling School of Artificial Intelligence, Renmin University of China
  • Task type: document-to-repository generation (long-horizon SWE tasks)
  • Scale: 4818 high-quality real task instances
  • Properties: executable, evaluable, trainable

Image

Paper at arxiv.org/abs/2606.10728, code at github.com/AweAI-Team/DeNovoSWE.

The Core Problem: From "Fixing Bugs" to "Building Repositories"

DeNovoSWE pushes task difficulty through a fundamental shift: no longer issue-level fixing, but whole-repository generation.

In a traditional SWE task, the Agent faces an existing repository and only needs to locate the bug, modify local code, and pass the tests. In DeNovoSWE, the Agent faces a scrubbed environment:

  • Original source code and tests are removed
  • Git history is reset
  • Every potential leakage channel is wiped: caches, leftover site-packages, pip wheels, temporary build artifacts

This means the Agent must genuinely rely on the documentation to rebuild the entire repository: plan the project structure, create module files, define public interfaces, implement cross-file interactions, handle dependencies and configuration, and keep fixing errors across rounds of editing and test feedback. Any deviation in an API signature, return field, exception type, or default behavior can fail the tests — and errors accumulate over the long horizon, since a poorly designed early module affects multiple downstream files and call chains.

Two Standards for High-Quality Task Documents

DeNovoSWE's core idea: documents must be readable, implementable, and verifiable at the same time. In document-to-repository generation, the document is not just a README — it is the Agent's sole task entry point for rebuilding the whole repository.

Standard 1: well-organized

Repository-level tasks are inherently complex, spanning multiple modules, interfaces, configurations, data structures, and interaction flows. If the document is just a heap of function descriptions, the Agent gets lost in fragmented information. A high-quality document should:

  • Open with a clear repository overview
  • Split into sections by capability or workflow
  • Map each part to a well-defined functional boundary

Standard 2: start from reliable evaluation

The document can't be too sparse (the task becomes under-defined, and the model passes evaluation only by guessing wildly) or too dense (it leaks implementation details and drains the task of challenge). A genuinely high-quality document should describe the key behaviors that evaluation depends on:

  • import paths
  • public APIs
  • inputs and outputs
  • default parameters
  • exception behavior
  • configuration options
  • pattern strings
  • return fields

In other words, the document must give the Agent enough to reproduce testable behavior, without becoming a copy of the implementation code.

Image

DeNovoSWE builds its high-quality dataset through Divide & Conquer and Critic & Repair mechanisms.

The Method: Divide & Conquer

DeNovoSWE does not hand-write its documents. It auto-builds high-quality instances through a sandboxed multi-agent workflow, in two stages.

The Divide Stage

The system analyzes the target repository and decomposes it into multiple repository capabilities, where each capability corresponds to one core ability or workflow in the repo (authentication and connection, data read/write, batch processing, export flows, and so on). At the same time it:

  • Runs the original unit tests and collects execution traces
  • Identifies which functions, classes, and interfaces actually affect evaluation
  • Separates components into three classes:
    • direct components: interfaces invoked directly by tests, which must be documented in detail
    • core indirect components: core indirect components that affect observable behavior, which must be covered
    • non-core indirect components: non-essential internal implementation, left for the Agent to decide freely

The Conquer Stage

A Draft-Critic-Repair mechanism generates documentation capability by capability:

  1. Draft agent: writes the first draft
  2. Critic agent: checks the document for missing key APIs, behavioral contracts, or structural information
  3. Repair agent: fixes the document based on the feedback
  4. Iterate in a loop until each capability section is clear, complete, and aligned with the evaluation

Finally, the per-capability documents are merged into one complete task document — the sole basis from which the Agent generates the repository from scratch.

Image

The Draft-Critic-Repair loop keeps documents strictly aligned with evaluation, avoiding both leakage and gaps.

Difficulty-Aware Filtering

To handle difficulty differences across repositories, DeNovoSWE introduces difficulty-aware trajectory filtering:

  • Easy tasks: demand a higher pass rate
  • Hard tasks: are not discarded wholesale just for falling short of a perfect score

Different filtering thresholds are set for different difficulty bands, based on structural complexity and LLM-judged difficulty. This matters especially for long-horizon tasks — the more complex a repository, the harder it is to pass all tests in one shot, yet the hard, low-scoring, partially successful trajectories still carry valuable long-horizon planning and implementation ability.

Experimental Results: Long-Horizon Data Is a Step Change

DeNovoSWE ultimately built 4818 high-quality document-to-repository task instances, and the experiments demonstrate how irreplaceable "long-horizon data" is.

Gains on Qwen3-30B-A3B-Instruct

DatasetBeyondSWE-Doc2RepoNL2RepoBench
Base model5.8%4.3%
Scale-SWE-Agent (trained on ordinary SWE data)29.2%18.3%
Trained on DeNovoSWE47.2%23.0%

Ordinary issue-level SWE data does transfer, but DeNovoSWE pushes the gains a big step further. The takeaway: data aimed at "fixing bugs" cannot fully replace long-horizon data aimed at "generating complete repositories."

Image

Experiments show DeNovoSWE delivers a significant boost to long-horizon repository generation.

Consistent Gains on a Stronger Backbone

On Qwen3.5-35B-A3B, DeNovoSWE again delivers steady gains:

  • BeyondSWE-Doc2Repo: 43.8% -> 50.0%
  • NL2RepoBench: 23.5% -> 27.1%

Further evidence that the benefit isn't an accident of one particular model but comes from the high-quality long-horizon data itself.

How to Use It

Resources

  • Paper: https://arxiv.org/pdf/2606.10728
  • Code repository: https://github.com/AweAI-Team/DeNovoSWE
  • Dataset (HuggingFace): https://huggingface.co/collections/AweAI-Team/denovoswe
  • Training Code Agents: add it as long-horizon SWE data to your existing training pipeline, complementing issue-level data like Scale-SWE
  • Evaluating models: benchmark your own model's repository-level generation on BeyondSWE-Doc2Repo and NL2RepoBench
  • Researching long-horizon tasks: the difficulty-aware trajectory filtering approach transfers to other long-horizon Agent tasks

Who It's For

  • LLM teams: training the next generation of Code Agents with long-horizon software engineering ability
  • Agent framework developers: using DeNovoSWE to benchmark how your framework performs on repository-level tasks
  • Academic research: a methodological reference for long-horizon tasks, verifiable tasks, and anti-leakage data construction

💡 Tip: if your Code Agent does fine on SWE-bench but falls apart when asked to "build a repository from scratch," DeNovoSWE is the most directly usable training data available right now. Start with a small-scale SFT run on its difficulty-aware filtered subset to validate, then decide whether to scale up.

Related articles

A Guide to Codex's Three Computer-Control Modes
AI Tutorials

A Guide to Codex's Three Computer-Control Modes

Codex's three control modes — Computer Use, the Chrome extension, and the in-app browser — each fit different scenarios. This piece breaks down the permission hierarchy and offers best practices.

Toolin Editorial Team
Codex Open-Source Mode: Plug In Local Models with One Line of Config
AI Products

Codex Open-Source Mode: Plug In Local Models with One Line of Config

OpenAI's Codex adds an OSS mode: a model_providers config connects Ollama, LM Studio, and other local model services, with switchable models to cut costs.

Toolin Editorial Team
Alibaba HappyHorse 1.1: Video Generation Upgraded Across Five Dimensions
AI Products

Alibaba HappyHorse 1.1: Video Generation Upgraded Across Five Dimensions

Alibaba's HappyHorse 1.1 video generation model improves five dimensions including motion expressiveness and subject consistency, cuts 1080P pricing by 25%, and is now live on the Bailian platform.

Toolin Editorial Team
MaineCoon: The Fastest Streaming Audio-Video Social Model Yet
AI Products

MaineCoon: The Fastest Streaming Audio-Video Social Model Yet

Catnip has unveiled MaineCoon, a 22B-parameter streaming audio-video model that hits 47.5 FPS on a single H100, runs at 1/2000th the cost of Veo 3, and supports 30+ minutes of synchronized audio-video output.

Toolin Editorial Team
Seko Infinite Canvas: From One Idea to a Wuxia Epic
AI Tutorials

Seko Infinite Canvas: From One Idea to a Wuxia Epic

Seko runs Seedance 2.0's all-in-one mode and uses Agent workflows to auto-generate plot, characters, and storyboards — 720P costs drop by 50%, with a finished video in about 10 minutes.

Toolin Editorial Team
Claude Science: A Claude Code for Scientific Research, Plus a Free Open-Source Alternative
AI Products

Claude Science: A Claude Code for Scientific Research, Plus a Free Open-Source Alternative

Anthropic has launched Claude Science, an AI workbench for research with 60+ built-in skills and fully reproducible outputs. There's also an open-source alternative, OpenScience, which supports DeepSeek/GLM.

Toolin Editorial Team