DeNovoSWE: The First Long-Horizon Doc2Repo Training Set, Teaching Code Agents to Build Repositories

·Toolin Editorial Team

The Gaoling School of AI at Renmin University of China has released DeNovoSWE, the first long-horizon training set for generating complete repositories from documents, with 4818 real task instances; Qwen3-30B improved from 5.8% to 47.2% on BeyondSWE-Doc2Repo.

DeNovoSWE: The First Long-Horizon Doc2Repo Training Set, Teaching Code Agents to Build Repositories

The Gaoling School of Artificial Intelligence at Renmin University of China has released DeNovoSWE — a dataset focused on long-horizon software engineering tasks, especially "repository-level code generation from scratch" (Doc2Repo). It answers a key question: now that Code Agents keep getting better at "fixing an issue" or "patching a few lines of bug code," how do you make one truly capable of starting from a document, planning the architecture, creating files, designing APIs, wiring up modules, and ultimately generating a complete, runnable repository? The answer — you need training environments purpose-built for long-horizon tasks. The dataset, code, and models are all open-sourced.

What Problem DeNovoSWE Solves

Over the past year, scaling on large-scale SWE data such as Scale-SWE has driven rapid progress for coding agents on tasks like SWE-bench. But frontier models still fare poorly on BeyondSWE-Doc2Repo and NL2RepoBench.

Real-world software development is often not fixing one function or adding one conditional check; it is:

  • Understanding requirements
  • Planning the architecture
  • Creating files
  • Designing APIs
  • Handling dependencies
  • Wiring modules together
  • Getting the whole repository to pass its tests

This is long-horizon repository-level generation. DeNovoSWE systematically constructs this goal into a trainable, verifiable, extensible dataset.

DeNovoSWE: rebuilding an entire repository from a single document

The Core Data

  • Scale: 4818 high-quality document-to-repository task instances
  • Properties: executable, evaluable, trainable long-horizon software engineering environments
  • Construction method: built automatically via a sandboxed multi-agent workflow, not human-written documents
BenchmarkBase modelScale-SWE trainingDeNovoSWE training
BeyondSWE-Doc2Repo5.8%29.2%47.2%
NL2RepoBench4.3%18.3%23.0%

The base model is Qwen3-30B-A3B-Instruct. This shows that data aimed at "fixing bugs" cannot fully substitute for long-horizon data aimed at "generating complete repositories."

On the stronger Qwen3.5-35B-A3B backbone, DeNovoSWE likewise delivers stable gains: BeyondSWE-Doc2Repo improves from 43.8% to 50.0%, and NL2RepoBench from 23.5% to 27.1% — further proof that the gains come from the high-quality long-horizon data itself.

Method: Divide & Conquer + Critic & Repair

The whole method has two steps.

The Divide Stage: Decomposing Repository Capabilities

The system analyzes the target repository and breaks it into multiple repository capabilities, each corresponding to a core capability or workflow (auth and connections, data read/write, batch processing, export pipelines, and so on). The originally massive repository-generation problem is split into several clearly structured document sections.

At the same time, DeNovoSWE runs the original unit tests and collects execution traces to identify which functions, classes, and interfaces actually affect evaluation, further classifying them into three groups:

  • direct components: interfaces called directly by the tests; must be documented in detail
  • core indirect components: core indirect components that affect observable behavior; must be covered
  • non-core indirect components: non-core internal implementation, left to the agent's discretion

The Conquer Stage: The Draft-Critic-Repair Loop

Documentation is generated capability by capability using a Draft-Critic-Repair mechanism:

  1. A Draft agent writes the initial draft
  2. A Critic agent checks for missing key APIs, behavioral contracts, or structural information
  3. A Repair agent fixes the document based on the feedback

The loop keeps iterating until each capability section is sufficiently clear, complete, and aligned with the evaluation. Finally, the different capability documents are merged into one complete task document.

Key Design: Two Criteria for High-Quality Task Documents

In document-to-repository generation, the document is neither a README nor a simple API list — it is the agent's only task entry point for rebuilding the entire repository. A high-quality document must satisfy at least two criteria.

1. It must be well-organized

Repository-level tasks are inherently complex. If the document just piles up function descriptions, the agent can easily get lost in fragmented information. The document should:

  • Open with a clear repository-level overview
  • Then split into sections by capability or workflow
  • Give every part a clear functional boundary

2. It must be grounded in reliable evaluation

The document can neither say too little (the task becomes an under-defined problem, and the model passes evaluation only through unbounded guessing) nor too much (implementation details leak directly, and the task loses its challenge).

A genuinely high-quality document should describe the key behaviors the evaluation depends on: import paths, public APIs, inputs and outputs, default parameters, exception behavior, configuration options, mode strings, returned fields, and so on. The document must be enough for the agent to reproduce testable behavior, but must not become a copy of the implementation code.

The Difficulty: Why This Is a Long-Horizon Task

DeNovoSWE's task difficulty comes from one fundamental shift: it is no longer issue-level fixing, but whole-repository generation.

The agent faces a cleaned environment:

  • Original source code and tests are removed
  • Git history is reset
  • Every potential leak channel — caches, leftover site-packages, pip wheels, temporary build artifacts — is purged

This means the agent must genuinely rely on the document to complete the entire repository rebuild: plan the project structure, create module files, define public interfaces, implement cross-file interactions, handle dependencies and configuration, and keep fixing errors through multiple rounds of editing and test feedback.

A single deviation in any API signature, returned field, exception type, or default behavior can fail the tests. Errors also accumulate over the long horizon — one poorly designed early module can affect many downstream files and call chains.

difficulty-aware trajectory filtering

To handle differences in repository difficulty, DeNovoSWE proposes difficulty-aware trajectory filtering: easy tasks are held to a higher pass rate, while hard tasks are not discarded wholesale just because they failed to reach a perfect score. Based on structural complexity and LLM difficulty judgments, different filtering thresholds are set for different difficulty ranges, balancing quality and diversity.

Use Cases

  • Training Code Agents: evolve a model from "repository maintainer" into "architect," mastering repository-level code generation
  • Long-horizon task evaluation: training on and improving benchmarks such as BeyondSWE-Doc2Repo and NL2RepoBench
  • SWE data scaling research: filling the long-horizon data gap for "generating complete repositories"
  • Agent workflow research: mechanisms like Divide & Conquer, Critic & Repair, and Draft-Critic-Repair can be borrowed directly

The next stage for coding agents is not just fixing individual issues faster, but being able to understand documents, plan architectures, organize modules, implement interfaces, and ultimately generate a complete, runnable software repository. DeNovoSWE systematically constructs this goal into a trainable, verifiable, extensible dataset.