FuseSearch: How a 4-Billion-Parameter Small Model Outguns Commercial LLMs at Code Localization
Ant Group's ACL 2026 work FuseSearch-4B uses an adaptive parallel search strategy to match Claude Haiku 4.5 on code localization — 93.6% faster with 68.9% fewer tokens


FuseSearch: How a 4-Billion-Parameter Small Model Outguns Commercial LLMs at Code Localization
Ant Group's ACL 2026 work FuseSearch-4B uses an adaptive parallel search strategy to match Claude Haiku 4.5 on code localization — 93.6% faster with 68.9% fewer tokens
In real-world AI coding, more than 50% of compute is burned on code search and localization. Ant Group's FuseSearch-4B proposes a counterintuitive solution: you don't need to pile on parameters — you just need to teach the model "how much to search, and when." This open-source model with only 4 billion parameters reaches 84.7% file-level F1 on SWE-bench Verified, matching Claude Haiku 4.5's localization ability.
The Problem: Why Code Localization Is So Expensive
When a coding agent hunts through a large project with hundreds of thousands of lines for which file and which function to change, existing approaches suffer from two pain points:
- Single-step serial search: each round can call only one tool to narrow the scope step by step, consuming an astonishing number of rounds
- Brute-force parallelism: a fixed 8 tool calls per round produces over 34.9% redundant calls and introduces noisy signals
The core tension: too little parallelism means not enough information; too much wastes resources. FuseSearch's insight — the key isn't how much parallelism you use, but when to parallelize heavily and when to hold back.

FuseSearch uses only three read-only tools: glob to find files, grep to search contents, and read_file to read details.
Core Innovations
A Three-Tool "Swiss Army Knife"
FuseSearch's toolbox is remarkably restrained — just three read-only tools:
- glob: find files by filename pattern
- grep: search within file contents
- read_file: read file details
Zero dependencies — it works out of the box. No code knowledge graph, no syntax parsers. Language-agnostic: it works on both Python and Java repositories.
Quantifying Search Quality with "Information Gain"
The paper introduces, for the first time, a Tool Efficiency metric:
Information gain = newly discovered code entities / total returned code entities
High efficiency means every search is exploring new territory; low efficiency means it's doing redundant work. This metric turns "search quality" directly into a quantifiable training objective.

Two-Stage Training
Stage one: supervised fine-tuning (SFT)
About 21,000 issue-patch pairs were extracted from 233 high-quality GitHub repositories, and search trajectories were generated with Kimi-K2-Instruct. The screening bar was twofold: localization accuracy >= 0.8 and tool efficiency >= 0.5. In the end, about 6,000 high-quality samples were curated from roughly 24,000 candidates.
Stage two: reinforcement learning (RL)
The reward function design is exquisite:
reward = 0.8 x localization accuracy + 0.2 x (localization accuracy x tool efficiency)Note that product term: the extra reward only lands when "finding it accurately" and "searching without waste" are satisfied simultaneously. If the localization is completely wrong, the reward is zero no matter how high the efficiency — the model cannot "fail efficiently."
Training Results: Cast a Wide Net, Then Reel It In
After RL training, the model taught itself a veteran-driver-style adaptive search pattern:
- Early stage: cast a wide net, covering the codebase quickly with high parallelism
- Middle stage: gradually narrow down, searching in depth along leads
- Late stage: precise verification, confirming key locations with low parallelism
This "breadth first, then depth" pattern was learned entirely by the model from reward signals, with no hand-written rules of any kind.
Experimental Data
Core Metrics (SWE-bench Verified, 386 instances)
| Metric | FuseSearch-4B | vs. prior methods |
|---|---|---|
| File-level F1 | 84.7% | Accuracy doubled |
| Speed | 93.6% faster | 16x faster |
| Token consumption | Down 68.9% | Nearly 70% saved |
Against Commercial Closed-Source Models
A 4B open-source model that can be deployed locally matches Claude Haiku 4.5 in localization ability while being faster and cheaper at the same time.
Plugging into Downstream Agents
Using FuseSearch-4B as Kimi-K2-Instruct's "front-end search engine" cuts costs nearly in half without affecting fix quality.
Practical Details
- Paper title: FuseSearch: Learning Adaptive Parallel Execution for Efficient Code Localization
- Venue: ACL 2026 Findings
- Authors' affiliation: Ant Group
- Paper link: https://github.com/sxthunder/FuseSearch
- Deployment cost: zero dependencies, three read-only tools, deployable instantly to any code repository
Who It's For
- Coding agent developers: anyone who needs to cut the token consumption and latency of code search
- Enterprise code repair: industrial-grade scenarios sensitive to cost and latency
- Local deployment needs: the 4B model runs on consumer GPUs