PerfEvolve: Teaching Agents to Tune Databases Like a Senior DBA
ISCAS open-sources the PerfEvolve framework, converting static tuning docs into executable procedural skills for Agents and delivering up to 58.9% performance gains on PostgreSQL v16 — a direct fix for LLM tuning failures.


PerfEvolve: Teaching Agents to Tune Databases Like a Senior DBA
ISCAS open-sources the PerfEvolve framework, converting static tuning docs into executable procedural skills for Agents and delivering up to 58.9% performance gains on PostgreSQL v16 — a direct fix for LLM tuning failures.
Automatic database tuning has long been a "looks perfect in theory, falls apart in practice" routine for LLM Agents. Feed them all the official docs and tuning guides you want — LLM Agents still stumble the moment they start: either they copy stale parameter values onto new hardware and crash, or parameters "fight" each other and performance degrades the more they tune.
The PerfEvolve framework, jointly launched by teams from the Institute of Software, Chinese Academy of Sciences (ISCAS) — its intelligent software research center and Foundation Software and Systems Key Laboratory — arrives at a counterintuitive conclusion: the problem was never that large models can't read documentation; it's that traditional tuning docs themselves are fatally flawed — they hand out "final answers" but never teach the "solution process." Conclusions expire; only processes transfer.
PerfEvolve doesn't make LLMs memorize parameter values; it converts static tuning documentation into procedural tuning skills that Agents can execute directly and apply autonomously. It achieves up to 58.9% performance improvement on PostgreSQL v16 and is open-sourced.

PerfEvolve's core: upgrading "reading the manual" to "knowing how to profile."
Before You Start
- Python 3.10+
- A reachable PostgreSQL v16 instance (spin up a clean test database with Docker — don't tune in production)
- Resources:
- Paper: https://arxiv.org/abs/2605.19988
- GitHub: https://github.com/ISCAS-OSLab/PerfEvolve
- Bundled: 23 executable tuning skills + 160 parameter characteristic profiles
Estimated time to run the minimal example: 1 hour.
Step 1: Understand the Three Fatal Flaws of Traditional Tuning
Know the problem first, and what PerfEvolve is solving becomes clear:
- Static docs lag badly: most database parameter recommendations date back to old versions and old hardware. Many of PostgreSQL's core parameters still carry official advice values from the spinning-disk era; drop them into an all-SSD environment and they don't just stop helping — they drag performance down.
- "Universal recommendations" flip completely across scenarios: the classic is
shared_buffers = 25% RAM— every DBA has seen it, yet this single misfit setting alone can cost 5%–16% in performance. - Parameters fight each other: each value looks perfectly sound in solo testing, but combined they crater performance or crash outright. Database tuning has never been 1+1=2.
PerfEvolve's diagnosis: documents that only hand out fixed parameter values cannot adapt to endlessly varying hardware, workloads, and system versions. The real bottleneck isn't "the Agent didn't read enough docs" — it's that "the docs gave the destination coordinates but no route."
Step 2: Clone the Repo and Load the PostgreSQL v16 Skill Pack
git clone https://github.com/ISCAS-OSLab/PerfEvolve.git
cd PerfEvolve
pip install -r requirements.txt
# Spin up a clean PG v16 test instance
docker run -d --name pg16-test \
-e POSTGRES_PASSWORD=test \
-p 5432:5432 postgres:16
# Load the official 23 tuning skills + 160 parameter characteristic profiles
python -c "
from perfevolve import SkillPack
pack = SkillPack.load('skills/postgres_v16')
print(f'{len(pack.skills)} skills, {len(pack.profiles)} profiles loaded')
"Each skill is not a set of recommended values but a structured, executable SOP: prerequisites → hands-on steps → decision criteria → post-verification → offline measured data.
Step 3: Wire PerfEvolve into Your Agent
PerfEvolve standardizes the Agent tuning process into one complete workflow:
Benchmark default-config performance → scan candidate parameter values → watch the throughput curve
→ locate the peak region → check for strong interactions with other parameters
→ if any, enter joint optimization → finally validate stability across workloadsIntegration approach (pseudocode, wired to your own LLM Agent):
from perfevolve import TuningWorkflow, AgentAdapter
class MyLLMAgent(AgentAdapter):
def decide_next_step(self, observation, history):
# observation: current perf curve, current parameter values, workload characteristics
# PerfEvolve has already structured "what to do next"
# the Agent just picks from structured options instead of inventing commands on instinct
return self.llm.choose_action(observation, history)
wf = TuningWorkflow(
db_url="postgresql://postgres:test@localhost:5432",
skill_pack="skills/postgres_v16",
workload="tpcc", # or sysbench / custom
)
result = wf.run(agent=MyLLMAgent(), iterations=10)
print(f"improvement: {result.improvement_pct}%")
print(f"best_config: {result.best_config}")Step 4: Understand Its Two Moves for Shrinking the Search Space
PerfEvolve drives down search cost with two mechanisms:
- Sensitivity dimensionality reduction: a quick scan first decides "which parameters are worth tuning, which can be left alone, where each parameter's safe range lies, and what shape its response curve takes." The 160 parameter characteristic profiles cut the "invalid search space" in advance.
- Parameter interaction topology map: identifies which parameters are strongly correlated and must be tuned together, avoiding "individually optimal, jointly disastrous." This is the key to solving "parameter infighting."

Sensitivity reduction + parameter interaction topology compress the search space from exponential to executable scale.
Results
When it finishes you get a tuning report:
Workload: tpcc (PG v16)
Default throughput: 1240 TPS
Optimized throughput: 1968 TPS
Improvement: 58.9%
Top changed params:
shared_buffers: 25%RAM → 38%RAM (memory-bound workload)
work_mem: 4MB → 64MB
effective_io_conc: 1 → 200 (NVMe detected)
max_worker_processes: 8 → 16
Cross-load validation: stable across tpcc/sysbench/read-heavyIf performance drops after the Agent tunes, don't blame the model first — check whether PerfEvolve's offline characteristic profiles cover your workload type; that's the foundation of this framework's correctness.
Common Issues
- Performance drops after Agent tuning: nine times out of ten it's "parameter infighting." Check the strongly correlated parameter groups flagged red in the interaction topology map, and make sure they're tuned jointly rather than individually.
shared_bufferskeeps getting tuned back to 25%: the classic symptom of an LLM reciting the docs' "standard answer." The PerfEvolve paper verified this specifically — even handing the LLM the entirely correct parameter answer can still make performance worse, because the number "anchors" the model cognitively. Switching from declarative knowledge to procedural skills fixes it.- Cross-load validation fails: the tuning result is overfit to a single workload. Raise the strictness of
cross_load_validationin the workflow and send the Agent back into joint optimization.
💡 Tip: PerfEvolve's core insight goes beyond "tuning databases" — it reveals a paradigm useful for all Agents: a hundred standard answers are worth less than one self-sufficient problem-solving ability. The same thinking applies to ops, network configuration, CI/CD pipelines, and any other domain where process matters more than conclusions.