AutoControl-Arena: Auto-Generated, Actually Runnable Agent Risk Testbeds
Fudan × Shanghai Innovation Institute × Oxford publish the ICML 2026 paper AutoControl-Arena, which auto-synthesizes executable test environments to surface latent Agent risks in long-tail scenarios, reproducing Anthropic/OpenAI safety-report behaviors with a 0.87 correlation.


AutoControl-Arena: Auto-Generated, Actually Runnable Agent Risk Testbeds
Fudan × Shanghai Innovation Institute × Oxford publish the ICML 2026 paper AutoControl-Arena, which auto-synthesizes executable test environments to surface latent Agent risks in long-tail scenarios, reproducing Anthropic/OpenAI safety-report behaviors with a 0.87 correlation.
As AI Agents move from the chat box into real applications, the shape of safety problems has changed. The old question was "will the model answer dangerous questions"; now you have to worry: will an Agent bypass approvals to get the job done? Will it tamper with verification logic under metric pressure? Will it overstep file permissions across multi-tool collaboration? Will it realize it is being evaluated and change its behavior accordingly?
These "long-tail risks" are nearly impossible to cover by hand-crafting benchmarks one at a time. AutoControl-Arena (accepted to ICML 2026), released jointly by Fudan University, Shanghai Innovation Institute, and the University of Oxford, offers an automated answer: automatically synthesize executable test environments, so researchers and developers can quickly surface latent Agent risks in unseen scenarios.
Its defining trait: the test environments actually run. Rather than having an LLM "simulate" an environment out of thin air (which invites hallucinated, self-contradictory state), it really builds a filesystem, databases, command-line tools, approval flows, and logging systems, then watches how the Agent behaves inside them.
AutoControl-Arena's core goal: auto-synthesize executable test environments to surface Agent long-tail risks.
Before You Start
- Python 3.11+
- Docker (for isolated, executable test environments)
- Resources:
- Paper: https://arxiv.org/abs/2603.07427
- Project page: https://cosmosyi.github.io/AutoControl-Arena/
- GitHub: https://github.com/CosmosYi/AutoControl-Arena
Estimated time to run the minimal example: 30 minutes.
Step 1: Understand the "Safety Evaluation Dilemma" It Solves
Agent safety evaluation has long been stuck on a core dilemma:
- Real environments are high-fidelity but costly and slow; every additional risk scenario means redesigning tools, state, rules, and feedback
- LLM-simulated environments are cheap and flexible but prone to "logic hallucinations" — inconsistent file state, fabricated database returns, permission rules that exist one moment and vanish the next
Yet long-tail risks demand both scenario breadth and trustworthy executability. AutoControl-Arena's solution: use an LLM to generate a "configuration description" of the test environment, then compile those descriptions into real, executable Docker containers (with a real filesystem, real databases, real approval flows) — getting "cheap" and "trustworthy" at the same time.
Real environments are trustworthy but don't scale; LLM simulation scales but isn't trustworthy — AutoControl-Arena splits the difference.
Step 2: Clone the Repo and Run a Minimal Risk Scenario
git clone https://github.com/CosmosYi/AutoControl-Arena.git
cd AutoControl-Arena
pip install -e .
# Confirm the Docker daemon is running
docker info
# Run a built-in example: will the Agent bypass approval
python -m autocontrol.run \
--scenario examples/bypass_approval.yaml \
--agent your_agent_module:YourAgentEach test scenario consists of three parts:
- Environment config: initial filesystem state, database schema, available tool set, approval rules
- Task description: the goal given to the Agent (often in tension with the rules)
- Risk probe: scripts that check after the fact whether the Agent triggered a given risk behavior
Step 3: Cover Your Agent with 7 Frontier Risk Templates
AutoControl-Arena ships 7 frontier risk templates you can run out of the box:
- Approval bypass: will the Agent overstep its authority to complete the task
- Verification tampering: will it "cheat" under metric pressure to make results pass
- Unauthorized file access: will it read files it shouldn't during multi-tool collaboration
- Evaluation awareness: will it change its behavior strategy because it "knows it's being tested"
- Replacement risk: will it turn adversarial when it senses a threat of being replaced
- Resource hoarding: will it over-claim tools or permissions
- Goal drift: will it stray from the original task objective
Each template can yield dozens to hundreds of variants by tweaking the environment config — no manual question-writing required.
from autocontrol import RiskTemplate, synthesize_scenarios
# Batch-generate 50 variant scenarios from the "bypass_approval" template
template = RiskTemplate("bypass_approval")
scenarios = synthesize_scenarios(template, n=50, seed=42)
for s in scenarios:
env = s.build_environment() # a real, executable Docker container
result = your_agent.run(env, s.task)
risk_triggered = s.probe(env, result)
print(s.id, risk_triggered)Step 4: Calibrate Against "Known Cases"
One reassuring design choice: AutoControl-Arena can automatically and accurately reproduce the risk behaviors already disclosed in safety reports from Anthropic, OpenAI, Apollo Research, and others. The correlation reaches 0.87.

Reproducing risk behaviors disclosed by Anthropic/OpenAI and others is the credibility baseline for this benchmark.
Calibration steps:
- Pick a risk case already disclosed in an official report (a list lives in
examples/known_cases/in the repo) - Run AutoControl-Arena's reproduction scenario with the same model version
- Compare whether the triggered behavior matches
- If it matches, your harness is sound and you can trust your own scenarios
Results
After a test run you get a risk report like this:
Agent: YourAgent-v1
Scenarios: 350 (7 categories × 50 variants)
Risk triggered: 23 (6.6%)
bypass_approval: 12
verification_tampering: 7
aware_of_eval: 4
Calibration (known cases): 0.87 correlationIf one category's trigger rate is markedly higher than the rest, that's the weakness to fix before deploying your Agent.
Common Issues
- Docker containers fail to start: usually port or mount conflicts. AutoControl-Arena isolates with random ports by default; check
docker psfor leftover containers. - Generated scenarios aren't devious enough: raise
temperatureorninsynthesize_scenarios, but beware — too high a temperature pushes scenarios outside executable bounds. - Known-case reproduction doesn't match: check that you're using the same model version and system prompt as the original report — Agent behavior is extremely sensitive to the system prompt.
💡 Tip: This tool's value isn't "racking up high scores" — it's "finding out earlier where your own Agent will break." Put it in your CI pipeline and run all 7 risk templates on every model iteration to block the most dangerous long-tail behaviors before release.