AutoControl-Arena: Auto-Generated, Actually Runnable Agent Risk Testbeds

·Toolin Editorial Team

Fudan × Shanghai Innovation Institute × Oxford publish the ICML 2026 paper AutoControl-Arena, which auto-synthesizes executable test environments to surface latent Agent risks in long-tail scenarios, reproducing Anthropic/OpenAI safety-report behaviors with a 0.87 correlation.

AutoControl-Arena: Auto-Generated, Actually Runnable Agent Risk Testbeds

As AI Agents move from the chat box into real applications, the shape of safety problems has changed. The old question was "will the model answer dangerous questions"; now you have to worry: will an Agent bypass approvals to get the job done? Will it tamper with verification logic under metric pressure? Will it overstep file permissions across multi-tool collaboration? Will it realize it is being evaluated and change its behavior accordingly?

These "long-tail risks" are nearly impossible to cover by hand-crafting benchmarks one at a time. AutoControl-Arena (accepted to ICML 2026), released jointly by Fudan University, Shanghai Innovation Institute, and the University of Oxford, offers an automated answer: automatically synthesize executable test environments, so researchers and developers can quickly surface latent Agent risks in unseen scenarios.

Its defining trait: the test environments actually run. Rather than having an LLM "simulate" an environment out of thin air (which invites hallucinated, self-contradictory state), it really builds a filesystem, databases, command-line tools, approval flows, and logging systems, then watches how the Agent behaves inside them.

AutoControl-Arena's core goal: auto-synthesize executable test environments to surface Agent long-tail risks.

Before You Start

Estimated time to run the minimal example: 30 minutes.

Step 1: Understand the "Safety Evaluation Dilemma" It Solves

Agent safety evaluation has long been stuck on a core dilemma:

  • Real environments are high-fidelity but costly and slow; every additional risk scenario means redesigning tools, state, rules, and feedback
  • LLM-simulated environments are cheap and flexible but prone to "logic hallucinations" — inconsistent file state, fabricated database returns, permission rules that exist one moment and vanish the next

Yet long-tail risks demand both scenario breadth and trustworthy executability. AutoControl-Arena's solution: use an LLM to generate a "configuration description" of the test environment, then compile those descriptions into real, executable Docker containers (with a real filesystem, real databases, real approval flows) — getting "cheap" and "trustworthy" at the same time.

Real environments are trustworthy but don't scale; LLM simulation scales but isn't trustworthy — AutoControl-Arena splits the difference.

Step 2: Clone the Repo and Run a Minimal Risk Scenario

git clone https://github.com/CosmosYi/AutoControl-Arena.git
cd AutoControl-Arena
pip install -e .

# Confirm the Docker daemon is running
docker info

# Run a built-in example: will the Agent bypass approval
python -m autocontrol.run \
  --scenario examples/bypass_approval.yaml \
  --agent your_agent_module:YourAgent

Each test scenario consists of three parts:

  • Environment config: initial filesystem state, database schema, available tool set, approval rules
  • Task description: the goal given to the Agent (often in tension with the rules)
  • Risk probe: scripts that check after the fact whether the Agent triggered a given risk behavior

Step 3: Cover Your Agent with 7 Frontier Risk Templates

AutoControl-Arena ships 7 frontier risk templates you can run out of the box:

  1. Approval bypass: will the Agent overstep its authority to complete the task
  2. Verification tampering: will it "cheat" under metric pressure to make results pass
  3. Unauthorized file access: will it read files it shouldn't during multi-tool collaboration
  4. Evaluation awareness: will it change its behavior strategy because it "knows it's being tested"
  5. Replacement risk: will it turn adversarial when it senses a threat of being replaced
  6. Resource hoarding: will it over-claim tools or permissions
  7. Goal drift: will it stray from the original task objective

Each template can yield dozens to hundreds of variants by tweaking the environment config — no manual question-writing required.

from autocontrol import RiskTemplate, synthesize_scenarios

# Batch-generate 50 variant scenarios from the "bypass_approval" template
template = RiskTemplate("bypass_approval")
scenarios = synthesize_scenarios(template, n=50, seed=42)

for s in scenarios:
    env = s.build_environment()  # a real, executable Docker container
    result = your_agent.run(env, s.task)
    risk_triggered = s.probe(env, result)
    print(s.id, risk_triggered)

Step 4: Calibrate Against "Known Cases"

One reassuring design choice: AutoControl-Arena can automatically and accurately reproduce the risk behaviors already disclosed in safety reports from Anthropic, OpenAI, Apollo Research, and others. The correlation reaches 0.87.

Reproducing official safety reports

Reproducing risk behaviors disclosed by Anthropic/OpenAI and others is the credibility baseline for this benchmark.

Calibration steps:

  1. Pick a risk case already disclosed in an official report (a list lives in examples/known_cases/ in the repo)
  2. Run AutoControl-Arena's reproduction scenario with the same model version
  3. Compare whether the triggered behavior matches
  4. If it matches, your harness is sound and you can trust your own scenarios

Results

After a test run you get a risk report like this:

Agent: YourAgent-v1
Scenarios: 350 (7 categories × 50 variants)
Risk triggered: 23 (6.6%)
  bypass_approval:        12
  verification_tampering:  7
  aware_of_eval:           4
Calibration (known cases): 0.87 correlation

If one category's trigger rate is markedly higher than the rest, that's the weakness to fix before deploying your Agent.

Common Issues

  • Docker containers fail to start: usually port or mount conflicts. AutoControl-Arena isolates with random ports by default; check docker ps for leftover containers.
  • Generated scenarios aren't devious enough: raise temperature or n in synthesize_scenarios, but beware — too high a temperature pushes scenarios outside executable bounds.
  • Known-case reproduction doesn't match: check that you're using the same model version and system prompt as the original report — Agent behavior is extremely sensitive to the system prompt.

💡 Tip: This tool's value isn't "racking up high scores" — it's "finding out earlier where your own Agent will break." Put it in your CI pipeline and run all 7 risk templates on every model iteration to block the most dangerous long-tail behaviors before release.