AgentDoG 1.5: An Open-Source Security Diagnostics and Guardrail Framework for AI Agents
A lightweight open-source agent security tool from Shanghai AI Lab that analyzes execution-trace risks with three-dimensional diagnostics and supports online guardrail deployment


AgentDoG 1.5: An Open-Source Security Diagnostics and Guardrail Framework for AI Agents
A lightweight open-source agent security tool from Shanghai AI Lab that analyzes execution-trace risks with three-dimensional diagnostics and supports online guardrail deployment
Once AI agents graduate from "answering questions" to "calling tools, running commands, operating systems," security is no longer just a content moderation problem. AgentDoG 1.5, released by Shanghai AI Laboratory, is an open-source agent security diagnostics framework that analyzes complete execution traces, pinpoints the source of risk, and deploys online guardrails. It's a fit for agent platform developers, security engineers, and AI infrastructure teams.
What Is AgentDoG 1.5
AgentDoG 1.5 is an open-source AI agent security diagnostics and online guardrail framework from Shanghai AI Laboratory. The core idea: don't judge the final output — inspect the complete execution trace.
An agent can look perfectly fine in its final reply while having already mis-called tools, leaked information, or executed dangerous commands along the way. AgentDoG 1.5 makes its security judgment by analyzing the full agent trajectory (user request, intermediate responses, tool calls, environment feedback, final reply).

Core Features
Three-Dimensional Security Diagnostics
AgentDoG 1.5 doesn't just label safe / unsafe; it outputs fine-grained diagnostics along three dimensions:
- Risk Source: where the risk comes from (tool descriptions? environment feedback? memory injection?)
- Failure Mode: how the agent failed (wrong tool call? approval bypass? goal drift?)
- Real-world Harm: what real damage it could cause (data leaks? file corruption? system compromise?)
An Extensible Taxonomy
Different agent platforms face entirely different risks. AgentDoG 1.5 keeps the three high-level dimensions fixed while extending the concrete categories per scenario:

For example:
- OpenClaw scenarios: persistent-session risks, approval bypass, plugin supply-chain attacks, cross-tool attack chains
- Codex scenarios: repository file injection, dependency supply chains, dangerous shell execution, destructive workspace modifications
The ATBench Family Benchmarks
The paper builds three benchmarks that share one framework:
- ATBench: general tool-use agents
- ATBench-Claw: OpenClaw cross-app execution scenarios
- ATBench-Codex: Codex code execution scenarios
Online Guardrail
AgentDoG 1.5 can be deployed as a Pre-Reply intervention mechanism: before the agent's final reply is sent to the user, it reads the complete execution trace and decides whether to let it through.
This design runs detection only once before the final reply, avoiding per-tool-call inspection and keeping latency low.
Performance Data
AgentDoG 1.5 trains its lightweight models (0.8B / 2B / 4B / 8B) on only about 1,000 high-quality samples, yet it punches well above its weight:
| Metric | AgentDoG 1.5-4B |
|---|---|
| R-Judge Accuracy | 92.2% |
| R-Judge F1 | 92.7% |
| ATBench Accuracy | 72.4% |
| ATBench F1 | 74.3% |
Online Guardrail Results
In OpenClaw online evaluations:
| Scenario | ASR Before Guardrail | ASR After Guardrail |
|---|---|---|
| ClawSafety | 56.25% | 18.75% |
| AgentHarm (Prompt Intelligence Theft) | 41.92% | 26.92% |
| CIK-Bench (retained) | 94.29% | 42.86% |

The Safety Training Pipeline
AgentDoG 1.5 is not just an evaluation model; it can also plug into agent training pipelines:
- SFT stage: filters for high-quality safe trajectories, cutting the AgentHarm harm score from 57.49% to 20.32%
- RL stage: builds a lightweight Python simulator environment supporting 10,000 concurrent environments with peak memory under 2.5GB
Resources
- Paper: https://arxiv.org/abs/2605.29801
- GitHub: https://github.com/AI45Lab/AgentDoG
- Hugging Face: https://huggingface.co/collections/AI45Research/agentdog15
All code, models, and data are open-sourced.
Use Cases
- Agent platform security teams: deploy it as an online guardrail to intercept dangerous agent behavior
- Agent developers: use AgentDoG during development to evaluate your agent's security
- AI safety researchers: use the ATBench Family to build and evaluate new agent security approaches
- Enterprise IT security: run security audits and risk assessments before deploying internal agents
Related articles

Astribot T1: An 89,900-Yuan Humanoid Robot Goes on Sale
Astribot launches the T1 humanoid robot starting at 89,900 yuan, with a three-in-one architecture of tendon-driven body + in-house AI model + embodied OS, shipping from June 1.

Volcano Engine AI Trust: A Three-Layer Architecture Guarding Agent Security
Volcano Engine launches the AI Trust security product family, covering trustworthy models, controllable agents, and AI-driven security operations, with 10 billion detection calls per day

Xiaomi MiMo API Prices Cut Up to 99% for Good: How Developers Can Grab the Deal
Xiaomi's MiMo-V2.5 series API gets a permanent price cut of up to 99%, Token Plan allowances grow 5-8x, and pricing now squarely matches DeepSeek

Darwin Skill 2.0: Let Your AI Skills Evolve on Their Own
darwin-skill 2.0 is an open-source Skill/Prompt auto-optimization tool that distills the best of two Microsoft papers, using multi-judge independent review plus human checkpoints to lift your AI skill documents from 80 points to over 90.

An Open-Source Social Card Skill: Say Goodbye to AI-Looking Images
guizang-social-card-skill is an open-source AI text-and-image card generator with 11 built-in content category adaptations, a magazine-grade layout engine, and free commercial-use image libraries, helping you one-click generate RedNote-grade card images.

Hy-Memory: Give Your AI Agent a Supercharged Memory
Hy-Memory is an OpenClaw memory plugin from Tencent. With a three-part architecture — a 6-layer memory framework, System1/System2 dual processing, and evolution chains — it lets an Agent truly remember your preferences, decisions, and history, cutting memory fragments by over 70%.