AgentDoG 1.5: An Open-Source Security Diagnostics and Guardrail Framework for AI Agents

·Toolin Editorial Team

A lightweight open-source agent security tool from Shanghai AI Lab that analyzes execution-trace risks with three-dimensional diagnostics and supports online guardrail deployment

AgentDoG 1.5: An Open-Source Security Diagnostics and Guardrail Framework for AI Agents

Once AI agents graduate from "answering questions" to "calling tools, running commands, operating systems," security is no longer just a content moderation problem. AgentDoG 1.5, released by Shanghai AI Laboratory, is an open-source agent security diagnostics framework that analyzes complete execution traces, pinpoints the source of risk, and deploys online guardrails. It's a fit for agent platform developers, security engineers, and AI infrastructure teams.

What Is AgentDoG 1.5

AgentDoG 1.5 is an open-source AI agent security diagnostics and online guardrail framework from Shanghai AI Laboratory. The core idea: don't judge the final output — inspect the complete execution trace.

An agent can look perfectly fine in its final reply while having already mis-called tools, leaked information, or executed dangerous commands along the way. AgentDoG 1.5 makes its security judgment by analyzing the full agent trajectory (user request, intermediate responses, tool calls, environment feedback, final reply).

AgentDoG 1.5 framework architecture

Core Features

Three-Dimensional Security Diagnostics

AgentDoG 1.5 doesn't just label safe / unsafe; it outputs fine-grained diagnostics along three dimensions:

  • Risk Source: where the risk comes from (tool descriptions? environment feedback? memory injection?)
  • Failure Mode: how the agent failed (wrong tool call? approval bypass? goal drift?)
  • Real-world Harm: what real damage it could cause (data leaks? file corruption? system compromise?)

An Extensible Taxonomy

Different agent platforms face entirely different risks. AgentDoG 1.5 keeps the three high-level dimensions fixed while extending the concrete categories per scenario:

The extensible taxonomy across different agent scenarios

For example:

  • OpenClaw scenarios: persistent-session risks, approval bypass, plugin supply-chain attacks, cross-tool attack chains
  • Codex scenarios: repository file injection, dependency supply chains, dangerous shell execution, destructive workspace modifications

The ATBench Family Benchmarks

The paper builds three benchmarks that share one framework:

  • ATBench: general tool-use agents
  • ATBench-Claw: OpenClaw cross-app execution scenarios
  • ATBench-Codex: Codex code execution scenarios

Online Guardrail

AgentDoG 1.5 can be deployed as a Pre-Reply intervention mechanism: before the agent's final reply is sent to the user, it reads the complete execution trace and decides whether to let it through.

This design runs detection only once before the final reply, avoiding per-tool-call inspection and keeping latency low.

Performance Data

AgentDoG 1.5 trains its lightweight models (0.8B / 2B / 4B / 8B) on only about 1,000 high-quality samples, yet it punches well above its weight:

MetricAgentDoG 1.5-4B
R-Judge Accuracy92.2%
R-Judge F192.7%
ATBench Accuracy72.4%
ATBench F174.3%

Online Guardrail Results

In OpenClaw online evaluations:

ScenarioASR Before GuardrailASR After Guardrail
ClawSafety56.25%18.75%
AgentHarm (Prompt Intelligence Theft)41.92%26.92%
CIK-Bench (retained)94.29%42.86%

Online guardrail evaluation results

The Safety Training Pipeline

AgentDoG 1.5 is not just an evaluation model; it can also plug into agent training pipelines:

  • SFT stage: filters for high-quality safe trajectories, cutting the AgentHarm harm score from 57.49% to 20.32%
  • RL stage: builds a lightweight Python simulator environment supporting 10,000 concurrent environments with peak memory under 2.5GB

Resources

All code, models, and data are open-sourced.

Use Cases

  • Agent platform security teams: deploy it as an online guardrail to intercept dangerous agent behavior
  • Agent developers: use AgentDoG during development to evaluate your agent's security
  • AI safety researchers: use the ATBench Family to build and evaluate new agent security approaches
  • Enterprise IT security: run security audits and risk assessments before deploying internal agents

Related articles

Astribot T1: An 89,900-Yuan Humanoid Robot Goes on Sale
AI Products

Astribot T1: An 89,900-Yuan Humanoid Robot Goes on Sale

Astribot launches the T1 humanoid robot starting at 89,900 yuan, with a three-in-one architecture of tendon-driven body + in-house AI model + embodied OS, shipping from June 1.

Toolin Editorial Team
Volcano Engine AI Trust: A Three-Layer Architecture Guarding Agent Security
AI Products

Volcano Engine AI Trust: A Three-Layer Architecture Guarding Agent Security

Volcano Engine launches the AI Trust security product family, covering trustworthy models, controllable agents, and AI-driven security operations, with 10 billion detection calls per day

Toolin Editorial Team
Xiaomi MiMo API Prices Cut Up to 99% for Good: How Developers Can Grab the Deal
AI Products

Xiaomi MiMo API Prices Cut Up to 99% for Good: How Developers Can Grab the Deal

Xiaomi's MiMo-V2.5 series API gets a permanent price cut of up to 99%, Token Plan allowances grow 5-8x, and pricing now squarely matches DeepSeek

Toolin Editorial Team
Darwin Skill 2.0: Let Your AI Skills Evolve on Their Own
AI Products

Darwin Skill 2.0: Let Your AI Skills Evolve on Their Own

darwin-skill 2.0 is an open-source Skill/Prompt auto-optimization tool that distills the best of two Microsoft papers, using multi-judge independent review plus human checkpoints to lift your AI skill documents from 80 points to over 90.

Toolin Editorial Team
An Open-Source Social Card Skill: Say Goodbye to AI-Looking Images
AI Products

An Open-Source Social Card Skill: Say Goodbye to AI-Looking Images

guizang-social-card-skill is an open-source AI text-and-image card generator with 11 built-in content category adaptations, a magazine-grade layout engine, and free commercial-use image libraries, helping you one-click generate RedNote-grade card images.

Toolin Editorial Team
Hy-Memory: Give Your AI Agent a Supercharged Memory
AI Products

Hy-Memory: Give Your AI Agent a Supercharged Memory

Hy-Memory is an OpenClaw memory plugin from Tencent. With a three-part architecture — a 6-layer memory framework, System1/System2 dual processing, and evolution chains — it lets an Agent truly remember your preferences, decisions, and history, cutting memory fragments by over 70%.

Toolin Editorial Team