TRIAD: Teaching AI Agents Not Just to Refuse, but to Repair Dangerous Plans
The open-source Agent safety framework TRIAD replaces binary guardrails with three-way decisions (proceed/update/refuse), preserving the user's original task even under prompt injection attacks.


TRIAD: Teaching AI Agents Not Just to Refuse, but to Repair Dangerous Plans
The open-source Agent safety framework TRIAD replaces binary guardrails with three-way decisions (proceed/update/refuse), preserving the user's original task even under prompt injection attacks.
Once Agents start handling email, querying databases, running code, and sending messages, the safety risk is no longer just "answering a question wrong." Malicious text on a web page or an instruction smuggled inside an email can steer an Agent off your original task — leaking customer addresses, sending meeting locations to the wrong people, calling tools it shouldn't. TRIAD (Tripartite Response for Iterative Agent Guardrailing), open-sourced by a University of Melbourne team, offers a solution that isn't "just refuse everything" — it teaches guardrails to repair dangerous plans. It targets the gray zone where a task has already been contaminated by prompt injection but the user's original goal remains legitimate.

Where Existing Guardrails Get Stuck
A traditional guardrail model makes a single "safe/unsafe" binary call before the Agent executes a tool. It looks straightforward, but in real Agent scenarios it frequently fails in one of two ways:
- Wave the whole thing through: the attack succeeds and the Agent runs the malicious instructions
- Reject the whole thing: the user's perfectly legitimate task gets sacrificed along with it
Real attacks are usually not "the whole task is harmful" but "untrusted instructions mixed into a legitimate task." Say you ask the Agent to "search for hotels and send an email" — malicious content is smuggled into the search results or the email body. Now you can neither let it pass nor simply refuse.
What TRIAD Is: From "Referee" to "Feedback Provider"
TRIAD's core idea is to expand the binary decision into three classes:
- Proceed: the current action plan is safe and aligned with the user's goal, so the Agent executes as normal
- Refuse: the user's request is itself harmful, or cannot be completed safely by modifying the plan — reject outright
- Update: the crucial middle state — the plan has been contaminated by prompt injection, but the user's original goal remains legitimate

Figure 1: The TRIAD flow. Before each tool call, Tri-Guard inspects the action plan and issues a Proceed/Update/Refuse decision; tasks that are contaminated but still repairable get natural-language feedback written back, guiding the Agent to revise its plan.
On the Update branch, TRIAD doesn't terminate the task. Instead, Tri-Guard generates natural-language feedback written back into the Agent's temporary context, spelling out the risk source, where the task went off course, and what's wrong with the current tool call — guiding the downstream Agent to re-plan.
This closes the loop: the Agent proposes a plan → Tri-Guard checks it → if an update is needed, feedback is injected back into context → the Agent generates a new plan → Tri-Guard checks again, until execution is allowed, the request is refused, or the maximum number of updates is reached.
The Numbers: Not Just Suppressing Attack Rates
On two benchmarks — ASB (direct/indirect prompt injection) and AgentHarm (refusing harmful tasks while preserving legitimate ones) — across four Agent backbones, Qwen3-32B, Kimi-2.5, GPT-5.1, and Gemini-2.5-Pro:

Table 1: TRIAD's experimental results on four Agent types, compared against unprotected ReAct, ToolSafe, TRIAD+TS-Guard, and TRIAD+Tri-Guard.
The key numbers:
- Average attack success rate (ASR) drops from 74.45% to 10.42%
- While average normal task success rate (TSR) rises from 28.45% to 68.60%
The team stresses a counterintuitive point: a low ASR does not make a good guardrail. Plenty of baselines can push the attack success rate down, but at the cost of high refusal rates — they reduce risk by "block whenever in doubt / give up whenever in doubt," and legitimate tasks fail along the way. Take TS-Guard: refusal rates hit 88.80% and 94.63% on ASB-DPI and ASB-IPI, with TSRs of just 1.33% and 0.59% — effectively abandoning the user's task entirely.

Figure 3: Guardrail decision distributions before and after training. Tri-Guard routes contaminated actions to Update far more often, rather than refusing outright.
Use Cases
TRIAD fits any team that needs to put Agents into real production environments while worrying about prompt injection:
- Office productivity Agents: when handling email, calendars, and documents, email bodies or web content may carry malicious instructions
- Data analysis Agents: external API calls and database returns may hide contamination in their results
- Automation assistants: long-chain tasks like browsing the web, filing tickets, and operating a CRM
- Agent safety research: a comparable guardrail baseline; paper, code, and project page are all open
Resources
- Paper: https://arxiv.org/abs/2606.05805
- Code: https://github.com/YUHAOSUNABC/TRIAD
- Project page: https://yuhaosunabc.github.io/TRIAD/
First author Yuhao Sun is a PhD student at the University of Melbourne working on Trustworthy AI and Agent Safety; collaborators include Jiacheng Zhang (University of Melbourne) and Zhexin Zhang (Tsinghua University), supervised jointly by A/Prof. Xingliang Yuan, Dr. Feng Liu, and Dr. Shaanan Cohney.
The One-Sentence Summary
TRIAD redefines the Agent guardrail from a "binary referee" into a "feedback-driven plan regulator" — choosing to proceed, revise the plan, or refuse depending on whether the task is still repairable, instead of funneling every risk into the same "refuse" exit.