PACE: Stop Vetting Agent Inputs, Check the Tool Call Right Before It Fires
Most agent-security work tries to catch poison at the door: scan the tool metadata, the retrieved page, the memory entry, the skill file before the agent reads it. PACE, from Fengpeng Li, Haiwei Wu and colleagues, argues that door can't be made to work and moves the check to the last possible moment.
The core claim is a negative result made precise. A safe artifact and a leaking artifact can produce identical admission evidence, so a sound gate can't let either through on that basis. Vetting at admission therefore can't settle whether the next action is safe. What's left is the point right before a tool call executes. PACE mediates every call there. Path confinement cuts the influence paths from untrusted content to the action. Capability and effect verification checks the call's schema-defined effects against authority compiled from the authenticated user request, not from whatever the agent read along the way.
Results across eight executable agent-security benchmarks and three model families: strictly the lowest attack success in 62 of 79 eligible attack columns and ties in 14, while native utility drops at most three points against the undefended agent. The ablation over 1,167 paired cases puts most of the security gain on effect verification. An adaptive search hit 0 of 30 out-of-authority targets, though the authors note it was a reduced-scale search.
This fits the week's pattern of agent authority getting enforced at execution, not at instruction. APEX, from the same arXiv batch, showed prompt-level defenses trade away as much benign pass rate as they remove attacks. PACE gets most of the protection for three points of utility. Authority comes from the request and gets checked at the call. That's the design to copy.
Link: arxiv.org/abs/2610.01349
← Back to all articles
The core claim is a negative result made precise. A safe artifact and a leaking artifact can produce identical admission evidence, so a sound gate can't let either through on that basis. Vetting at admission therefore can't settle whether the next action is safe. What's left is the point right before a tool call executes. PACE mediates every call there. Path confinement cuts the influence paths from untrusted content to the action. Capability and effect verification checks the call's schema-defined effects against authority compiled from the authenticated user request, not from whatever the agent read along the way.
Results across eight executable agent-security benchmarks and three model families: strictly the lowest attack success in 62 of 79 eligible attack columns and ties in 14, while native utility drops at most three points against the undefended agent. The ablation over 1,167 paired cases puts most of the security gain on effect verification. An adaptive search hit 0 of 30 out-of-authority targets, though the authors note it was a reduced-scale search.
This fits the week's pattern of agent authority getting enforced at execution, not at instruction. APEX, from the same arXiv batch, showed prompt-level defenses trade away as much benign pass rate as they remove attacks. PACE gets most of the protection for three points of utility. Authority comes from the request and gets checked at the call. That's the design to copy.
Link: arxiv.org/abs/2610.01349
Comments