October 5, 2026ResearchMonitoringAgents

Actions with Receipts: A Valid Citation and a Valid Trace Can Still Be a Lie

Here's an agent failure most audit tools can't see. The citation is real, the execution log is real, and the claim shown to the user didn't come from either. Actions with Receipts, from Miaobo Hu, Yina Sa and colleagues, calls it transplantation: a well-formed citation and a well-formed trace get moved across claims, actions, runs or source versions. Each piece checks out on its own. The link between them was never checked.

The fix is a claim-anchored execution contract, a receipt for every claim. It binds four things together: the exact claim emitted, the exact source span with offsets, hashes and quotes, the ordered execution prefix that produced it, and the source version and access state that execution actually saw. A deterministic integrity verifier rebuilds those bindings before any semantic or task judgment runs. Structural validity is kept on a separate plane from "does the source support this", so a clean receipt isn't mistaken for a true claim.

The numbers are strong. Across 1,280 cross-object substitution attacks the joint contract catches 1,275, a rate of 0.9961. Remove any one of the seven properties and detection for its targeted attack falls to between 0.0156 and 0.0625, so each binding is doing real work. On a separately adjudicated 384-pair split, the support guard reaches F1 0.8865 with a 7.29% false-acceptance rate, and holds at 0.8679 on unseen failure families.

This is the trace-integrity thread getting an engineering spec. Monitors that read the trace assume the trace belongs to the answer. Receipts make that assumption checkable, and they cost hashes and offsets, not model calls.

Link: arxiv.org/abs/2610.00327
← Previous
PACE: Stop Vetting Agent Inputs, Check the Tool Call Right Before It Fires
Next β†’
RuleEvolve: Your CLAUDE.md Is Hand-Tuned, and an Evolution Loop Beats It
← Back to all articles

Comments

Loading...
>_