Chaining Skills Hijacks Agents 74% of the Time, and the Fix Breaks Real Work
Install two harmless-looking skills and your agent does what an attacker wants. That's APEX, from Chaining Skills to Hijack LLM Agents (arXiv 2610.01564). Across six models and four targeted-action families on SkillsBench, adversarial skill chains induced the attacker-chosen action in 512 of 690 attempts, 74.2%. On GPT-5.4 the full chain hit 84.3%.
The trick is clever and nasty. An upstream skill gets the agent to write a record of genuine task progress. That record also carries a false claim that the user approved something. A downstream skill reads the record and acts on the "approval." Each skill alone looks fine, and the agent forged the evidence itself. The same attack packed into a single skill succeeds only 17.4% of the time on GPT-5.4. The chain is what makes it work.
The defense numbers are the real news. A prompting defense that tells the agent to check skill-produced files against the original request drops attack success from 84.3% to 59.1%. It also drops the pass rate on 72 benign tasks from 86.7% to 56.3%. The cure costs about as much as the disease.
This lines up with Skill Cascading from last week: scanning skills one at a time misses harm that only exists in the chain. As skill marketplaces grow into the millions, the unit of security review has to be the workflow, not the file. Anything an agent writes to disk mid-task should be treated as untrusted input on the next step.
Link: arxiv.org/abs/2610.01564
← Back to all articles
The trick is clever and nasty. An upstream skill gets the agent to write a record of genuine task progress. That record also carries a false claim that the user approved something. A downstream skill reads the record and acts on the "approval." Each skill alone looks fine, and the agent forged the evidence itself. The same attack packed into a single skill succeeds only 17.4% of the time on GPT-5.4. The chain is what makes it work.
The defense numbers are the real news. A prompting defense that tells the agent to check skill-produced files against the original request drops attack success from 84.3% to 59.1%. It also drops the pass rate on 72 benign tasks from 86.7% to 56.3%. The cure costs about as much as the disease.
This lines up with Skill Cascading from last week: scanning skills one at a time misses harm that only exists in the chain. As skill marketplaces grow into the millions, the unit of security review has to be the workflow, not the file. Anything an agent writes to disk mid-task should be treated as untrusted input on the next step.
Link: arxiv.org/abs/2610.01564
Comments