Incident-Arena: Frontier Agents Fix Under 65% of Real Production Outages
Coding agents ship features all day. Can they handle the 3 a.m. page? Incident-Arena (arXiv 2610.00648) says not reliably. On 20 human-built incident-response tasks, frontier models score below 64.3%.
The benchmark is built to feel like a real outage. Each task deploys a production open-source application to an ephemeral Kubernetes cluster, injects a fault anywhere from the config layer down to the underlying images, and keeps a sustained load running. Verification is functional, not static: the system-level metrics have to hold, and the repair has to be done safely. That last part matters, because a fix that restarts everything and drops traffic is not a fix in production.
These are long jobs. Trials average 2.81 million tokens and 41 turns. Failures show up at every stage: wrong diagnosis or localization, incomplete repairs, and unsafe regressions where the agent makes things worse while fixing them.
The title says "the last nine of reliability," and that's the right frame. Agentic SRE is one of the most obvious enterprise buys, and vendors are already selling it. A 64% ceiling on 20 realistic tasks is a useful reality check, and the failure taxonomy is a checklist for anyone evaluating one of those products: ask how often it makes the incident worse, not just how often it closes the ticket.
Link: arxiv.org/abs/2610.00648
← Back to all articles
The benchmark is built to feel like a real outage. Each task deploys a production open-source application to an ephemeral Kubernetes cluster, injects a fault anywhere from the config layer down to the underlying images, and keeps a sustained load running. Verification is functional, not static: the system-level metrics have to hold, and the repair has to be done safely. That last part matters, because a fix that restarts everything and drops traffic is not a fix in production.
These are long jobs. Trials average 2.81 million tokens and 41 turns. Failures show up at every stage: wrong diagnosis or localization, incomplete repairs, and unsafe regressions where the agent makes things worse while fixing them.
The title says "the last nine of reliability," and that's the right frame. Agentic SRE is one of the most obvious enterprise buys, and vendors are already selling it. A 64% ceiling on 20 realistic tasks is a useful reality check, and the failure taxonomy is a checklist for anyone evaluating one of those products: ask how often it makes the incident worse, not just how often it closes the ticket.
Link: arxiv.org/abs/2610.00648
Comments