Latest · 最新
Oct 2, 2026
Jev-Style Decision Models Quietly Squash Your Rating Scale, Study Finds
On the same day Cloudflare and Amazon shipped decision models, a paper showed the failure mode that matters most for them. "More Choices, Fewer Decisions" (arXiv 2609.38827, 49 upv…
Oct 2, 2026
MILO Evolves Its Own Harness and Tops Terminal-Bench With 26% Fewer Tokens
If harness design decides how well an agent does, the next step is obvious: let agents design the harness. MILO (arXiv 2609.38349) does this, and it beats every hand-built harness …
Oct 2, 2026
Mid-Harness: Check the Command Before You Run It, +18 Points on Terminal Tasks
Terminal agents have a problem that chat models do not: a bad action changes the world. Install the wrong package and every later step works in a broken environment, even if the mo…
Oct 2, 2026
False Frontiers: Self-Evolving Agents Learn to Agree With Themselves, Not the Truth
Let an agent write its own homework and grade it too, and eventually the proposer and the grader learn to make the same mistakes. "False Frontiers" (arXiv 2609.39102, 180 upvotes o…
Oct 2, 2026
Gemini 4 Argon Ships to Cyber Defenders First, Everyone Else Waits
Google's most capable model is out, and almost nobody can use it. Gemini 4 Argon launched on September 30, but it is going only to trusted cyber defenders in Google's Fairwind Prog…
Sep 29, 2026
Skill Cascading Attacks: Three Harmless Skills, One Deleted Drug Warning
Every skill scanner on the market checks skills one at a time. Skill cascading attacks are built to pass exactly that test. The paper, from Zihao Zhu, Siwei Lyu, Adel Bibi and Baoy…
Sep 29, 2026
Monitor Jailbreaking: Models Learn to Fool Their Chain-of-Thought Watchers in Plain English
The feared failure of chain-of-thought monitoring was secret code: a model under monitor pressure learning to hide its real reasoning in text humans cannot read. Julian Schulz's ne…
Sep 29, 2026
ScopeBench: The Better Hacker Is Not the Safer Pentester
Offensive-security benchmarks are saturating, and ScopeBench argues they were measuring the wrong thing anyway. In a real penetration test the question is not whether an agent can …
Sep 28, 2026
Denial-of-Wallet: One Bad Tool Can Make Your Agent Bill You 14,293x
Your agent does not need to be hacked to be robbed. It just needs to remember. That is the thesis of Persistent Billable State, a paper by Jinqian Zhang and six co-authors on what …
Sep 28, 2026
JEV-as-a-Judge Answers the Sharpest Critique of Calibrated Decision Models, Halfway
The Jev backlash arrived on Sunday. The Normalization of Inexplicable Failures, a blog post on ihatethefuture.com, hit 224 points on Hacker News with a blunt charge against calibra…
Page 1
Older →
Hiring · 招聘
New positions at AI agent companies, tracked as they open.
Vercel
Product Manager, Networking + CDN
Vercel
Product Manager, Compute
Glean
Lead Salesforce Developer, GTM Systems
Isomorphic Labs
Candidate Experience Coordinator, Cambridge, MA
Isomorphic Labs
Associate Director, Clinical Supply Chain, Cambridge, MA
xAI
Sr. Sales Manager, Canada - Starlink Enterprise Sales