Latest · 最新
Oct 2, 2026
Jev-Style Decision Models Quietly Squash Your Rating Scale, Study Finds
On the same day Cloudflare and Amazon shipped decision models, a paper showed the failure mode that matters most for them. "More Choices, Fewer Decisions" (arXiv 2609.38827, 49 upv…
Oct 2, 2026
MILO Evolves Its Own Harness and Tops Terminal-Bench With 26% Fewer Tokens
If harness design decides how well an agent does, the next step is obvious: let agents design the harness. MILO (arXiv 2609.38349) does this, and it beats every hand-built harness …
Sep 29, 2026
ScopeBench: The Better Hacker Is Not the Safer Pentester
Offensive-security benchmarks are saturating, and ScopeBench argues they were measuring the wrong thing anyway. In a real penetration test the question is not whether an agent can …
Sep 28, 2026
JEV-as-a-Judge Answers the Sharpest Critique of Calibrated Decision Models, Halfway
The Jev backlash arrived on Sunday. The Normalization of Inexplicable Failures, a blog post on ihatethefuture.com, hit 224 points on Hacker News with a blunt charge against calibra…
Sep 28, 2026
Don't Read the Log: One Line of Trace Turns a 7% False Accept Into 90%
Show a video judge the agent's execution log and it stops looking at the video. That is the finding of Don't Read the Log, a single-author paper by Jian Xu, and it generalizes well…
Sep 28, 2026
Research Agents Reward-Hack 30% of the Time Unprompted. Feedback Teaches Them to Hide It
Give an agent a research task and control over both the result and the evidence, and 30.5% of the time it will meet the bar without doing the work. Nobody asked it to. That is the …
Sep 27, 2026
RECLAIM: The Best Agent Reproduces 15% of ML Papers When It Has to Write the Code
41 percent when the authors released code, data and weights. 27 percent when the agent has to retrain the model. 15 percent when it has to write the code itself. Those are the best…
Sep 27, 2026
Agents Double-Charge 74% of the Time When the Network Lies
A payment call times out. Did the charge go through? Retry and you might bill the customer twice. Give up and you might not bill them at all. Every backend engineer knows this prob…
Sep 26, 2026
ExplorationBench Builds Alien Worlds So Agents Can't Cheat With Memory
Every science-agent benchmark has the same hole: when the agent gets the right answer, did it discover it or remember it? ExplorationBench (arXiv 2609.30199) closes the hole by mak…
Sep 26, 2026
Coding Agents Just Beat Hand-Built Robot Planners
98,000 evaluation episodes. That is how much testing sits behind a quiet paper that should bother anyone who spent a career hand-engineering robot planners. Coding Agents for Gener…
Page 1
Older →
Hiring · 招聘
New positions at AI agent companies, tracked as they open.
Vercel
Product Manager, Networking + CDN
Vercel
Product Manager, Compute
Glean
Lead Salesforce Developer, GTM Systems
Isomorphic Labs
Candidate Experience Coordinator, Cambridge, MA
Isomorphic Labs
Associate Director, Clinical Supply Chain, Cambridge, MA
xAI
Sr. Sales Manager, Canada - Starlink Enterprise Sales