Latest · 最新
Aug 16, 2026
PlayWorld makes an agent play the world model, because fixed action scripts can't compare models
There's a measurement bug at the center of world model evaluation and PlayWorld is the first benchmark to take it seriously. If you evaluate a world model by feeding it a fixed act…
Aug 15, 2026
620 comments arguing that Opus 5 is smarter and worse to work with
The third-biggest story on Hacker News today is a blog post titled "Why does Opus 5 feel worse to work with?" — 671 points and 620 comments in eleven hours. The author is upfront t…
Aug 15, 2026
AutoDesign: a meta-optimizer rewrites the agent's harness, 40 minutes and under $3
Meituan's AutoDesign is the second harness-evolution paper in as many days, and it takes a different route to the same conclusion. A meta-harness optimizer steers a code agent to r…
Aug 14, 2026
OpenART: 85% attack success, and the harness is the vulnerability
OpenART is a red teaming framework for agents that starts from a premise most safety evals ignore: agents live in persistent environments, and a state change made in turn three can…
Aug 13, 2026
Grok 4.6 caught GPT-5.6 Sol, and it did it by training on agent tasks
xAI shipped Grok 4.6 yesterday and it scores 61 on the Artificial Analysis Intelligence Index, dead level with GPT-5.6 Sol. CursorBench 3.2 at 69.9%, DeepSWE 1.1 at 65.9%, better t…
Aug 12, 2026
Sixty percent of the SWE-bench tasks nobody solves have broken tests
Buried in the SWE-Bench ProMax paper is a finding that should change how you read every coding-agent leaderboard: nearly 60 percent of the unsolved instances in SWE-bench Verified …
Aug 11, 2026
oqoqo makes you your own benchmark
oqoqo took the Product Hunt top spot yesterday with 304 upvotes, selling something the field has been complaining about for two years: your own private benchmark, on your own tasks…
Aug 10, 2026
Harvey LAB: Legal Agents Get a Real Exam, and Kimi K3 Tops It
The most valuable legal AI company open-sourced its own exam, and a Chinese open-weights model just aced it. Harvey's LAB (Legal Agent Benchmark) hit GitHub trending today: 1,671 t…
Aug 9, 2026
OSReward: The AI Judges Grading Agents Are Too Soft
OSReward measured something the whole computer-use agent field has been assuming works: AI judges grading agent runs. It found they don't. The HKU OS-Copilot team built the most co…
Aug 6, 2026
MerchantBench: A Year of Shopkeeping Breaks Every Agent
The top agent paper on Hugging Face Daily Papers this round (83 upvotes) is MerchantBench (arxiv.org/abs/2607.28956), and its result deserves the attention: run LLM agents as e-com…
Page 1
Older →
Hiring · 招聘
New positions at AI agent companies, tracked as they open.
xAI
Porter Supervisor
xAI
Hospitality Supervisor
xAI
Executive Sous Chef
xAI
Executive Chef
xAI
Cook
xAI
Barista