Latest · 最新
Sep 27, 2026
Most of Your Agent Bill Is Multiple-Choice Questions
Four dollars and seventy-three cents versus two cents. Same eighteen-turn agent loop, same goal score of 0.93, same work done underneath. The only change was who made the decisions…
Sep 21, 2026
The Harness Is the Product. Its Scoreboard Is the Weak Spot.
Eight dollars and seventy-five cents to thirteen dollars and fifty cents an hour. That is what NVIDIA, NTU and MIT say you save by running a coding agent through a better wrapper i…
Sep 20, 2026
Offense Cost $4.65 This Week
Eleven targets. Code execution on every single one. Median four minutes and thirty-eight seconds per box. Total bill for the accepted runs: four dollars and sixty-five cents.
That…
Sep 14, 2026
Your CLAUDE.md Is Written for a Model That No Longer Exists
The biggest self-improvement result of the week was a deletion.
Teknium pointed 110 subagents at the Hermes codebase and let them run for fifteen hours. They made 111,352 tool cal…
Sep 13, 2026
The Check Is the Product
Eleven days.
Claude ran nearly unattended for eleven days and produced a Lean formalization of Fermat's Last Theorem: 13 million lines, 29,511 intermediate theorems, every one of …
Sep 6, 2026
Who Audits the Loop? The Week Verification Became the Product
75% of the failed Claude Code runs said "task completed successfully."
That number comes from Frontier Challenge, a benchmark released this week that asks agents to finish real mu…
Aug 30, 2026
The Harness Ate the Model
A developer posted one sentence this week. He'd pay an absurd amount of money for a single window that shows all his coding agents at once. Sixteen hundred likes, four hundred and …
Aug 23, 2026
Deep Dive: The Memory Layer Is the Real Moat
A year ago GitHub Copilot was the product that changed how code got written. Now nobody even thinks of it as an AI company. Claude Code took the crown as the game changer, and it's…
Aug 19, 2026
DeepSeek's double drop: the harness, not the model, was the main event
On Aug 13, 2026, DeepSeek shipped two things the same day: the V4-Pro model went GA, and it open-sourced its own agent harness (dsh). I pulled the full Twitter propagation of both …
Aug 16, 2026
The model didn't change. The score tripled.
13.3% to 38.3% on ARC-AGI-3. Same weights, same GPT-5.6 Sol, nothing retrained. OpenAI just let it keep its own reasoning between steps instead of throwing it away, and added conte…
Page 1
Older →
Hiring · 招聘
New positions at AI agent companies, tracked as they open.
Vercel
Product Manager, Networking + CDN
Vercel
Product Manager, Compute
Glean
Lead Salesforce Developer, GTM Systems
Isomorphic Labs
Candidate Experience Coordinator, Cambridge, MA
Isomorphic Labs
Associate Director, Clinical Supply Chain, Cambridge, MA
xAI
Sr. Sales Manager, Canada - Starlink Enterprise Sales