Latest · 最新
Aug 16, 2026
Spatial Memory Agent: frozen weights, no tools, and it still beats fine-tuning
Third paper in three days concluding the same thing from a different angle: capability lives at the system layer now, not in the weights. Spatial Memory Agent takes a frozen vision…
Aug 16, 2026
PlayWorld makes an agent play the world model, because fixed action scripts can't compare models
There's a measurement bug at the center of world model evaluation and PlayWorld is the first benchmark to take it seriously. If you evaluate a world model by feeding it a fixed act…
Aug 16, 2026
Alaya-EVOKE takes the memory out of the model, and the cost curve goes flat
Interactive world models have a boring, fatal problem: the longer you interact, the more history there is, and if that history lives in the model's working memory then cost grows q…
Aug 16, 2026
CLI-Anything: stop teaching agents to click, generate them a command line instead
Agents are bad at professional software, and everyone has quietly agreed to work around it in three unsatisfying ways. Drive the GUI with vision, which breaks the moment a button m…
Aug 16, 2026
A 232x kernel from 1,500 submissions: what loop engineering actually looks like
Sankalp Shubham entered GPU Mode's qr_v2 problem — batched square compact-Householder QR factorization, FP32 CUDA on B200 — and came out with a 232x speedup over baseline. Geomean …
Aug 16, 2026
Claude now watermarks everything it writes, and code is the part it can't touch
Anthropic published the mechanics of Claude's text watermark on August 14, and the interesting part isn't the watermark. It's where the watermark fails.
Here's the method. Claude …
Aug 15, 2026
620 comments arguing that Opus 5 is smarter and worse to work with
The third-biggest story on Hacker News today is a blog post titled "Why does Opus 5 feel worse to work with?" — 671 points and 620 comments in eleven hours. The author is upfront t…
Aug 15, 2026
Mole caps your research agent's spend before it runs, and overshot by zero
Show HN today: Mole, a deep-research agent that lives in your terminal. Written in Go, Apache-2.0, static binaries for Linux and macOS, works against Anthropic, OpenAI, or any Open…
Aug 15, 2026
spec-kit is at 128k stars and shipping a release almost every day
GitHub's spec-kit picked up 1,147 stars in a day on its way past 128,000, which is remarkable for a repo whose entire proposition is that you should write the spec first. The pitch…
Aug 15, 2026
Anthropic published the token math on Claude Code, and /clear is the cheapest habit you have
Anthropic put out a post on August 14 breaking down what actually costs money in a Claude Code session, and it hit Hacker News fast because it answers a question every heavy user h…
Page 1
Older →
Hiring · 招聘
New positions at AI agent companies, tracked as they open.
xAI
Porter Supervisor
xAI
Hospitality Supervisor
xAI
Executive Sous Chef
xAI
Executive Chef
xAI
Cook
xAI
Barista