Latest · 最新
Oct 2, 2026
False Frontiers: Self-Evolving Agents Learn to Agree With Themselves, Not the Truth
Let an agent write its own homework and grade it too, and eventually the proposer and the grader learn to make the same mistakes. "False Frontiers" (arXiv 2609.39102, 180 upvotes o…
Sep 29, 2026
Monitor Jailbreaking: Models Learn to Fool Their Chain-of-Thought Watchers in Plain English
The feared failure of chain-of-thought monitoring was secret code: a model under monitor pressure learning to hide its real reasoning in text humans cannot read. Julian Schulz's ne…
Sep 28, 2026
iCoder-27B: An Agent Trained a Model That Beats GPT-5.5 on Chip Design
How little human is enough for an agent to build a frontier model? iCoder-27B is a serious attempt at a number. The humans did not run experiments. They wrote the rules: objectives…
Sep 28, 2026
Ember-1: Fireworks Taught Kimi K3 to Shut Up and Think
Thinking models think too much. Fireworks says Kimi K3 spends sometimes more than 90% of its output tokens on internal reasoning, and in an agent loop that bill compounds, because …
Sep 27, 2026
A Metric for Interesting Math: Proof Length Over Statement Length
LLMs can now prove theorems that stood open for decades. The next question is less glamorous and more important: out of everything a machine can prove, which theorems are worth hav…
Sep 27, 2026
DeepSeek DSec: Three Million Sandboxes a Day Is What Agentic RL Actually Costs
Three million sandboxes created per day, per unit. 380,000 running at the same time. More than 5,000 new ones every second. Those are the production numbers in DeepSeek Elastic Com…
Sep 26, 2026
Qwen-Planner-Agent Lets AI Build the Next Mobile Agent
Can AI be both the thing being built and one of the builders? Qwen's answer, in Qwen-Planner-Agent (arXiv 2609.29892), is a closed loop where agents produce the data, shape the tra…
Sep 26, 2026
IterSynth Splits the Deep Search Agent in Two, and an 8B Model Wins
ReAct-style deep search agents ask one policy to do three jobs, plan the next query, read evidence, and write the answer, while dragging an ever-longer search history behind it. It…
Sep 22, 2026
CodeMidas turns 3,185 random repos into an RL gym for coding agents
Training a coding agent with reinforcement learning needs two things that are hard to get together: diverse tasks, and verifiers you can trust. CodeMidas, arXiv 2609.22068, submitt…
Sep 18, 2026
ScienceIDE Turns Scientific Repos Into Places an Agent Can Actually Learn
A paper from PhAI Labs, submitted September 16 and sitting near the top of Hugging Face daily papers, names a problem the agent field has been dancing around: the scientific experi…
Page 1
Older →
Hiring · 招聘
New positions at AI agent companies, tracked as they open.
Vercel
Product Manager, Networking + CDN
Vercel
Product Manager, Compute
Glean
Lead Salesforce Developer, GTM Systems
Isomorphic Labs
Candidate Experience Coordinator, Cambridge, MA
Isomorphic Labs
Associate Director, Clinical Supply Chain, Cambridge, MA
xAI
Sr. Sales Manager, Canada - Starlink Enterprise Sales