Latest · 最新
Aug 11, 2026
Why RL trains many skills at once and SFT can't
Top paper on Hugging Face today, and it answers a question anyone training an agent on more than one task has hit face-first: why does supervised fine-tuning on a mixed task set ma…
Aug 10, 2026
EnvACE: Tencent Teaches Agents to Rehearse the World in Their Heads
The most expensive part of agent RL is not the GPU, it is the environment — real APIs, real sandboxes, real latency, real flakiness. EnvACE, a new paper from a Tencent-affiliated t…
Aug 8, 2026
AgentOPSD Finds the Three Turns That Actually Mattered
The top agent paper on Hugging Face today (66 upvotes) is AgentOPSD: Recursive Self-Distillation for Agentic Reinforcement Learning, from a Tsinghua-led team (arxiv.org/abs/2608.05…
Aug 7, 2026
ABSeeker: A 4B Search Agent That Fights Like a 30B
Top agent paper on Hugging Face Daily Papers for August 6, from Shanghai Jiao Tong University: ABSeeker, Training Long-Horizon Search Agents via Answer-Backtracked Credit Assignmen…
Aug 6, 2026
Prime Agent Rewrites Its Own Harness While Running
Prime Intellect released Prime Agent on August 5 (primeintellect.ai/blog/prime-agent): an open-source coding and research agent built around one idea — the agent should be allowed …
Aug 4, 2026
RLSVR: A Spy Game Instead of a Reward Model
Top paper on Hugging Face daily papers with 143 upvotes: From RLVR to RLSVR, Task Transformation Induces Self-Verifiable Rewards for Open-Ended LLM Self-Improvement. The problem is…
Jul 31, 2026
SkillRise: Agents That Grow Skills Instead of Relearning
A new paper called SkillRise (arXiv 2607.26784) goes after one of the more honest weaknesses of today's agents: they relearn everything from scratch. Solve a task, throw away what …
Jul 28, 2026
NVIDIA Molt: an RL framework small enough for the agent to read itself
NVIDIA's NeMo team dropped Molt today and it shot to the top of Hugging Face papers with 647 upvotes. It's a PyTorch-native training framework for agentic reinforcement learning, a…
Jul 22, 2026
DeepSearch-World: A 9B Search Agent That Taught Itself, No Teacher
The default recipe for a good search agent right now is: take a small model, distill trajectories out of a frontier model, ship it. HKUST's DeepSearch-World, arXiv 2607.07820, does…
Jul 19, 2026
LongStraw: everyone serves 1M context, almost nobody can train on it
Here's an embarrassing asymmetry in the agent stack: inference systems happily serve million-token contexts, but RL post-training — the thing that actually makes agents better at l…
Page 1
Older →
Hiring · 招聘
New positions at AI agent companies, tracked as they open.
xAI
Porter Supervisor
xAI
Hospitality Supervisor
xAI
Executive Sous Chef
xAI
Executive Chef
xAI
Cook
xAI
Barista