October 3, 2026ResearchAgents

PoS: Stop Giving Agents Memory, Give Them a Belief State

Most long-horizon agents treat context management as a storage problem: keep the history, compress it when it gets long, retrieve the right bits. PoS (Progression of States), the top agent paper on today's Hugging Face board at 68 upvotes, argues that is the wrong question. Remembering what happened is not the same as understanding where you are.

PoS builds and keeps updating an explicit belief state that becomes the agent's decision context. Each belief holds two things: the agent's current estimate of the world, and the task requirements that are still unresolved. Simply put, "here's what I think is true" plus "here's what I still need to find out or do." The framework checks that belief for internal consistency and watches progress to catch what the authors call Belief Trapping, where the agent keeps taking actions without getting any closer to the goal. When it detects a trap, the recovery is tailored to both the pattern of the trap and the type of requirement that's still open.

On four benchmarks covering execution and diagnosis tasks, PoS gets the best overall score on every one, with all three LLM backbones tested. Ablations show consistency validation and recovery do the heavy lifting, and context-scaling runs show it holds up as context grows.

The interesting part is the framing. Belief Trapping is the failure every agent builder has watched: forty turns of busy, confident, useless tool calls. Memory systems don't catch it because nothing is missing from memory. A running list of "what's still unresolved" is a cheap way to make that stall visible, and it's an idea you can bolt onto your own harness today. Code at github.com/luoyu100/PoS.

Link: arxiv.org/abs/2610.01415
← Previous
Your Agent Leaderboard Is Ranking Harnesses, Not Models
Next β†’
ActiveSaddler: Harness Optimizers Were Training on the Wrong Problems
← Back to all articles

Comments

Loading...
>_