October 5, 2026loop

Loop Daily: 2026-10-05

The weekend's loops split cleanly into two camps. One camp ran them on hard targets with an external judge: an autoresearch loop that wrote a formally verified Ethereum encoding library faster than the Go original while a separate auditor fleet hunted for proof gaps until it gave up, and a Meta pipeline where Claude Code and Codex review each other's research changes and beat either product alone by 17 points. The other camp argued about what the loop even is. Alexandr Wang reduced it to markdown files, cron jobs and one metric, OpenRouter's Alex Atallah called the whole agent loop stack table stakes, and several builders pushed back that the moat is the boring part nobody demos: stop conditions, budgets, permissions and recovery. Outside code, loops reached a 3D printer, a genetics-inspired search for circle packing and 36 physics manuscripts that still needed human experts to sign off.
πŸ’‘#1
@GiulioRebuffo
https://x.com/GiulioRebuffo/status/2106866511440724276
GiulioRebuffo spent about 12 days building a formally verified implementation of SSZ, Ethereum's consensus-layer encoding, in the Bend language, generating data structures from YAML along with proofs for their setters, encoders, decoders and hashers. The initial spec and generators were written by hand with AI help; implementation and optimization were then handed to Karpathy-style autoresearch loops. A third stage wired auditor agents running GLM-5.3 to mutate the implementation and check whether the proofs caught it, and after eight straight hours they could no longer find gaps and gave up on the codec paths. The result is faster than Go's fastssz on encoding and decoding (hash_tree_root is still 5x slower), and it runs inside the Prysm client via FFI.
πŸ’‘#2
@dair_ai
https://x.com/dair_ai/status/2106529236676976773
dair_ai highlighted RankEvolve, a Meta paper on making auto-research agents reliable when a single silent bug, like leaked evaluation data or a disconnected gradient, can invalidate hours of training and every iteration built on top. It enforces each research phase and gate through a compiled protocol and runs Claude Code and Codex as separate nodes that review and repair each other's changes. At a matched budget, combining the two raises execution accuracy from 45.8% for the best single product to 62.5%. Over twelve iterations on the open-source HSTU recommender it improved NDCG@10 on MovieLens-20M by 4.48% over the published result.
πŸ’‘#3
@mafolebaraka
https://x.com/mafolebaraka/status/2106851976151781584
mafolebaraka transcribed the part of Alexandr Wang's Y Combinator talk describing how Wang's own team runs agents: markdown files, cron jobs and identifying the metric. A plain text file holds the goal and constraints, a scheduled job wakes the agents, and a clearly defined success measure with evals catches drift; run it, review, refine the instructions, run again. Wang's claim is that the right loop lets a swarm outdo a team of 100 engineers, and that the hard part is no longer code but the metric, because a swarm will efficiently hit the wrong target if success is poorly defined. The writeup notes the rest of the talk is a Meta pitch, but this section is an operating description.
πŸ’‘#4
@abishekv
https://x.com/abishekv/status/2106769492277961036
abishekv argued that the code around an agent should define the environment (context, memory, tools, permissions, guardrails, evals) while the model decides how to reason, because encoding the reasoning path makes the system inherit its author's blind spots. The example is Karpathy's autoresearch, where the agent was given a goal and an eval and discovered that halving the batch size improved results, because it could take more optimization steps inside the fixed compute budget. Nobody told it to look there. The missing piece, in this view, is simulation, which would give agents somewhere safe to find what they were never told to look for.
πŸ’‘#5
@XSeyvion
https://x.com/XSeyvion/status/2106504630301560972
XSeyvion pointed to Matthew Schwartz's report that the open-source BootLoops harness plus Claude produced 36 manuscripts across 18 fields, pushing models through precise scientific calculations. Human experts still had to step in before the results became scientifically valuable. The division of labor is the takeaway: agents draft and compute, experts validate and sign off.
πŸ’‘#6
@kannappan
https://x.com/kannappan/status/2106827786396541196
kannappan built an autoresearch evolution framework at the London AI Science hackathon that applies Mendelian genetics to code synthesis, treating candidate programs as organisms whose traits are inherited and recombined. It was pointed at two classic search problems: circle packing and the no-five-on-a-sphere problem. It is a small example of the trend of swapping the keep-or-discard step for a population-based search.
πŸ’‘#7
@MWsatware
https://x.com/MWsatware/status/2106804840898970086
MWsatware runs every step of converting 3D objects for a 3D printer through an agentic loop with vision attached and a direct connection to the printer. The output included planetary gears and a gearbox that, by the author's research, did not exist in that form before a few weeks ago. The model was trained to interpret tool results and create from them, and the next experiment is knitting a pullover.
πŸ’‘#8
@a16z
https://x.com/a16z/status/2106444310187290762
a16z clipped OpenRouter co-founder Alex Atallah answering the complaint that everyone is building the same thing: an agent loop with notifications, connectors, context management, memory, sandboxes, agentic web search and an always-on agent on top. Atallah's answer is that these are the new table-stakes primitives, like the database, users table and login page every 2005 web app needed, and that real differentiation sits above them. With nearly half a million views, it reframed the agent loop from a product into infrastructure and the replies immediately asked who will absorb it into frameworks the way 2005's web stack was absorbed.
πŸ’‘#9
@AGTPinsights
https://x.com/AGTPinsights/status/2106515522305302920
AGTPinsights summarized Atallah's longer argument on the same podcast: ten specialized agents beat one universal agent, because with one job each you can see exactly where an agent is failing, while a universal agent hides failures in cross-domain joins. The proposed shape is vertically focused sub-agents with their own quality checks and a chief-of-staff agent coordinating them. Atallah also named decision models like Jev as a possible alignment layer that checks a tool call or an agent-to-agent message. Harrison Chase said the LangChain team is arguing the same question internally.
πŸ’‘#10
@free_ai_guides
https://x.com/free_ai_guides/status/2106398901985317191
free_ai_guides laid out a stop-condition-first approach to loops: an agent that keeps looping without a definition of done is not persistent, just expensive. Done is defined by four things up front (goal, evidence, constraints, context), and the cycle of plan, execute, verify, learn repeats until it hits a real gate. If verify decides the task is impossible, it escalates instead of looping forever; subjective checks go to an independent grader; and a cap on passes and cost sits under the whole thing.
πŸ’‘#11
@rai_anshu81786
https://x.com/rai_anshu81786/status/2106629074106081385
rai_anshu81786 read Meta's proactive memory management paper, which addresses the way memory and state decay over long-horizon agent tasks. The approach runs a background cron job over memory and lets it intervene in the agentic loop, rather than waiting for the agent to retrieve what it needs. The paper reports improved accuracy on long-horizon tasks, and it fits a pattern of moving memory from a passive store to an active participant in the loop.
πŸ’‘#12
@dinodaizovi
https://x.com/dinodaizovi/status/2106717462465065321
dinodaizovi proposed decomposing the monolithic harness into three parts: the agentic loop, which only needs the inference API; session data, which owns the context window and carries an identity based on the provenance of the tokens appended to it; and tools, triggered by tool-use requests in context. Tools would then be authorized against that session identity and isolated individually by what they do. It is a security-first reading of the loop, where permissions attach to provenance rather than to the process.
πŸ’‘#13
@kbsingh
https://x.com/kbsingh/status/2106752689161851029
kbsingh hit a failure in OpenCode with DeepSeek 4.1 Flash: during tool calls the model broke out of the tool context and emitted raw DSML tags into the output area, which terminated the whole agentic loop without updating the context. Asking it to continue where it left off did not work, because the model had no working loop memory and fell back to the prompt before last. A reply described the same symptom with other quantized models, where tool calls get garbled after a while and loops run for hours, which is a reminder that tool-call formatting is the weak link in long local loops.
πŸ’‘#14
@cristian_le0
https://x.com/cristian_le0/status/2106817965706658202
cristian_le0 recommends that anyone running auto research add a routine where the agent interviews the human with a few questions to validate assumptions and evaluation criteria before it starts. The point is to steer the auto researcher in the right direction early and avoid compounding errors over many iterations. It is the research-loop version of the spec interview that coding users now run before handing off a task.
πŸ’‘#15
@0xtomdaniel
https://x.com/0xtomdaniel/status/2106551884739936624
0xtomdaniel says the real unlock for autonomy came from realizing that escalations from agents to the human could be resolved by another agent. Only then did the system actually become autonomous rather than a queue of pings. The author plans to fold a newly released skill into a Karpathy-style auto research loop next.
πŸ’‘#16
@BTCMax21
https://x.com/BTCMax21/status/2106336680701776302
BTCMax21 listed what makes a fully recursive agentic loop hard to keep running in practice. One bot keeps asking for information it already has, like a phone number, because secrets are stored per bot instead of centrally. Agents also get confused when websites move controls since they last looked. These are mundane problems, but they are the ones that break loops after the demo.
πŸ’‘#17
@undefinedKi
https://x.com/undefinedKi/status/2106775535787323649
undefinedKi's explainer of Karpathy's autoresearch, now at 92,000 stars, highlights the details people tend to skip. The agent is told never to stop and never ask whether to keep going, because the human might be asleep; the human only edits program.md, which Karpathy calls the code of the research org; and a simplicity rule throws out tiny gains that add 20 lines of hacky code while keeping equal scores with less code. The author's generalization: anything with one number to beat, one file to edit and a fixed time per run can grind overnight this way.
πŸ’‘#18
@chelsearustrum
https://x.com/chelsearustrum/status/2106474396433223698
chelsearustrum shared what was learned building a generate, evaluate, repair loop: the evaluator checks every rule, and the repair step receives those evaluations and tries to fix them. Practical lessons include hygiene, explicit structure for roles, voice and output, balanced reasoning, and not overcorrecting with NEVER rules, plus the heuristic that if a human can understand an instruction, the model probably can too. The loop also turned into a cost lever, letting the author choose between cheaper models with more passes and stronger models with fewer.
πŸ’‘#19
@nickvernij
https://x.com/nickvernij/status/2106755793764721008
nickvernij is building autoresearch loops that look for opportunities where specialized models can beat frontier models on domain tasks at a better price. The loop's target is not a single benchmark but a market gap: find a domain where a smaller trained model wins, then train it. It came in reply to the same podcast where Amjad Masad suggested models could train their own smaller domain replacements on the fly.
πŸ’‘#20
@Unplugged_kk
https://x.com/Unplugged_kk/status/2106746215345922457
Unplugged_kk argues autonomous ML research needs a budget limit before it needs more GPUs. The list is short: cap GPU hours, pin code and data, save artifacts, and stop when eval gains stall. Without a trusted eval harness, the post says, you are just automating noise.
πŸ’‘#21
@michael_chomsky
https://x.com/michael_chomsky/status/2106574123480551598
michael_chomsky takes the opposite position: pre-made training APIs are becoming antiquated in the era of autoresearch, and even convenient managed fine-tuning is too constrained. The proposal is to give models raw GPUs, let them spin up 8, 12 or more on zero notice and choose the base model themselves. It frames the debate this week between loops with tight budgets and loops with open-ended compute.
πŸ’‘#22
@shaundevan_
https://x.com/shaundevan_/status/2106773711160242630
shaundevan_ uses auto-research personally but argues production agents, especially in the enterprise, should not have free rein. In tricky production environments, a system that follows the reasoning its author encoded is the point, because you know how it will behave. It is a useful counterweight to the let-the-agent-find-it view: exploration loops and production loops want different degrees of freedom.
πŸ’‘#23
@Ja5mineEgg
https://x.com/Ja5mineEgg/status/2106597559523229759
Ja5mineEgg has given Claude tasks where it estimated the work would take two to three months, and then an hour and a long agentic loop later the result was flawless. The anecdote shows how badly models still estimate their own throughput: the estimate is borrowed from human timelines, while the loop compresses the work.
πŸ’‘#24
@de1lymoon
https://x.com/de1lymoon/status/2106745564410867933
de1lymoon walked through OpenAI's new hosted agent sessions, where one sessions.create call defines the task, model, tools and environment and OpenAI runs the harness. The runtime can be OpenAI-hosted (the same sandbox as Codex and ChatGPT), self-hosted, or none; tool search loads definitions on demand; subagents run up to six in parallel; vaults carry credentials to MCP tools; and older context is compacted automatically with lessons saved to MEMORY.md. The pitch is that the agent loop itself becomes a managed service, leaving the developer to choose tools and environment.
πŸ’‘#25
@Cisco_research
https://x.com/Cisco_research/status/2106791644620054579
Cisco_research flagged that DeepSeek Harness is now in public preview worldwide as an open-source agent runtime where everything is a plugin, including the agent loop itself. A creator mode can write and install a plugin from a chat, on desktop or web. Together with Claude Code mods and OpenAI's hosted sessions, it means all three camps now expose the loop as something you can swap or customize.
πŸ’‘#26
@MarcosRGjr
https://x.com/MarcosRGjr/status/2106492451175239965
MarcosRGjr read Claude Code mods from an operations angle: small TypeScript functions that hook the agent loop to rewrite prompts, block or retry tool calls, approve or deny permissions, redact secrets from tool output and swap UI, with built-ins like /diff now replaceable. Team and Enterprise load a security default first so user mods cannot override permission deny rules. The conclusion is that the agent runtime just became a programmable platform and also an unsandboxed plugin surface, so extensibility without trust boundaries is the new attack surface for agent fleets.
πŸ’‘#27
@techNmak
https://x.com/techNmak/status/2106765347793858563
techNmak published a technical handbook on harness engineering, built on the argument that a capable model by itself is not an agent: something has to decide what context the model sees, which tools it can use, what state survives between steps, which actions need approval, when to retry, when to stop and how to verify success. It covers the agent loop, agent-computer interfaces, MCP, skills, compaction, sandboxes, credentials, prompt injection, execution budgets, evaluator loops, idempotency and long-running recovery. Its sharpest point is that more tools or scaffolding can make an agent worse if they add ambiguity a newer model no longer needs.
πŸ’‘#28
@sengorakualpha
https://x.com/sengorakualpha/status/2106300230064914510
sengorakualpha is shipping microloop next, a supervised agent loop for local models with hard-coded guardrails so that risky actions require the user's OK. It follows docwrite, a README generator, from the same builder. Local-model loops with explicit approval gates are a recurring small-builder project this month.
πŸ’‘#29
@CodyBontecou
https://x.com/CodyBontecou/status/2106432192947875972
CodyBontecou built a proof of concept seven months ago that runs an on-device model through llama.cpp and reads and writes the phone's filesystem in an agentic loop. It resurfaced in a thread about mobile agents, as phones become a target for always-on personal agents. The interesting part is that the loop runs entirely on the device, with no server in the middle.
πŸ’‘#30
@leonardoalt
https://x.com/leonardoalt/status/2106677678564577419
leonardoalt noted that verifying Rust code is still extremely difficult and a big open problem, but that Rust-level performance is achievable in Lean through autoresearch. The suggestion is to write in a verifiable language and let an optimization loop close the speed gap, accepting a hit where performance is not critical. It is the same bet GiulioRebuffo made with Bend this week.
πŸ’‘#31
@bhaskarmelkani
https://x.com/bhaskarmelkani/status/2106757637920510432
bhaskarmelkani describes a simple improvement loop for everyday work: people correct AI constantly, but most of those lessons vanish when the chat ends. Turn repeated corrections into reusable skills, test them, and make task N plus one better than task N. It is autoresearch's keep-or-discard rule applied to the user's own feedback instead of a benchmark.
πŸ“‘ Eco Products Radar
Eco Products Radar

Karpathy's autoresearch: still the reference loop, cited in about half the substantive posts, from formal verification to batch-size discoveries.
Claude Code and Codex: increasingly paired inside one loop as author and reviewer, most clearly in Meta's RankEvolve.
OpenRouter and the Atallah podcast: the source of the weekend's argument over table-stakes loops and ten specialized agents versus one.
GPT-6.1 Sol and OpenAI hosted agent sessions: the managed-loop option, pitched as cheap enough to run the loop itself.
Claude Code Mods and DeepSeek Harness: the two plugin layers that now expose the agent loop for rewriting.
← Previous
Super User Daily: 2026-10-05
Next β†’
Ideas Radar: 2026-10-05
← Back to all articles

Comments

Loading...
>_