October 3, 2026loop

Loop Daily: 2026-10-03

The headline loop of the day ran for three months and produced five million lines of verified Lean. A team of mathematicians with Codex and Claude Code agents, talking through a mailbox they set up, classified all 15,973 semigroups of order 6 and wrote down an explicit basis for every one that has one, something no human had done for most of them. That is the clearest proof yet that auto-research works when the objective is a proof checker that cannot be argued with. The rest of the feed kept circling the same point from different sides: someone measured that humans change direction on 25% of steps where Codex does on 9%, a Jev cookbook found a claimed 193.6x speedup was really 3.6x once measured, and a Meta research team built an auto-research harness where Claude Code and Codex cross-check each other for bugs. Claude Code Mods also arrived, and within a day someone had ported pi-autoresearch into it, so the try-benchmark-keep-or-revert loop now runs inside the most common coding agent.
πŸ’‘#1
@nasqret
https://x.com/nasqret/status/2106136282061574579
nasqret announced that after three months of nonstop work by several AI agents, Codex and Claude Code, the team has classified all 15,973 semigroups of order 6. Every one of the 15,969 that has a finite basis now has it written down explicitly, the remaining 4 are proved to have none, and for most of them mathematicians had previously only known that a basis exists. The Lean formalization, orchestrated with a multi-agent approach, passed five million lines of verified code, one of the largest auto-formalization projects to date, and the paper was accepted at the MATH-AI workshop at NeurIPS 2026. The setup used agents communicating through a mailbox plus a custom bootstrapping and auto-research method; the last 390 semigroups took 23 more days and the very last needed a full structural analysis before its proof could even be written. In a follow-up the author adds that every project needs its own setup, and that tuning it is a skill requiring real domain expertise.
πŸ’‘#2
@_reachsumit
https://x.com/_reachsumit/status/2105516947953926452
_reachsumit shared RankEvolve from Meta, a multi-agent auto-research harness for evolving a generative recommender. The design choice worth noticing is that coding agents like Claude Code and Codex cross-check each other to catch bugs, rather than one agent grading its own changes. Recommender systems are an ideal target for this kind of loop: a clear offline metric, lots of small architectural knobs, and expensive human iteration. Cross-family review inside an auto-research loop is quickly becoming the standard answer to agents fooling themselves.
πŸ’‘#3
@LacorteMichele
https://x.com/LacorteMichele/status/2105678957404361071
LacorteMichele pointed to TraceML, which put coding agents on the same Kaggle contests as humans and logged how they search. People changed direction on 25% of steps; OpenAI Codex did on 9%, and went back to an abandoned version only once. The practical lesson is to give your agent loop an explicit way to reopen old branches, or it just keeps tuning the last one. It is a good quantitative description of the most common failure in overnight optimization runs: local hill-climbing that never backtracks.
πŸ’‘#4
@fluixoo
https://x.com/fluixoo/status/2105574912773509213
fluixoo collected five public Jev tests that each come with numbers and limitations, and the most useful one deflates a hype claim: on 791 labeled decisions in an automation harness, the measured speed gap was 3.6x, not 193.6x. The autoresearch cookbook turned 2,000 wine reviews into 67 numeric columns and moved CatBoost from 3.09 RMSE to 1.77, batching 13 decisions over one 53,777-character document cost $0.000497 versus $0.006090 separately, and skill suggestion across 182 skills cut wrong loads from 16.8% to 7.3% while breaking seven previously correct requests. The recommended stack is retrieve with code, let the small model answer bounded questions, inspect confidence, escalate uncertainty, log corrections and keep irreversible actions behind policy. The point is that the interface is small enough to test and replace, which is what a decision component inside a loop should be.
πŸ’‘#5
@rronak_
https://x.com/rronak_/status/2106093884069867883
rronak_ explained how Trajectory onboards new open-source models quickly. They optimize for two metrics, perfect numerics and weight-sync speed, express them as a series of unit tests, and then attack those tests with autoresearch agent pipelines. This is auto-research used on infrastructure rather than on model quality, and it fits the pattern that works: a hard, cheap-to-check pass or fail signal and a lot of tedious porting work in between.
πŸ’‘#6
@_minuteman3
https://x.com/_minuteman3/status/2105758538324684900
_minuteman3 had been using pi-autoresearch from davebcn87 and tobi but moved back to Claude Code for personal projects, so in honor of Mods shipping they ported the workflow into Claude Code. A separate post describes the result: type /autoresearch make sort.js faster and the mod starts looping, trying an idea, benchmarking it, keeping wins and reverting regressions, an almost one-to-one port built on the new mods API. Loops that used to require a separate harness can now live inside the agent people already use daily. The interesting next step is whether mods also get used to enforce the guardrails, like immutable benchmarks, rather than just run the loop.
πŸ’‘#7
@cromwellian
https://x.com/cromwellian/status/2105506077592916013
cromwellian says Apple's M5 Studio Ultra really cooks for training. Since the day before, they had run dozens of training experiments with auto-research, training up to five 0.5B models at a time while simultaneously running CPU-based integration and validation tests on the results, all while using the machine as a normal desktop. A single desktop running five concurrent small-model experiments plus validation is the hardware shape that makes overnight auto-research a personal hobby rather than a lab budget.
πŸ’‘#8
@graykevinb
https://x.com/graykevinb/status/2105841168474943922
graykevinb runs 24/7 autoresearch to optimize their renderer on the ChatGPT Plus plan, not Pro, and says they sometimes hit limits but usually do not. That is a useful data point against the assumption that continuous loops need top-tier plans. Rendering performance is also a good fit, with an objective frame-time metric and a large space of small code changes.
πŸ’‘#9
@eliebakouch
https://x.com/eliebakouch/status/2105866163800637585
eliebakouch, working on autoresearch, says the real bottleneck is understanding the output of the different models the loop produces, and that Opus 5.5 is excellent at helping with that. Their current best format is interactive HTML, though models still struggle to build the right level of abstraction around the plots, and the examples took many iterations without fully satisfying them. It is a rarely discussed cost of auto-research: generating experiments is cheap, but a human still has to read them, and the reporting layer is now the constraint.
πŸ’‘#10
@kaixin_tai
https://x.com/kaixin_tai/status/2105757083442594036
kaixin_tai's favorite talk at Modal Runtime came from the cofounder of Niteshift, who argues agent sandboxes should be treated as a lossy cache, not the source of truth. Persist the essential state, let agents rebuild after failures, and keep the speed of running the harness inside the sandbox, which buys reliability without sacrificing latency. The same talk covered how Listen Labs builds autoresearch agents and hill-climbs on its evals with Niteshift. For long loops, deciding what state must survive a crash is the design question that matters most.
πŸ’‘#11
@stretchcloud
https://x.com/stretchcloud/status/2106069701248028731
stretchcloud reported that DeepSeek's newly packaged Harness beat Claude Code head to head on identical tasks with DeepSeek V4 Pro: 20 of 30 passed versus 19 of 30, at $0.028 per success versus $0.074, because Claude Code burned 7.3 times the tokens, mostly a caching penalty from running off its home platform. Harness is MIT licensed and built on a plugin framework where the model adapter, tool registry, sandbox and even the agent loop are swappable pieces. It launched with about 80 plugins against Claude Code's thousands of community skills, and DeepSeek calls it a developer preview with breaking changes. The read is that model quality converges while switching cost does not, until a free and extensible runtime removes it.
πŸ’‘#12
@kabelsalat_info
https://x.com/kabelsalat_info/status/2105928620057444799
kabelsalat_info ran the DeepSeek Harness v0.1 preview through Ollama in August and found the swappable agent loop the most interesting idea in it. What stopped them was mundane: two retries against an overloaded backend and the run was dead. Their question for v0.2 is whether it got more patient with 503s. Retry policy is the unglamorous part of any long-running loop, and it decides whether an overnight run survives to morning.
πŸ’‘#13
@alxai_
https://x.com/alxai_/status/2106056479736226078
alxai_ kicked off Star Skirmish, an experiment benchmarking auto-research tendencies by having Opus 5.5 in Claude Code and Astra in Codex write C++ bots to play StarCraft: Brood War. For 15 years humans have been writing algorithms to manage whole Brood War armies against each other, so there are human records to compare against. The framing matters: more AI is being given a goal and told to keep going, so understanding each model's strategies and tendencies on open-ended optimization is as important as raw benchmark scores. In a reply they note the focus is code optimization and auto-research, not collaborative play.
πŸ’‘#14
@virtual_rf
https://x.com/virtual_rf/status/2105952819937231111
virtual_rf argues that every self-improving loop, GEPA, autoresearch, reinforcement learning, climbs something, and whoever defines what better means owns the loop. Andon Labs started with one question, can an AI keep a simulated vending machine running for months, and one number, the money it ends with; that became Vending-Bench, then a real shop in Anthropic's office, then an Andon-run store in San Francisco. The flip side is CEO-Bench, where a plain script with no AI beat all 16 frontier models because the score rewarded something nobody intended and the search found it. Their budget advice: a score built from your real work, experts to label what good looks like, and someone whose job is to cheat it before the agents do.
πŸ’‘#15
@RylanSchaeffer
https://x.com/RylanSchaeffer/status/2106082988702552385
RylanSchaeffer took aim at a myth baked into many auto-research setups: that bits per byte is agnostic to the tokenizer and vocabulary size. Nearly every paper using BPB claims it, including Karpathy's autoresearch repo, which calls it a vocab size-independent evaluation metric, and the thread argues this is not true. If the metric the loop is climbing moves with tokenizer choices, an agent that changes the tokenizer can look like it improved the model when it did not. It is exactly the kind of objective-function bug auto-research is best at exploiting.
πŸ’‘#16
@Kizuno18
https://x.com/Kizuno18/status/2106139670564462872
Kizuno18 put the rule bluntly: the red-to-green loop only works if the test file is immutable. If an agent writes both the test and the code, it cheats by relaxing assertions to force green. Their setup locks the test suite upfront as a frozen objective function, lets the agent loop from red to green against exit codes, and forbids test edits. It is the coding-agent version of the same lesson every auto-research team keeps relearning: the agent must not be able to touch the scoreboard.
πŸ’‘#17
@maheshboda006
https://x.com/maheshboda006/status/2106085467976577322
maheshboda006 broke down the Karpathy loop and the part most people miss. The basic loop: give the agent one feature, write checks before building, lock them, let it build, keep successes, undo failures and store every round. The problem is that a normal loop forgets, so the same mistakes return with every new feature; the fix is a second loop that reviews previous results, finds repeated mistakes and rewrites the workflow instructions. In their example a feature passed all checks while the app was not connected to the new database, and the outer loop learned to require wiring every feature into the app in the same round the checks pass.
πŸ’‘#18
@MatthewGunnin
https://x.com/MatthewGunnin/status/2105790614985929023
MatthewGunnin summed up the core risk of self-improving agents in one line: every one is a student grading its own homework. In reply to a reported result, they said the only number they believe is the +4.7 that survived tasks the system had never seen. Held-out evaluation is the whole game for these loops, and it is notable how often the headline number and the held-out number differ.
πŸ’‘#19
@andyzworks
https://x.com/andyzworks/status/2106108844107518240
andyzworks introduced ProgressCompass, which pairs two models that are each weak alone inside an agentic loop. A process reward model scores a step well once it knows which step it is, but cannot work it out from history; a vision-language model is a poor progress estimator but good at understanding a task and its steps. So an Orienter VLM states the current step and expected transition, the frozen PRM scores progress within that step, a Verifier VLM checks the step is really done, and a text-only Navigator runs the loop and keeps the verified plan. Splitting judgment across components with different blind spots is the same idea as cross-family review, applied to robot and UI progress tracking.
πŸ’‘#20
@0xbobaaa
https://x.com/0xbobaaa/status/2105753401187553518
0xbobaaa flagged a cost trap in agent loops: prompt layout is a budget line. Putting one timestamp at the top of the prompt breaks the cache prefix on every call, and that alone turns a $0.37 agent loop into $6.25 on Sol. Anything that changes per call belongs at the end of the prompt, not the start, and this is the kind of bug that never shows up in a quality metric, only on the invoice.
πŸ’‘#21
@MagaShawn
https://x.com/MagaShawn/status/2106031769766207877
MagaShawn describes putting payment inside the agent loop so the agent is the customer, with one hard requirement before USDC moves on Base: spend authority. The owner signs an EIP-191 grant that the facilitator verifies, checking budget left, being under a no-confirm ceiling and not revoked; the seller returns 402, the grant is checked, payment settles and the tools continue. Private keys stay with the payer while small API calls clear without a human click each time, at $0.10 per query and $0.02 per escrow settlement. A signed, revocable, budget-capped grant is the right primitive for loops that spend money unattended.
πŸ’‘#22
@DTXNaidu
https://x.com/DTXNaidu/status/2106143072908087340
DTXNaidu raised the governance consequence of harnesses where even the agent loop is a plugin. The loop stops being internals and becomes a change surface: whoever ships a plugin edit ships new decision rules. If the loop is a dependency, version-pin it, and treat an eval result as valid only for the loop hash it ran against. It is a sharp point now that both DeepSeek Harness and Claude Code Mods let third parties rewrite the loop.
πŸ’‘#23
@iCleanAI
https://x.com/iCleanAI/status/2106139219798409221
iCleanAI relayed company-reported results from TCBS after adopting Kiro: nearly 25% faster delivery from backlog to release. The firm merged developer and QA roles into one engineer who owns code, testing and deployment, and about 70% of AI-generated code passes review on the first attempt, with failures going back into the agent loop. The organizational change is the more interesting half: a review failure is now an input to the loop, not a ticket for a different team.
πŸ’‘#24
@samweinstein
https://x.com/samweinstein/status/2105623684610211961
samweinstein transcribed a talk by Alexandr Wang, who said there is still astronomical opportunity in agentic looping: systems that spend 1,000x or a million x more on tokens to drive an outcome through a continuous feedback loop. Companies are themselves large feedback loops with humans operating each edge, and Wang says Meta has seen internal cases where, with the right agentic loop and the right eval or metric to optimize, a swarm of agents accomplished more than a team of 100 engineers very easily. Mechanically it is mundane: figure out the metric, then skills, markdown files, cron jobs and a goal. The emphasis on the metric over the machinery matches everything else in today's feed.
πŸ’‘#25
@nickvernij
https://x.com/nickvernij/status/2106042278204842247
nickvernij has autoresearch agents continuously reading papers and improving an on-device PII redaction model. It is a small, well-scoped example of the pattern working outside frontier labs: a narrow model, a measurable task, and a loop that turns the literature into experiments without a human in the middle. On-device privacy models are also exactly the kind of niche that would never get a dedicated research team.
πŸ’‘#26
@HyeonggyuC
https://x.com/HyeonggyuC/status/2105544941162098966
HyeonggyuC offered the first categorization of auto-research approaches into two families. Solution-driven systems treat implemented code as the primary unit of search, get strong optimization from direct evaluation feedback and have low interpretability. Idea-driven systems search over hypotheses, are more interpretable and explore more broadly, but how to manage the ideas remains unclear. It is a useful vocabulary for the split visible all over the feed between overnight benchmark climbers and research-agent swarms.
πŸ’‘#27
@polydao
https://x.com/polydao/status/2105977846095040587
polydao audited one real multi-file refactor in Claude Code and found 31 hidden checks inside it: did it pass, is this the right file, merge, retry or flag a person. Their estimate is that 25 to 40% of model calls in any agent loop with a review step are such checks, which can go to a small decision model at about 348 ms each while the frontier model writes the code and the review. The playbook is reasonable: mark every yes-or-no moment in one real transcript, write the schema down, move one decision first, and hand anything below the confidence threshold back to the big model. The honest caveat is included: a small model cannot hallucinate text because it never writes any, but a bad schema lets it pick the wrong option with clean confidence.
πŸ’‘#28
@generussai
https://x.com/generussai/status/2106101738562552252
generussai, responding to a long agent-builder checklist, said the cost kill-switch is the one they would build first, having watched an agent loop torch credits while away from the keyboard. A hard cap per query beats finding out later. It is the most repeated lesson from people who have actually left loops running: the budget guard is not an optimization, it is the first feature.
πŸ“‘ Eco Products Radar
Eco Products Radar

autoresearch / auto-research (37 mentions): the pattern itself, from overnight renderer tuning to a five-million-line Lean formalization, now also running inside Claude Code via a pi-autoresearch port.
Claude Code (15): the most common host for loops this week, and with Mods (8) the loop can now live inside the agent.
Codex (11): the other half of most multi-agent setups, often paired with Claude Code for cross-checking.
Jev (10): small decision models used for the yes-or-no checks inside loops, with public measured tests.
MCP (9): the tool bridge most loop builders name when wiring agents to external systems.
Opus 5.5 (8) and Karpathy's loop (7): the model and the recipe most often cited behind the runs.
← Previous
Super User Daily: 2026-10-03
Next β†’
Ideas Radar: 2026-10-03
← Back to all articles

Comments

Loading...
>_