October 3, 2026ResearchAgentsOpen Source

AutoGUIWorld: Train Computer-Use Agents on Screens That Never Existed

Training a computer-use agent usually means actually running the software: installing apps, configuring them, putting them in the right state, then recording clicks. Every niche application adds setup cost, so the data ends up concentrated on the handful of apps that are easy to host. AutoGUIWorld (HF 40 upvotes) skips the software entirely and lets an image generator play the operating system.

The setup works like this. A structured spec describes the OS context, visual appearance and interface state, and the system samples a starting screenshot from it. A task gets generated for that scene. A planner then picks each atomic action and describes what should visibly happen as a result, and an image generator edits the current screenshot to produce the next one. Action grounding and transition-level quality filters clean it up. The output is 79,266 step-level training samples with spatial annotations across Ubuntu, Windows, macOS and Chrome, none of them recorded from a real running app.

It works better than it sounds like it should. Fine-tuning Qwen3.5-35B-A3B on AutoGUIWorld trajectories lifts the OSWorld mean task score from 33.0% to 40.8%. On ScienceBoard, which tests specialized scientific software that is especially painful to deploy, task success goes from 14.0% to 32.2%, more than double.

The ScienceBoard jump is the real story. The long tail of professional software was the hardest place to collect agent data, and that is exactly where hallucinated screens help most. Image models have become good enough to serve as world models for GUIs. Code at github.com/ImYangC7/AutoGUIWorld.

Link: arxiv.org/abs/2610.01215
← Previous
GraphForge: 2,169 Trajectories Push a 27B Open Model Up 65 Points on GDPVal
Next β†’
Argo-Bench: 7.5 Billion Rows, and the Agent Gets Graded on What Happens Next
← Back to all articles

Comments

Loading...
>_