CUA-SWE: When Computer-Use Agents Meet Visual Software Engineering

Prince Zizhuang Wang*†, Chenhao Liang*, Zelong Xu*, Aojie Yuan*, Xiaolin Zhou*, Haiyue Zhang, Yue Zhao, Xiyang Hu, Shuli Jiang

Carnegie Mellon University · University of Southern California · University of Wisconsin–Madison · Arizona State University · AWS Agentic AI  ·  *Equal contribution  ·  †Project lead

CUA-SWE provides a benchmark and environment where coding agents need to repair the software by using it and visually inspecting it. Each of the 105 tasks gives the agent an editable project and the running application; deterministic, task-specific tests then judge the repaired software on a clean copy of the project. Across Web, Game, DevOps and Mobile, adding computer use to coding raises task success for all eight API models in the paired comparison.

CUA-SWE overview: task domains and buggy software, the edit, check, screenshot, interact and revise loop, and the independent deterministic verifier
Agents inspect screenshots, interact with the application and edit code; an independent deterministic verifier checks the requested behavior and preserved functionality on a clean copy. Figure 1 (left) in the paper.

Web Allocation Ring Studio

Allocation Ring Studio baseline: the live ring chart overflows its cardBug

Ring chart overflows its card in the Stacked layout

After the reference repair the ring fits its cardRepaired

Ring fits every layout profile

Game Vector Relay

Vector Relay before the repair: relay armed, orb replaced, score 0Bug

Return fractured, relay stays armed, score 0

Vector Relay after the repair: circuit complete, score 500Repaired

Circuit complete, +500

DevOps Counter Order Console

Counter Order Console baseline: a rolling restart renders as a 12,516 requests-per-second spikeBug

Rolling restart rendered as a 12,516 req/s spike

After the reference repair the false spike is gone and the real burst remainsRepaired

False spike removed, real burst retained

Mobile Mural Desk

Mural Desk baseline: the artwork is projected onto the wall plan in the wrong orientationBug

Artwork projected onto the wall plan the wrong way

After the reference repair the artwork is aligned to the wall planRepaired

Artwork aligned to the wall plan

One task per domain, before and after the repair. Web, DevOps and Mobile show reference-repair replays; Game shows two screenshots from a recorded GPT-5.6 Sol attempt.

Results

Task success (%) on one selected attempt per task, model and condition. Code-only uses the coding interface: read and edit source, run commands. Hybrid adds computer use: screenshots of the running application and mouse and keyboard actions. Both conditions share the same tasks, budgets and hidden tests, so the gap isolates what visual feedback and interaction contribute.

a Overall: four-domain mean for the eight frontier models evaluated under both conditions in every domain
c Success and development time: mean Hybrid task success against mean agent time, nine models and systems
Mean task success against mean agent time in minutes for nine models and systems; GPT-6 Astra is highest at about 4 minutes
b By domain: Code-only and Hybrid for every evaluated model and agent system

lighter bar: Code-only    solid bar: Hybrid (computer use + coding)    +pp gain from computer use    Hue identifies the model: GPT · Claude · Grok · agent systems

Panel a weights the four domains equally and matches the right panel of Figure 1 in the paper; Opus 5 is omitted because its Mobile evaluation is deferred. Panel c is Figure 6a of the paper (circles: frontier models; open square: Codex). Codex and Claude Code are evaluated as complete agent systems in the Hybrid condition only. Every frontier model gains with Hybrid access; the gains concentrate on tasks whose requirements include application materials (see Analysis).

Single-attempt task success (%) across four domains

Column maxima are in bold. Hover a cell for the number of solved tasks. Each domain retains its original scoring rules: Web success also requires the required visual evidence; DevOps and Mobile exclude infrastructure errors and count genuine agent failures. Table 1 in the paper; Table 4 there reports a three-attempt Hybrid evaluation of GPT-6 Astra, GPT-5.6 Sol and Fable 5.

Examples

Recorded runs, step by step: what the agent saw, what it changed, and what the verifier decided. This one follows GPT-5.6 Sol through a Vector Relay attempt; players for Web, DevOps and Mobile follow the same layout.

    Screenshots, patch and verifier result come from this attempt. In the task’s three matched trials, Code-only solved 0/3 and Hybrid solved 1/3.

    Tasks

    Every task starts from a verified bug in a working application and specifies the failure to fix, the behavior to implement and the existing functionality to preserve. Related variants share application code and test different interactions or configurations.

    Web 36 tasks · 20 applications

    Interface geometry, selection and application contracts in Fabric.js (9 tasks), MapLibre (7), React Spectrum, Saleor, JupyterLab, Quill, Tiptap, Handsontable, ECharts, Recharts and business workflows such as account recovery, tax rounding and warehouse cutoffs.

    Make the account recovery page display the channel selected by the live identity risk policy for the account being viewed, including ACC-731, while preserving the recovery decision trace.

    Game 29 tasks · 12 games

    Temporal behavior, physics and state transitions in Captain Callisto (7), Boxel Rebound (5), Flappy Bird (4), 2048, Hextris, Minesweeper, Vector Relay, Astray, Core Ball, Letter Arc, Maze Signal and Phase Relay. Tests set initial states, apply inputs and restart with fixed seeds.

    The deterministic Level 3 benchmark build contains a gameplay regression. Play the pulse lane and compare the start, release, and restart of several short jetpack bursts.

    DevOps 20 tasks · 10 contracts

    Operational observations tied to configuration and data-processing semantics: rate before aggregation, alert inhibition, log severity normalization, pending-alert identity, SLO rollups, StatsD gauge modes, tail sampling, trace clock skew and trace context across awaits, on Prometheus, Alertmanager, Jaeger and OpenTelemetry code.

    After the collector rollout, failed operations are missing from the log workbench’s severe view while a routine event appears as an error. Investigate the live records and receiver fleet, then repair severity classification.

    Mobile 20 tasks · 12 families

    Visually specified relationships and editing behavior in mobile-web applications run in Chromium at 400×800: nine Gantt variants plus Mural Desk, Foldnote, Pass wallet, Lamp courier, Northline, Tideglass, Marble Post, Lantern, Campus Pocket and inspection and plate-integration workflows.

    Foldnote is a mobile web packaging-proof review mockup. Review markers no longer stay correctly placed and oriented when switching between Flat sheet and Individual panel.

    Specification basis

    Tasks differ in where their requirements are specified. In source-specified tasks (42), the instruction and readable project code define the target behavior; running the application supports diagnosis and checking. In application-material-dependent tasks (63), part of the behavior is specified only through application materials such as an imported drawing, a displayed service contract or a graphical reference card, which the code-only interface does not expose. Both groups require implementing the behavior and verifying the resulting software.

    Two example tasks

    Game example: repairing inventory updates in Captain Callisto
    Game. Captain Callisto: one pickup changes the inventory from 0/2 to 2/2. The agent reproduces the fault with game controls and screenshots and edits the inventory logic; independent checks test collection and level completion under two replay scenarios. Figure 3 in the paper.
    Mobile example: repairing imported routes in Campus Pocket
    Mobile. Campus Pocket: the provider report identifies a closed lift and the travel costs used to build a step-free route; the faulty route uses the closed H6 lift, the repaired one avoids it. Figure 4 in the paper.

    Browse all 105 tasks with their original instructions, budgets and tests

    Analysis

    Computer use improves task success on every model

    Task-paired contribution of computer use for nine API models across four domains
    Task-paired outcomes under the two access conditions. Each bar partitions a model's matched tasks by the joint Code-only and Hybrid outcome; the value at right is the Hybrid gain in percentage points: tasks solved only with computer use minus tasks solved only with code. Figure 20 in the paper.

    Every frontier model in the four-domain comparison achieves higher aggregate task success with Hybrid access; GPT-6 Astra gains 48.6 percentage points, from 11.3% to 59.9%. The gains concentrate on tasks whose requirements include application materials, while source-specified subsets have near-zero domain-average changes and substantial variation across models.

    Success by task information requirement (%), averaged over the eight frontier models. Table 9 in the paper.
    DomainTypeTasksCode-onlyHybridΔ pp
    WebS2223.923.9+0.0
    WebM148.950.9+42.0
    GameS2010.610.0−0.6
    GameM92.826.4+23.6
    DevOpsM200.648.8+48.1
    MobileM200.639.4+38.8

    In material-dependent (M) tasks the agent must recover requirements from drawings or displayed behavior, implement them and preserve the surrounding workflow. Source-level execution tests the implementation; application observations supply requirements and feedback that those executions alone may not expose. The same figure also shows the limits: a few Web tasks are solved only with code, and most Game tasks remain unsolved under both conditions.

    How agents turn runtime observations into correct changes

    Four released GPT-6 Astra attempts, one per domain, as reviewed in the paper (Figures 28 to 31). Each card shows what the agent observed, the consequential source change, and the evaluator's verdict. Click a card to enlarge.

    Web case: preserve nested selection after resizing
    Web · following the consequences of a code change. After repairing resizing, the agent drags the object beyond its containing frame; when selection is lost it extends hit testing to selectable descendants and repeats the interaction. GPT-5.6 Sol and Claude Opus 5 repair the resize but leave the selection defect.
    Game case: restore checkpoint state from geometry
    Game · state updates jointly restore the scene. The checkpoint stores geometry; the agent reconstructs lift phase from a faded reference frame, restores camera progress in course coordinates, and reattaches the player at its saved height.
    DevOps case: infer each producer's gauge-update semantics
    DevOps · live behavior identifies the operative contract. A pulse–pulse–settle sequence in the gauge console reveals each endpoint's update mode; the repair relearns the mode at each epoch and preserves receipt order.
    Mobile case: recover gear connections from the received drawings
    Mobile · recovering relations from drawings. Enlarged maker drawings expose shared axes and belt routing; the agent reconstructs all 23 driving relations and implements their angle, phase and editing behavior. GPT-5.6 Sol chooses the wrong driving rim at two compound assemblies.

    Where repairs still fail

    Core Ball's final retained agent screenshot still shows the red collision screen
    A local check passes; the repair fails. In Core Ball, the agent's scripted replay completes four attachments and the build passes, but its final browser screenshot still shows a collision. The hidden tests reject the patch: impact is resolved one simulation step late, and a rejected pin still occupies a position. Task ↗
    Letter Arc screenshot 17: the revealed ELEGY rowLetter Arc screenshot 18: revealed tile faces turn dark mid-animation
    Replay reveals the next defect. In Letter Arc, an initial CSS edit exposes a second bug: revealed tiles turn dark mid-animation. Only further edits that separate the animation layers pass all 24 browser checkpoints. Task ↗

    Evaluation

    The CUA-SWE development and evaluation pipeline: task package, matched comparison of a code-only agent and a hybrid computer-use agent, and the shared evaluator-private deterministic verifier
    The development and evaluation pipeline. A task package gives the agent a bug report and an editable repository. The Code-only agent reads code, edits files and runs commands; the Hybrid agent additionally views screenshots, clicks, types and scrolls in the running application. The same evaluator-private verifier applies the permitted edits to a clean baseline, rebuilds the application and runs the task-specific tests. Figure 2 in the paper.

    What counts as a success

    Task check

    Does the requested behavior work? Tests execute the rebuilt application under controlled interaction sequences and assert on the rendered interface, runtime behavior or internal application state.

    Regression check

    Does designated existing behavior still work? The same tests cover the functionality the task says must be preserved, so a fix that breaks normal play, editing or navigation fails.

    Permitted-change check

    Did the patch stay inside the allowed scope? Edits outside the files the task permits fail regardless of test results.

    These three checks give patch correctness. Reported task success additionally requires a valid trajectory: the run used only the interfaces allowed by its condition and, for Hybrid runs, actually observed the screenshots. Correctness depends on the behavior of the submitted program, not on matching the reference patch; the evaluator may inspect privileged state such as a game board that the agent only sees rendered.

    How tasks are built and validated

    Each task starts from a software behavior that can be reproduced in the running application and modified in the codebase. An LLM author constructs the instruction, initial codebase, runtime setup, gold reference repair, plausible negative repairs and protected tests; human review checks that the requirement is meaningful and that the relevant runtime evidence is observable through the benchmark interface. The verifier must reject the broken project and every negative repair and accept the gold repair. In the masked-date task, for example, clearing the field should remove the selected date while invalid input should keep it; a repair that clears the selection in both cases fails. Reviewed construction-time agent trials then guide neighboring variants that extend each task family's difficulty range.

    Web example: preserving distinct masked-date input behaviors, with the reference replay, the required-behavior table and the replay audit
    Validation of one task. The masked-date input must stay empty after clearing and blurring, report no selected date, and keep the previous date on invalid text. The reference replay shows the gold repair meeting all three requirements; the required-behavior table is what the verifier checks, and the replay audit confirms the behavior is observable through the benchmark interface. A negative repair that clears the selection in both cases fails. Figure 5 in the paper.

    Citation

    CUA-SWE: When Computer-Use Agents Meet Visual Software Engineering. arXiv:2609.32600, 2026. Abstract ↗ · PDF ↗ · Code and task bundles ↗ · Dataset on Hugging Face ↗ · Recorded trajectories ↗

    @misc{wang2026cuaswe,
      title         = {CUA-SWE: When Computer-Use Agents Meet Visual Software Engineering},
      author        = {Wang, Prince Zizhuang and Liang, Chenhao and Xu, Zelong and Yuan, Aojie and Zhou, Xiaolin and Zhang, Haiyue and Zhao, Yue and Hu, Xiyang and Jiang, Shuli},
      year          = {2026},
      eprint        = {2609.32600},
      archivePrefix = {arXiv},
      url           = {https://arxiv.org/abs/2609.32600}
    }