GUI agents · reinforcement learning research

Seeing What ChangedTransition-Guided Credit Assignment
for GUI Agents

A successful task does not mean every action deserves credit. OTAD uses what changed on screen to assign the verified outcome more precisely across turns.

A completed chart in LibreOffice Calc
An edited image with transparent background in GIMP
From verified successful desktop trajectories
52.1%
OSWorld-Verifiedtask success rate
+9.9 ppover budget-matched GRPO
8Bparameters · 50-step limit
01 / The problem

The outcome says
whether it worked. The screen
shows what changed.

Trajectory-level training broadcasts the same advantage across every turn. A mistaken click, its correction, and the decisive action can all be reinforced equally. Before-and-after screenshots offer finer evidence.

The verifier determines the sign of the advantage; observable state changes determine each turn’s weight. The scorer is used only during training.

02 / Method

Align every update
with visible evidence.

Observable-Transition Advantage Decomposition (OTAD) makes hindsight ordinal judgments about screenshot changes, then maps them to non-negative weights on the trajectory-level advantage rather than adding new rewards.

OTAD training loop: task sampling, group rollout and verification, observable-transition advantage decomposition, and budgeted turn selection
01

Observe the change

Compare screenshots before and after each action, conditioned on the goal, context, and verified outcome.

02

Allocate advantage

The verifier fixes the sign. Non-negative transition weights change only its magnitude, allowing detours in successful runs to be downweighted.

03

Spend the budget well

The same evidence selects informative training turns and prioritizes tasks near the capability frontier.

03 / Results

Sharper credit assignment.
Higher task success.

Across 361 evaluated OSWorld-Verified tasks, OTAD and budget-matched GRPO share the same SFT starting point, 50-step interaction cap, and training budget. Success rates are means over three independent passes.

OSWORLD-VERIFIED52.1%

overall success for OTAD-Qwen3-VL-8B

+9.9 ppover budget-matched GRPO
Matched-budget trainingSuccess rate ↑
+ SFT + GRPO42.2%
+ SFT + OTAD52.1%

Same 8B backbone; 50 interaction steps per trajectory and at most 16 selected training turns per trajectory.

Across task categories

OTAD · success rate
OS66.7%
Office52.7%
Daily64.1%
Professional75.5%
Workflow25.1%

Single-turn GUI grounding

Accuracy
ScreenSpot-Pro60.3%GRPO: 57.4%
OSWorld-G69.1%GRPO: 66.7%

Source: Tables 1, 2, and 3 of the paper. OSWorld-Verified covers 361 eligible tasks; results are means over three passes, with ±1.0 standard deviation for OTAD overall. Scorer compute is not matched in the GRPO baseline; see the paper.

04 / Verified trajectories

Watch the work unfold.

These four cases come from the provided evaluation records. Both the result file and trajectory metadata mark each run as successful. Watch the original recording or inspect screenshots and action logs turn by turn.

Verified success · 1.0
Original task instruction

Action log for this step

05 / Full paper

From a change on screen
to a better policy update.

Read the method, experimental setup, full comparisons, and theoretical appendix.

Open paper PDF
06 / Citation

Cite this work

This manuscript is under anonymous review. Please use the final author and publication details when the paper becomes public.

BibTeX
@misc{otad2026,
  title  = {Seeing What Changed: Transition-Guided Credit Assignment for GUI Agents},
  author = {{Anonymous Authors}},
  year   = {2026},
  note   = {Manuscript under review at ICLR 2027}
}

Provisional citation: authors, publication year, and public link will be updated after release.