MobileDreamer: Generative Sketch World Model for GUI Agent
TLDR
MobileDreamer proposes a generative sketch world model for mobile GUI agents to forecast post-action states and improve long-horizon task performance.
Reasoning
Strengths include a novel textual sketch world model with order-invariant learning for spatial preservation and a rollout imagination strategy for action selection. Weaknesses are the domain specificity to GUI agents and lack of detailed efficiency analysis.
Read-first score
Read-first score 59.4, weighted from topical fit, citation, graph, method, reproducibility, and recency signals. Original total remains 48.
Field roles
Rank sensitivity
Stability: volatile; rank range: 410.
Keyword Scores
Deep Analysis
Innovations
- Textual sketch world model that transforms digital images into key task-related sketches
- Order-invariant learning strategy to preserve spatial information of GUI elements
- Rollout imagination strategy for GUI agent to optimize action selection using world model predictions
Methodology
MobileDreamer proposes a world-model-based lookahead framework consisting of a textual sketch world model and a rollout imagination strategy. The world model learns to convert digital images into task-related sketches while using an order-invariant learning approach to maintain spatial awareness. The rollout imagination leverages the world model's predictions to guide the agent's action selection process.
Key Results
On Android World, MobileDreamer achieves state-of-the-art performance with a 5.25% improvement in task success rate. World model evaluations confirm that the textual sketch modeling accurately forecasts key GUI elements.