Awesome Reinforcement Learning for GUI Agents

July 26, 2026 ยท View on GitHub

Awesome License: MIT Contributions Welcome

This repository provides a comprehensive and curated list of research papers, datasets, and tools focused on Reinforcement Learning (RL) in GUI Agents. GUI agents are intelligent systems that perceive graphical interfaces visually and execute tasks through human-like inputs (click, swipe, type).

๐Ÿ“„ Based on the survey: GUI Agents with Reinforcement Learning: Toward Digital Inhabitants


๐Ÿ“‹ Table of Contents

Survey Structure
Overview of the survey structure. We organize our analysis into three main pillars: RL Methods, Key Dimensions, and Training Resources.


๐Ÿ”” News

  • [2026-04-30] ๐Ÿ“„ Our survey "GUI Agents with Reinforcement Learning: Toward Digital Inhabitants" is now available on arXiv!
  • [2026-04-19] ๐Ÿš€ Repository created! Stay tuned for more updates on RL-based GUI Agents.

๐ŸŒŸ Introduction

Reinforcement Learning for GUI agents addresses the core difficulties of GUI automation: long-horizon credit assignment under sparse rewards, distribution shift across evolving interfaces, and safe exploration. We organize the landscape into three methodological paradigms:

  • Offline RL: Learning from static datasets without environment interaction.
  • Online RL: Refinement through continuous trial and error in dynamic environments.
  • Hybrid Strategies: Bridging pre-training and adaptation via semi-online methods and world models.

RL Training Pipeline
Overview of the RL training pipeline for GUI agents. The agent perceives the GUI environment through screenshots, reasons about the task, and executes actions. RL optimizes the policy through reward signals derived from task completion, visual grounding accuracy, and intermediate reasoning quality.

GUI Agent Timeline
Timeline of GUI Agent Development from rule-based systems to the multimodal LLM era.


PaperVenue / Year
Gui agents: A surveyFindings of ACL 2025
Gui agents with foundation models: A comprehensive surveyarXiv 2024
Large language model-brained gui agents: A surveyarXiv 2024
Llm-powered gui agents in phone automation: Surveying progress and prospectsarXiv 2025
A survey on (m) llm-based gui agentsarXiv 2025
A Comprehensive Survey of Agents for Computer Use: Foundations, Challenges, and Future DirectionsarXiv 2025
A survey of webagents: Towards next-generation ai agents for web automation with large foundation modelsKDD
Api agents vs. gui agents: Divergence and convergencearXiv 2025
Survey on large language model-enhanced reinforcement learning: Concept, taxonomy, and methodsTNNLS
Reinforcement learning enhanced llms: A surveyarXiv 2024
The landscape of agentic reinforcement learning for llms: A surveyarXiv 2025
A survey of reinforcement learning for large reasoning modelsarXiv 2025
Llm-based multi-agent reinforcement learning: Current and future directionsarXiv 2024
Os agents: A survey on mllm-based agents for computer, phone and browser useACL
From system 1 to system 2: A survey of reasoning large language modelsarXiv 2025
Lifelong learning of large language model based agents: A roadmapTPAMI

๐Ÿค Contributing

Contributions are welcome! Please see CONTRIBUTING.md for guidelines.


๐Ÿ—๏ธ RL Methods

RL Training Pipeline
Overview of the RL training pipeline for GUI agents. The agent perceives the GUI environment through screenshots, reasons about the task, and executes actions. RL optimizes the policy through reward signals derived from task completion, visual grounding accuracy, and intermediate reasoning quality.

Frontier Models

PaperVenue / Year
Agent S: an Open Agentic Framework That Uses Computers Like a HumanarXiv 2024
Agent S2: a Compositional Generalist-specialist Framework for Computer Use AgentsarXiv 2025
Constitutional Ai: Harmlessness from Ai FeedbackarXiv 2022
Digirl: Training In-the-wild Device-control Agents with Autonomous Reinforcement LearningNeurIPS 2024
Qwen3-vl Technical Report2025
Gui-eyes: Tool-augmented Perception for Visual Grounding in Gui AgentsarXiv 2026
Mano Technical ReportarXiv 2025
Gui Exploration Lab: Enhancing Screen Navigation in Agents Via Multi-turn Reinforcement Learning2025
Navigating the Digital World as Humans Do: Universal Visual Grounding for Gui AgentsarXiv 2024
Seed1. 5-vl Technical ReportarXiv 2025
Cogagent: a Visual Language Model for Gui AgentsCVPR 2024
Clickagent: Enhancing Ui Location Capabilities of Autonomous AgentsSIGDIAL 2025
Spiritsight Agent: Advanced Gui Agent with One LookCVPR 2025
Efficient Multi-turn Rl for Gui Agents Via Decoupled Training and Adaptive Data CurationarXiv 2025
Showui: One Vision-language-action Model for Gui Visual AgentCVPR 2025
Infigui-g1: Advancing Gui Grounding with Adaptive Exploration Policy OptimizationAAAI 2026
Infiguiagent: a Multimodal Generalist Gui Agent with Native Reasoning and ReflectionarXiv 2025
Ui-s1: Advancing Gui Automation Via Semi-online Reinforcement LearningarXiv 2025
Ui-r1: Enhancing Efficient Action Prediction of Gui Agents by Reinforcement LearningarXiv 2025
Gui-r1: a Generalist R1-style Vision-language Action Model for Gui AgentsarXiv 2025
Visual Test-time Scaling for Gui Agent GroundingICCV 2025
Computer-using Agent2025
Ui-tars: Pioneering Automated Gui Interaction with Native AgentsarXiv 2025
Falcon-ui: Understanding Gui Before Following User InstructionsarXiv 2024
Coact-1: Computer-using Agents with Coding as ActionsarXiv 2025
Gui-g$^22025
Magicgui: a Foundational Mobile Gui Agent with Scalable Data Pipeline and Reinforcement Fine-tuningarXiv 2025
Kimi-vl Technical ReportarXiv 2025
Internvl3. 5: Advancing Open-source Multimodal Models in Versatility, Reasoning, and EfficiencyarXiv 2025
Opencua: Open Foundations for Computer-use AgentsarXiv 2025
Ponder & Press: Advancing Visual Gui Agent Towards General Computer ControlFindings of ACL 2025
Ui-tars-2 Technical Report: Advancing Gui Agent with Multi-turn Reinforcement LearningarXiv 2025
Os-copilot: Towards Generalist Computer Agents with Self-improvementarXiv 2024
Backtrackagent: Enhancing Gui Agent with Error Detection and Backtracking MechanismarXiv 2025
Vsc-rl: Advancing Autonomous Vision-language Agents with Variational Subgoal-conditioned Reinforcement LearningarXiv 2025
Aguvis: Unified Pure Vision Agents for Autonomous Gui InteractionarXiv 2024
Step-gui Technical ReportarXiv 2025
Aria-ui: Visual Grounding for Gui InstructionsFindings of ACL 2025
Gta1: Gui Test-time Scaling AgentarXiv 2025
Mobile-agent-v3: Fundamental Agents for Gui AutomationarXiv 2025
Se-gui: Enhancing Visual Grounding for Gui Agents Via Self-evolutionary Reinforcement LearningN/A
Uitron: Foundational Gui Agent with Advanced Perception and PlanningarXiv 2025
Agentcpm-gui: Building Mobile-use Agents with Reinforcement Fine-tuningEMNLP 2025
Phi-ground Tech Report: Advancing Perception in Gui GroundingarXiv 2025
Ufo2: the Desktop AgentosarXiv 2025
Omegause: Building a General-purpose Gui Agent for Autonomous Task ExecutionarXiv 2026
Mai-ui Technical Report: Real-world Centric Foundation Gui AgentsarXiv 2025
Introducing computer use, a new Claude 3.5 Sonnet, and Claude 3.5 HaikuN/A
Introducing the Gemini 2.5 Computer Use modelN/A
Computer-Using AgentN/A
Os-atlas: A foundation action model for generalist gui agentsICLR
Agent q: Advanced reasoning and learning for autonomous ai agentsarXiv 2024

Reinforcement Learning Paradigms

Offline RFT Methods

PaperVenue / Year
Direct Preference Optimization: Your Language Model Is Secretly a Reward ModelNeurIPS 2023
Deepseekmath: Pushing the Limits of Mathematical Reasoning in Open Language ModelsarXiv 2024
Conservative q-learning for offline reinforcement learningNeurIPS
Offline reinforcement learning with implicit q-learningICLR
Decision transformer: Reinforcement learning via sequence modelingNeurIPS
Behavioral cloning from observationIJCAI
Implicit behavioral cloningConference on robot learning
Advantage-weighted regression: Simple and scalable off-policy reinforcement learningarXiv 2019

Representative Methods

PaperVenue / Year
Digirl: Training In-the-wild Device-control Agents with Autonomous Reinforcement LearningNeurIPS 2024
Dynaweb: Model-based Reinforcement Learning of Web AgentsarXiv 2026
Ui-agile: Advancing Gui Agents with Effective Reinforcement Learning and Precise Inference-time GroundingarXiv 2025
Ui-s1: Advancing Gui Automation Via Semi-online Reinforcement LearningarXiv 2025
Hiper: Hierarchical Reinforcement Learning with Explicit Credit Assignment for Large Language Model AgentsarXiv 2026
Probabilistic Subgoal Representations for Hierarchical Reinforcement LearningarXiv 2024
Hi-agent: Hierarchical Vision-language Agents for Mobile Device ControlarXiv 2025
Ultracua: a Foundation Model for Computer Use Agents with Hybrid ActionarXiv 2025
ARPO: End-to-End Policy Optimization for GUI Agents with Experience ReplayarXiv 2025
Mobilerl: Online agentic reinforcement learning for mobile gui agentsarXiv 2025
Digi-q: Learning q-value functions for training device-control agentsarXiv 2025

Emerging Directions

PaperVenue / Year
Group-in-group Policy Optimization for Llm Agent TrainingarXiv 2025
Enhancing Cooperative Multi-agent Reinforcement Learning with State Modelling and Adversarial ExplorationarXiv 2025
Wcsac: Worst-case Soft Actor Critic for Safety-constrained Reinforcement LearningAAAI 2021
Constrained Reinforcement Learning with Smoothed Log Barrier FunctionarXiv 2024
CGL: Advancing Continual GUI Learning via Reinforcement Fine-TuningarXiv 2026
Continual GUI AgentsarXiv 2026
Autonomous Continual Learning of Computer-Use Agents for Environment AdaptationarXiv 2026
Online continual learning for interactive instruction following agentsarXiv 2024

Algorithmic Advances: Exploration and Multi-Turn Optimization

PaperVenue / Year
Agentic Entropy-balanced Policy OptimizationarXiv 2025
Nested Browser-use Learning for Agentic Information SeekingarXiv 2025
Infigui-g1: Advancing Gui Grounding with Adaptive Exploration Policy OptimizationAAAI 2026
Webrl: Training Llm Web Agents Via Self-evolving Online Curriculum Reinforcement LearningICLR 2024
Gui-g$^22025

๐ŸŽจ Key Dimensions

Reward Engineering

The process of defining objective feedback signals for GUI tasks.

Reward Engineering Pyramid
The Reward Engineering Pyramid balances accuracy and generality for GUI Agents: rule-based rewards offer precision, while learned rewards and LLM-as-judge enable semantic depth.

Reward Engineering

PaperVenue / Year
Gui Agents: a SurveyFindings of ACL 2025
Rlthf: Targeted human feedback for llm alignmentarXiv 2025
Rlaif vs. rlhf: Scaling reinforcement learning from human feedback with ai feedbackNeurIPS
Curriculum-rlaif: Curriculum alignment with reinforcement learning from ai feedbackarXiv 2025
Agentprm: Process reward models for llm agents via step-wise promise and progressarXiv 2025
Process reinforcement through implicit rewardsarXiv 2025
Ovm, outcome-supervised value models for planning in mathematical reasoningFindings of ACL 2024

Rule-Based Rewards

PaperVenue / Year
Mind2web: Towards a Generalist Agent for the WebNeurIPS 2023
Infigui-g1: Advancing Gui Grounding with Adaptive Exploration Policy OptimizationAAAI 2026
Ui-r1: Enhancing Efficient Action Prediction of Gui Agents by Reinforcement LearningarXiv 2025
Gui-g$^22025
Lpo: Towards Accurate Gui Agent Interaction Via Location Preference OptimizationarXiv 2025
Btl-ui: Blink-think-link Reasoning Model for Gui AgentarXiv 2025

LLM-as-Judge Rewards

PaperVenue / Year
Smartsnap: Proactive Evidence Seeking for Self-verifying AgentsarXiv 2025
Prore: a Proactive Reward System for Gui Agents Via Reasoner--actor CollaborationarXiv 2025
Webrl: Training Llm Web Agents Via Self-evolving Online Curriculum Reinforcement LearningICLR 2024
Zerogui: Automating Online Gui Learning at Zero Human CostarXiv 2025

Learned Rewards

PaperVenue / Year
Infigui-g1: Advancing Gui Grounding with Adaptive Exploration Policy OptimizationAAAI 2026
Gui-g$^22025

Data Efficiency

Synthetic Data via World Models

PaperVenue / Year
Scaling Agent Learning Via Experience SynthesisarXiv 2025
Simura: a World-model-driven Simulative Reasoning Architecture for General Goal-oriented AgentsarXiv 2025
Websynthesis: World-model-guided Mcts for Efficient Webui-trajectory SynthesisarXiv 2025
Llms as Scalable, General-purpose Simulators for Evolving Digital Agent TrainingarXiv 2025
Webworld: a Large-scale World Model for Web Agent TrainingarXiv 2026
Code2world: a Gui World Model Via Renderable Code GenerationarXiv 2026
Is your llm secretly a world model of the internet? model-based planning for web agentsarXiv 2024

Enhancement of Human Demonstrations

PaperVenue / Year
Gui-rewalk: Massive Data Generation for Gui Agent Via Stochastic Exploration and Intent-aware ReasoningarXiv 2025
Watch and Learn? Using Edpuzzle to Enhance the Use of Online VideosManagement Teaching Review 2019
Watch and Learn: Learning to Use Computers from Online VideosarXiv 2025
Os-genesis: Automating Gui Agent Trajectory Construction Via Reverse Task SynthesisACL 2025
Agenttrek: Agent Trajectory Synthesis Via Guiding Replay with Web TutorialsarXiv 2024
Prune4web: Dom Tree Pruning Programming for Web AgentarXiv 2025

Iterative Self-Improvement

PaperVenue / Year
Gui-r1: a Generalist R1-style Vision-language Action Model for Gui AgentsarXiv 2025
Co-epg: a Framework for Co-evolution of Planning and Grounding in Autonomous Gui AgentsarXiv 2025

Technical Innovations

Multimodal Perception: Active and Adaptive Visual Grounding

PaperVenue / Year
Gui-eyes: Tool-augmented Perception for Visual Grounding in Gui AgentsarXiv 2026
GroundCUA: Grounding Computer Use Agents on Human DemonstrationsICLR 2026
Infigui-g1: Advancing Gui Grounding with Adaptive Exploration Policy OptimizationAAAI 2026
Gui-g$^22025
Gui-actor: Coordinate-free Visual Grounding for Gui AgentsarXiv 2025
Gui-aima: Aligning Intrinsic Multimodal Attention with a Context Anchor for Gui GroundingarXiv 2025
Seeclick: Harnessing gui grounding for advanced visual gui agentsACL
Mapping natural language commands to web elementsEMNLP 2018
Understanding html with large language modelsFindings of EMNLP 2023
Attacking vision-language computer agents via pop-upsACL

Memory and Planning: Sustaining Context over Long Horizons

PaperVenue / Year
Mga: Memory-driven Gui Agent for Observation-centric InteractionarXiv 2025
Memr $^2025
Plan-and-act: Improving Planning of Agents for Long-horizon TasksarXiv 2025
Magnet: Towards Adaptive Gui Agents with Memory-driven Knowledge EvolutionarXiv 2026
Agentprog: Empowering Long-horizon Gui Agents with Program-guided Context ManagementarXiv 2025
History-aware Reasoning for Gui AgentsarXiv 2025
Webagent-r1: Training Web Agents Via End-to-end Multi-turn Reinforcement LearningEMNLP 2025
Auto-scaling Continuous Memory for Gui AgentarXiv 2025
Memsearcher: Training Llms to Reason, Search and Manage Memory Via End-to-end Reinforcement LearningarXiv 2025
MemR ห†3\^{} 3: Memory Retrieval via Reflective Reasoning for LLM AgentsarXiv 2025

๐Ÿ“Š Training Resources

Datasets

Training Pipeline Pyramid
This pyramid depicts a four-stage data-training pipeline for agent capability, progressing from static data imitation to offline RL, synthetic simulation, and online RL.

Demonstration and Trajectory Datasets

PaperVenue / Year
Mind2web: Towards a Generalist Agent for the WebNeurIPS 2023
Omniact: a Dataset and Benchmark for Enabling Multimodal Generalist Autonomous Agents for Desktop and WebECCV 2024
On the Effects of Data Scale on Ui Control AgentsNeurIPS 2024
Androidinthewild: a Large-scale Dataset for Android Device ControlNeurIPS 2023
Guiodyssey: A comprehensive dataset for cross-app gui navigation on mobile devicesICCV
Learnact: Few-shot mobile gui agent with a unified demonstration benchmarkarXiv 2025

Perception and Grounding Datasets

PaperVenue / Year
Unveiling the Tricks: Automated Detection of Dark Patterns in Mobile ApplicationsUIST 2023
Rico: a Mobile App Dataset for Building Data-driven Design ApplicationsUIST 2017
Widget Captioning: Generating Natural Language Description for Mobile User Interface ElementsEMNLP 2020
Ferret-ui 2: Mastering Universal User Interface Understanding Across PlatformsICLR 2024
Screenspot-pro: Gui Grounding for Professional High-resolution Computer UseACM MM 2025
GroundCUA: Grounding Computer Use Agents on Human DemonstrationsICLR 2026
Screen2words: Automatic Mobile Ui Summarization with Multimodal LearningUIST 2021
Ferret-ui: Grounded Mobile Ui Understanding with Multimodal LlmsECCV 2024

Synthetic and RL-Generated Corpora

PaperVenue / Year
Agent-x: Evaluating Deep Multimodal Reasoning in Vision-centric Agentic TasksarXiv 2025
Gui-bee: Align Gui Action Grounding to Novel Environments Via Autonomous ExplorationarXiv 2025
End-to-end Navigation with Vision Language Models: Transforming Spatial Reasoning Into Question-answeringarXiv 2024
Visualwebarena: Evaluating Multimodal Agents on Realistic Visual Web TasksACL 2024
Explorer: Scaling Exploration-driven Web Trajectory Synthesis for Multimodal Web AgentsFindings of ACL 2025
Webcanvas: Benchmarking Web Agents in Online EnvironmentsarXiv 2024
Ui-tars: Pioneering Automated Gui Interaction with Native AgentsarXiv 2025
Scaling Synthetic Task Generation for Agents Via ExplorationarXiv 2025
Androidworld: a Dynamic Benchmarking Environment for Autonomous AgentsICLR 2024
Bearcubs: a Benchmark for Computer-using Web AgentsarXiv 2025
Beyond Browsing: Api-based Web AgentsFindings of ACL 2025
Webwalker: Benchmarking Llms in Web TraversalACL 2025
Osworld: Benchmarking Multimodal Agents for Open-ended Tasks in Real Computer EnvironmentsNeurIPS 2024
Theagentcompany: Benchmarking Llm Agents on Consequential Real World TasksarXiv 2024
Appagent: Multimodal Agents as Smartphone UsersCHI 2025
Webarena: a Realistic Web Environment for Building Autonomous AgentsICLR 2023

Interactive Environments

Web and Browser Environments

PaperVenue / Year
The Browsergym Ecosystem for Web Agent ResearcharXiv 2024
Mind2web: Towards a Generalist Agent for the WebNeurIPS 2023
A Data-driven Approach for Learning to Control ComputersarXiv 2022
Visualwebarena: Evaluating Multimodal Agents on Realistic Visual Web TasksACL 2024
Reinforcement Learning on Web Interfaces Using Workflow-guided ExplorationICLR 2018
Webchorearena: Evaluating Web Browsing Agents on Realistic Tedious Web TasksarXiv 2025
Webshop: Towards Scalable Real-world Web Interaction with Grounded Language AgentsNeurIPS 2022
Webarena: a Realistic Web Environment for Building Autonomous AgentsICLR 2023
Webgpt: Browser-assisted question-answering with human feedbackarXiv 2021
Webvoyager: Building an end-to-end web agent with large multimodal modelsACL
World of bits: An open-domain platform for web-based agentsICML
ClawBench: Can AI Agents Complete Everyday Online Tasks?arXiv 2026

Desktop and OS Environments

PaperVenue / Year
Computerrl: Scaling End-to-end Online Reinforcement Learning for Computer Use AgentsarXiv 2025
Screenagent: a Vision Language Model-driven Computer Control AgentIJCAI 2024
Osworld: Benchmarking Multimodal Agents for Open-ended Tasks in Real Computer EnvironmentsNeurIPS 2024
Windows agent arena: Evaluating multi-modal os agents at scalearXiv 2024
Ui-vision: A desktop-centric gui benchmark for visual perception and interactionarXiv 2025

Mobile Environments

PaperVenue / Year
Androidworld: a Dynamic Benchmarking Environment for Autonomous AgentsICLR 2024
Mobilegui-rl: Advancing Mobile Gui Agent Through Reinforcement Learning in Online EnvironmentarXiv 2025
Androidenv: a Reinforcement Learning Platform for AndroidarXiv 2021
Uisim: an Interactive Image-based Ui Simulator for Dynamic Mobile EnvironmentsarXiv 2025
Benchmarking Mobile Device Control Agents across Diverse ConfigurationsCoLLAs 2025
MobileSafetyBench: Evaluating Safety of Autonomous Agents in Mobile Device ControlarXiv 2024
Mobile-env: a Universal Platform for Training and Evaluation of Mobile InteractionCoRR 2023
Mai-ui Technical Report: Real-world Centric Foundation Gui AgentsarXiv 2025
PaperVenue / Year
Osworld-mcp: Benchmarking Mcp Tool Invocation in Computer-use AgentsarXiv 2025
Mcpworld: a Unified Benchmarking Testbed for Api, Gui, and Hybrid Computer Use AgentsarXiv 2025
Os-harm: A benchmark for measuring safety of computer use agentsarXiv 2025

RL Infrastructure

Distributed RL Architecture
An asynchronous distributed architecture for GUI RL agent training, decoupling slow environment interaction from fast GPU learning.

VLM-RL Algorithm Libraries and Framework Evolution

PaperVenue / Year
Openrlhf: an Easy-to-use, Scalable and High-performance Rlhf FrameworkEMNLP 2024
Efficient Memory Management for Large Language Model Serving with PagedattentionSOSP 2023
Real: Efficient Rlhf Training of Large Language Models with Parameter ReallocationMLSys 2025
Hybridflow: a Flexible and Efficient Rlhf FrameworkEuroSys 2025
Megatron-lm: Training Multi-billion Parameter Language Models Using Model ParallelismarXiv 2019
Rewarddance: Reward Scaling in Visual GenerationarXiv 2025
Pytorch Fsdp: Experiences on Scaling Fully Sharded Data ParallelVLDB 2023

Distributed Rollout and Training Architectures

PaperVenue / Year
Areal: a Large-scale Asynchronous Reinforcement Learning System for Language ReasoningarXiv 2025
Distrl: an Asynchronous Distributed Reinforcement Learning Framework for On-device Control AgentsarXiv 2024
Agent. Xpu: Efficient Scheduling of Agentic Llm Workloads on Heterogeneous SocarXiv 2025
Hethub: a Distributed Training System with Heterogeneous Cluster for Large-scale ModelsarXiv 2024

Reward Engineering and Verification Systems

PaperVenue / Year
Agentic Reward Modeling: Verifying Gui Agent Via Online Proactive InteractionarXiv 2026
Mano Technical ReportarXiv 2025
Infigui-g1: Advancing Gui Grounding with Adaptive Exploration Policy OptimizationAAAI 2026
Progrm: Build Better Gui Agents with Progress RewardsarXiv 2025

Memory Management and Long-Horizon Reasoning

PaperVenue / Year
Mga: Memory-driven Gui Agent for Observation-centric InteractionarXiv 2025
Hi-agent: Hierarchical Vision-language Agents for Mobile Device ControlarXiv 2025
Memory-r1: Enhancing Large Language Model Agents to Manage and Utilize Memories Via Reinforcement LearningarXiv 2025

Integration and Ecosystem Standardization

PaperVenue / Year
The Browsergym Ecosystem for Web Agent ResearcharXiv 2024
Openhands: an Open Platform for Ai Software Developers as Generalist AgentsICLR 2024
Autogen: Enabling Next-gen Llm Applications Via Multi-agent ConversationsCOLM 2024

๐Ÿ“ Citation

If you find this repository or our survey useful, please consider citing:

@article{hu2026gui,
  title={GUI Agents with Reinforcement Learning: Toward Digital Inhabitants},
  author={Hu, Junan and Liu, Jian and Lai, Jingxiang and Hu, Jiarui and Sheng, Yiwei and Chen, Shuang and Li, Jian and Du, Dazhao and Guo, Song},
  journal={arXiv preprint arXiv:2604.27955},
  year={2026}
}