AppAgent: Application Execution Agent

November 12, 2025 · View on GitHub

AppAgent is the core execution runtime in UFO, responsible for carrying out individual subtasks within a specific Windows application. Each AppAgent functions as an isolated, application-specialized worker process launched and orchestrated by the central HostAgent.


What is AppAgent?

![AppAgent Architecture](../../img/appagent2.png)
AppAgent Architecture: Application-specialized worker process for subtask execution

AppAgent operates as a child agent under the HostAgent's orchestration:

  • Isolated Runtime: Each AppAgent is dedicated to a single Windows application
  • Subtask Executor: Executes specific subtasks delegated by HostAgent
  • Application Expert: Tailored with deep knowledge of the target app's API surface, control semantics, and domain logic
  • Hybrid Execution: Leverages both GUI automation and API-based actions through MCP commands

Unlike monolithic Computer-Using Agents (CUAs) that treat all GUI contexts uniformly, each AppAgent is tailored to a single application and operates with specialized knowledge of its interface and capabilities.


Core Responsibilities

graph TB
    subgraph "AppAgent Core Responsibilities"
        SR[Sense:<br/>Capture Application State]
        RE[Reason:<br/>Analyze Next Action]
        EX[Execute:<br/>GUI or API Action]
        RP[Report:<br/>Write Results to Blackboard]
    end
    
    SR --> RE
    RE --> EX
    EX --> RP
    RP --> SR
    
    style SR fill:#e3f2fd
    style RE fill:#fff3e0
    style EX fill:#f1f8e9
    style RP fill:#fce4ec
ResponsibilityDescriptionExample
State SensingCapture application UI, detect controls, understand current stateScreenshot Word window → Detect 50 controls → Annotate UI elements
ReasoningAnalyze state and determine next action using LLM"Table visible with Export button [12] → Click to export data"
Action ExecutionExecute GUI clicks or API calls via MCP commandsclick_input(control_id=12) or execute_word_command("export_table")
Result ReportingWrite execution results to shared BlackboardWrite extracted data to subtask_result_1 for HostAgent

ReAct-Style Control Loop

Upon receiving a subtask and execution context from the HostAgent, the AppAgent initializes a ReAct-style control loop where it iteratively:

  1. Observes the current application state (screenshot + control detection)
  2. Thinks about the next step (LLM reasoning)
  3. Acts by executing either a GUI or API-based action (MCP commands)
sequenceDiagram
    participant HostAgent
    participant AppAgent
    participant Application
    participant Blackboard
    
    HostAgent->>AppAgent: Delegate subtask<br/>"Extract table from Word"
    
    loop ReAct Loop
        AppAgent->>Application: Observe (screenshot + controls)
        Application-->>AppAgent: UI state
        AppAgent->>AppAgent: Think (LLM reasoning)
        AppAgent->>Application: Act (click/API call)
        Application-->>AppAgent: Action result
    end
    
    AppAgent->>Blackboard: Write result
    AppAgent->>HostAgent: Return control

The MCP command system enables reliable control over dynamic and complex UIs by favoring structured API commands whenever available, while retaining fallback to GUI-based interaction commands when necessary.


Execution Architecture

Finite State Machine

AppAgent uses a finite state machine with 7 states to control its execution flow:

  • CONTINUE: Continue processing the current subtask
  • FINISH: Successfully complete the subtask
  • ERROR: Encounter an unrecoverable error
  • FAIL: Fail to complete the subtask
  • PENDING: Wait for user input or clarification
  • CONFIRM: Request user confirmation for sensitive actions
  • SCREENSHOT: Capture and re-annotate the application screenshot

State Details: See State Machine Documentation for complete state definitions and transitions.

4-Phase Processing Pipeline

Each execution round follows a 4-phase pipeline:

graph LR
    DC[Phase 1:<br/>DATA_COLLECTION<br/>Screenshot + Controls] --> LLM[Phase 2:<br/>LLM_INTERACTION<br/>Reasoning]
    LLM --> AE[Phase 3:<br/>ACTION_EXECUTION<br/>GUI/API Action]
    AE --> MU[Phase 4:<br/>MEMORY_UPDATE<br/>Record Action]
    
    style DC fill:#e1f5ff
    style LLM fill:#fff4e6
    style AE fill:#e8f5e9
    style MU fill:#fce4ec

Strategy Details: See Processing Strategy Documentation for complete pipeline implementation.


Hybrid GUI–API Execution

AppAgent executes actions through the MCP (Model-Context Protocol) command system, which provides a unified interface for both GUI automation and native API calls:

# GUI-based command (fallback)
command = Command(
    tool_name="click_input",
    parameters={"control_id": "12", "button": "left"}
)
await command_dispatcher.execute_commands([command])

# API-based command (preferred when available)
command = Command(
    tool_name="word_export_table",
    parameters={"format": "csv", "path": "output.csv"}
)
await command_dispatcher.execute_commands([command])

Implementation: See Hybrid Actions for details on the MCP command system.


Knowledge Enhancement

AppAgent is enhanced with Retrieval Augmented Generation (RAG) from heterogeneous sources:

Knowledge SourcePurposeConfiguration
Help DocumentsApplication-specific documentationLearning from Help Documents
Bing SearchLatest information and updatesLearning from Bing Search
Self-DemonstrationsSuccessful action trajectoriesExperience Learning
Human DemonstrationsExpert-provided workflowsLearning from Demonstrations

Knowledge Substrate Overview: See Knowledge Substrate for the complete RAG architecture.


Command System

AppAgent executes actions through the MCP (Model-Context Protocol) command system:

Application-Level Commands:

  • capture_window_screenshot - Capture application window
  • get_control_info - Detect UI controls via UIA/OmniParser
  • click_input - Click on UI control
  • set_edit_text - Type text into input field
  • annotation - Annotate screenshot with control labels

Command Details: See Command System Documentation for complete command reference.


Control Detection Backends

AppAgent supports multiple control detection backends for comprehensive UI understanding:

UIA (UI Automation):
Native Windows UI Automation API for standard controls

  • ✅ Fast and accurate
  • ✅ Works with most Windows applications
  • ❌ May miss custom controls

OmniParser (Visual Detection):
Vision-based grounding model for visual elements

  • ✅ Detects icons, images, custom controls
  • ✅ Works with web content
  • ❌ Requires external service

Hybrid (UIA + OmniParser):
Best of both worlds - maximum coverage

  • ✅ Native controls + visual elements
  • ✅ Comprehensive UI understanding

Control Detection Details: See Control Detection Overview.


Input and Output

AppAgent Input

InputDescriptionSource
User RequestOriginal user request in natural languageHostAgent
Sub-TaskSpecific subtask to executeHostAgent delegation
Application ContextTarget app name, window infoHostAgent
Control InformationDetected UI controls with labelsData collection phase
ScreenshotsClean, annotated, previous step imagesData collection phase
BlackboardShared memory for inter-agent communicationGlobal context
Retrieved KnowledgeHelp docs, demos, search resultsRAG system

AppAgent Output

OutputDescriptionConsumer
ObservationCurrent UI state descriptionLLM context
ThoughtReasoning about next actionExecution log
ControlLabelSelected control to interact withAction executor
FunctionMCP command to execute (click_input, set_edit_text, etc.)Command dispatcher
ArgsCommand parametersCommand dispatcher
StatusAgent state (CONTINUE, FINISH, etc.)State machine
Blackboard UpdateExecution resultsHostAgent

Example Output:

{
    "Observation": "Word document with table, Export button at [12]",
    "Thought": "Click Export to extract table data",
    "ControlLabel": "12",
    "Function": "click_input",
    "Args": {"button": "left"},
    "Status": "CONTINUE"
}

Detailed Documentation:

Core Features:

Tutorials:


API Reference

:::agents.agent.app_agent.AppAgent


Summary

AppAgent Key Characteristics:

Application-Specialized Worker: Dedicated to single Windows application
ReAct Control Loop: Iterative observe → think → act execution
Hybrid Execution: GUI automation + API calls via MCP commands
7-State FSM: Robust state management for execution control
4-Phase Pipeline: Structured data collection → reasoning → action → memory
Knowledge-Enhanced: RAG from docs, demos, and search
Orchestrated by HostAgent: Child agent in hierarchical architecture

Next Steps:

  1. Deep Dive: Read State Machine and Processing Strategy for implementation details
  2. Learn Features: Explore Core Features for advanced capabilities
  3. Hands-On Tutorial: Follow Creating AppAgent guide