🤖 Agent API Reference
March 13, 2026 · View on GitHub
Overview
The Agent module provides specialized functionality for the AgentBay cloud platform. It includes various methods and utilities to interact with cloud services and manage resources.
📚 Tutorial
Learn about agent modules and custom agents
Agent
An Agent to manipulate applications to complete specific tasks.
⚠️ Note: Currently, for agent services (including ComputerUseAgent, BrowserUseAgent, and MobileUseAgent), we do not provide services for overseas users registered with alibabacloud.com.
Constructor
public Agent(Session session)
Methods
getComputer
public Computer getComputer()
Get the Computer agent for desktop task execution.
Returns:
Computer: Computer agent instance
getBrowser
public Browser getBrowser()
Get the Browser agent for browser task execution.
Returns:
Browser: Browser agent instance
getMobile
public Mobile getMobile()
Get the Mobile agent for mobile device task execution.
Returns:
Mobile: Mobile agent instance
SchemaHelper
Methods
generateJsonSchema
public static String generateJsonSchema(Class<?> schemaClass)
Computer
An Agent to perform tasks on the computer.
⚠️ Note: Currently, for agent services (including ComputerUseAgent, BrowserUseAgent, and MobileUseAgent), we do not provide services for overseas users registered with alibabacloud.com.
Constructor
public Computer(Session session)
Methods
executeTask
public ExecutionResult executeTask(String task)
Execute a task in human language without waiting for completion (non-blocking).
This is a fire-and-return interface that immediately provides a task ID. Call getTaskStatus to check the task status. You can control the timeout of the task execution in your own code by setting the frequency of calling getTaskStatus.
Parameters:
task(String): Task description in human language
Returns:
ExecutionResult: ExecutionResult containing success status, task ID, task status, and error message if any
executeTaskAndWait
public ExecutionResult executeTaskAndWait(String task, int timeout)
Execute a specific task described in human language synchronously.
This is a synchronous interface that blocks until the task is completed or an error occurs, or timeout happens. The default polling interval is 3 seconds.
Parameters:
task(String): Task description in human languagetimeout(int): Maximum time to wait for task completion in seconds
Returns:
ExecutionResult: ExecutionResult containing success status, task ID, task status, task result, and error message if any
getTaskStatus
public QueryResult getTaskStatus(String taskId)
Get the status of the task with the given task ID.
Parameters:
taskId(String): The ID of the task to query
Returns:
QueryResult: QueryResult containing success status, task status, task action, task product, and error message if any
terminateTask
public ExecutionResult terminateTask(String taskId)
Terminate a task with a specified task ID.
Parameters:
taskId(String): The ID of the running task to terminate
Returns:
ExecutionResult: ExecutionResult containing success status, task ID, task status, and error message if any
Browser
Browser agent for browser automation with natural language.
⚠️ Note: Currently, for agent services (including ComputerUseAgent, BrowserUseAgent, and MobileUseAgent), we do not provide services for overseas users registered with alibabacloud.com.
Constructor
public Browser(Session session)
Methods
initialize
public boolean initialize(BrowserOption option)
Initialize the browser on which the agent performs tasks. You are supposed to call this API before executeTask is called, but it's not optional. If you want perform a hybrid usage of browser, you must call this API before executeTask is called.
Parameters:
option(BrowserOption): the browser initialization options. If {@code null}, default options will be used
Returns:
boolean: {@code true} if the browser is successfully initialized, {@code false} otherwise
executeTask
public ExecutionResult executeTask(String task, boolean useVision, Object outputSchema, boolean fullPageScreenShot)
Execute a task described in human language on a browser without waiting for completion (non-blocking).
This is a fire-and-return interface that immediately provides a task ID. Call get_task_status to check the task status. You can control the timeout of the task execution in your own code by setting the frequency of calling get_task_status.
Parameters:
task(String): Task description in human languageuseVision(boolean): Whether to use vision to performe the taskoutputSchema(Object): The schema of the structured outputfullPageScreenShot(boolean): Whether to take a full page screenshot. This only works when use_vision is true. When use_vision is enabled, we need to provide a screenshot of the webpage to the LLM for grounding. There are two ways of screenshot: 1. Full-page screenshot: Captures the entire webpage content, including parts not currently visible in the viewport. 2. Viewport screenshot: Captures only the currently visible portion of the webpage. The first approach delivers all information to the LLM in one go, which can improve task success rates in certain information extraction scenarios. However, it also results in higher token consumption and increases the LLM's processing time. Therefore, we would like to give you the choice—you can decide whether to enable full-page screenshot based on your actual needs.
Returns:
ExecutionResult: ExecutionResult Result object containing success status, task ID, task status, and error message if any
executeTaskAndWait
public ExecutionResult executeTaskAndWait(String task, int timeout, boolean useVision, Object outputSchema, boolean fullPageScreenShot)
Execute a specific task described in human language synchronously.
This is a synchronous interface that blocks until the task is completed or an error occurs, or timeout happens. The default polling interval is 3 seconds.
Parameters:
-
task(String): Task description in human language -
timeout(int): Maximum time to wait for task completion in seconds -
useVision(boolean): Whether to use vision in the task -
outputSchema(Object): Optional Zod schema for a structured task output if you need -
fullPageScreenShot(boolean): Whether to take a full page screenshot, this only works if useVision is trueWhen use_vision is enabled, we need to provide a screenshot of the webpage to the LLM for grounding. There are two ways of screenshot: 1. Full-page screenshot: Captures the entire webpage content, including parts not currently visible in the viewport. 2. Viewport screenshot: Captures only the currently visible portion of the webpage. The first approach delivers all information to the LLM in one go, which can improve task success rates in certain information extraction scenarios. However, it also results in higher token consumption and increases the LLM's processing time. Therefore, we would like to give you the choice—you can decide whether to enable full-page screenshot based on your actual needs.
Returns:
ExecutionResult: ExecutionResult containing success status, task ID, task status, task result, and error message if any
getTaskStatus
public QueryResult getTaskStatus(String taskId)
Get the status of the task with the given task ID.
Parameters:
taskId(String): The ID of the task to query
Returns:
QueryResult: QueryResult containing success status, task status, task action, task product, and error message if any
terminateTask
public ExecutionResult terminateTask(String taskId)
Terminate a task with a specified task ID.
Parameters:
taskId(String): The ID of the running task to terminate
Returns:
ExecutionResult: ExecutionResult containing success status, task ID, task status, and error message if any
Mobile
Mobile agent for mobile device automation with natural language. Uses execute_task, get_task_status, terminate_task MCP tools for task execution.
Constructor
public Mobile(Session session)
Methods
executeTask
public TaskExecution executeTask(String task)
public TaskExecution executeTask(String task, MobileTaskOptions options)
Execute a mobile task in human language without waiting for completion (non-blocking).
When options has streaming callbacks, uses WebSocket streaming for real-time events. Otherwise uses MCP call and polling.
Parameters:
task(String): Task description in human languageoptions(MobileTaskOptions): Optional MobileTaskOptions (maxSteps, streaming callbacks)
Returns:
TaskExecution: TaskExecution handle for the running task
executeTaskAndWait
public ExecutionResult executeTaskAndWait(String task, int timeout)
public ExecutionResult executeTaskAndWait(String task, int timeout, MobileTaskOptions options)
Execute a task synchronously with optional WebSocket streaming. When {@code options} has streaming params, uses WS streaming for real-time events.
Parameters:
task(String): Task description in human languagetimeout(int): Maximum time to wait for task completion in secondsoptions(MobileTaskOptions): Optional MobileTaskOptions (maxSteps, streaming callbacks)
Returns:
ExecutionResult: ExecutionResult containing success status, task ID, task status, task result, and error message if any
getTaskStatus
public QueryResult getTaskStatus(String taskId)
Get the status of the task with the given task ID.
Parameters:
taskId(String): The ID of the task to query
Returns:
QueryResult: QueryResult containing success status, task status, task action, task product, stream, error, and error message if any
terminateTask
public ExecutionResult terminateTask(String taskId)
Terminate a task with a specified task ID.
Parameters:
taskId(String): The ID of the running task to terminate
Returns:
ExecutionResult: ExecutionResult containing success status, task ID, task status, and error message if any
AgentEvent
Represents a streaming event from an Agent execution.
Event types map directly to LLM output field names:
- "reasoning": from LLM reasoning_content (model's internal reasoning/thinking)
- "content": from LLM content (model's text output, intermediate analysis or final answer)
- "tool_call": from LLM tool_calls (tool invocation request)
- "tool_result": tool execution result
- "error": execution error
The {@code result} field in tool_result events carries an agent-defined structure that the SDK passes through without parsing. Typical fields include {@code isError} (boolean), {@code output} (string), and optionally {@code screenshot} (base64 string). The final task outcome is delivered via the {@link com.aliyun.agentbay.model.ExecutionResult} return value of {@code executeTaskAndWait}.
Constructor
public AgentEvent()
public AgentEvent(String type, int seq, int round)
Methods
getSeq
public int getSeq()
setSeq
public void setSeq(int seq)
getRound
public int getRound()
setRound
public void setRound(int round)
getContent
public String getContent()
setContent
public void setContent(String content)
getToolCallId
public String getToolCallId()
setToolCallId
public void setToolCallId(String toolCallId)
getToolName
public String getToolName()
setToolName
public void setToolName(String toolName)
getArgs
public Map<String, Object> getArgs()
setArgs
public void setArgs(Map<String, Object> args)
getResult
public Map<String, Object> getResult()
setResult
public void setResult(Map<String, Object> result)
getError
public Map<String, Object> getError()
setError
public void setError(Map<String, Object> error)
StreamOptions
Options for WebSocket streaming execution of agent tasks.
When any callback is set, the SDK uses the WebSocket streaming channel for real-time event delivery instead of HTTP polling.
Constructor
public StreamOptions()
Methods
getOnReasoning
public Consumer<AgentEvent> getOnReasoning()
Returns the callback for reasoning events (LLM reasoning_content).
setOnReasoning
public void setOnReasoning(Consumer<AgentEvent> onReasoning)
getOnContent
public Consumer<AgentEvent> getOnContent()
Returns the callback for content events (LLM content output).
setOnContent
public void setOnContent(Consumer<AgentEvent> onContent)
getOnToolCall
public Consumer<AgentEvent> getOnToolCall()
Returns the callback for tool_call events.
setOnToolCall
public void setOnToolCall(Consumer<AgentEvent> onToolCall)
getOnToolResult
public Consumer<AgentEvent> getOnToolResult()
Returns the callback for tool_result events.
setOnToolResult
public void setOnToolResult(Consumer<AgentEvent> onToolResult)
getOnError
public Consumer<AgentEvent> getOnError()
Returns the callback for error events.
setOnError
public void setOnError(Consumer<AgentEvent> onError)
hasStreamingParams
public boolean hasStreamingParams()
Returns true if streaming should be used (any callback is set).
builder
public static Builder builder()
Builder for StreamOptions.
Builder
Methods
onReasoning
public Builder onReasoning(Consumer<AgentEvent> onReasoning)
onContent
public Builder onContent(Consumer<AgentEvent> onContent)
onToolCall
public Builder onToolCall(Consumer<AgentEvent> onToolCall)
onToolResult
public Builder onToolResult(Consumer<AgentEvent> onToolResult)
onError
public Builder onError(Consumer<AgentEvent> onError)
build
public StreamOptions build()
MobileTaskOptions
Options for mobile task execution, including streaming callbacks. Extends StreamOptions with mobile-specific options like maxSteps and onCallForUser.
Constructor
public MobileTaskOptions()
Methods
getMaxSteps
public int getMaxSteps()
setMaxSteps
public void setMaxSteps(int maxSteps)
getOnCallForUser
public Function<AgentEvent, String> getOnCallForUser()
setOnCallForUser
public void setOnCallForUser(Function<AgentEvent, String> onCallForUser)
hasStreamingParams
public boolean hasStreamingParams()
mobileBuilder
public static MobileBuilder mobileBuilder()
MobileBuilder
Methods
maxSteps
public MobileBuilder maxSteps(int maxSteps)
onReasoning
public MobileBuilder onReasoning(java.util.function.Consumer<AgentEvent> cb)
onContent
public MobileBuilder onContent(java.util.function.Consumer<AgentEvent> cb)
onToolCall
public MobileBuilder onToolCall(java.util.function.Consumer<AgentEvent> cb)
onToolResult
public MobileBuilder onToolResult(java.util.function.Consumer<AgentEvent> cb)
onError
public MobileBuilder onError(java.util.function.Consumer<AgentEvent> cb)
onCallForUser
public MobileBuilder onCallForUser(Function<AgentEvent, String> cb)
build
public MobileTaskOptions build()
TaskExecution
Represents a running task that can be waited on for its final result. Returned by Mobile.executeTask() when the task is started.
Constructor
public TaskExecution(String taskId, CompletableFuture<ExecutionResult> resultFuture)
public TaskExecution(String taskId, CompletableFuture<ExecutionResult> resultFuture, Runnable cancelFn)
Methods
getTaskId
public String getTaskId()
wait
public ExecutionResult wait(int timeout)
Block until the task finishes or the timeout (in seconds) is reached.