Open Chrome CLI (Agent Browser Bridge)

May 24, 2026 · View on GitHub

中文版 | English

License Platform

A powerful bridge connecting AI Agents (Claude CodeCodexOpenclaw/Hermes.) with your live Chrome browser. Unlike Puppeteer-based solutions, this project runs directly in your real browser profile, allowing AI to assist you with genuine sessions, logins, and extensions.

📺 Demo Video

Watch the Demo on Bilibili

🚀 Key Features

  • Full Automation Toolkit: Over 25+ professional-grade tools for interaction, navigation, and debugging.
  • Genuine Context: Operates within your existing browser session (retains logins, cookies, and history).
  • Dual-Engine DOM Analysis: Uses both Accessibility Tree (AXTree) and DOM injection to provide AI with a perfect "semantic map" of any webpage.
  • Expert Debugging: Access to Network requests, Console logs, and deep Memory Heap analysis.
  • SEO & Performance: Built-in device emulation (Mobile, Googlebot) and full-page auditing.

⚖️ Tool Comparison

DimensionPlaywrightBrowser-UseChrome DevTools MCPOpenCLI (jackwener)OpenChromeCLI (justlovemaki)
Developer BackgroundMicrosoft Official Team.Independent community open-source team.Google Chrome DevTools Official Team.Senior database expert, Apache Software Foundation Top-Level Project PMC member.Independent Developer (justlovemaki / He Xi).
Core ArchitectureNode.js/TS driver relay based on CDP protocol.Python Agent framework (supports Playwright or native CDP).Native CDP mapped to standard MCP (Model Context Protocol) service.Lightweight browser bridge extension + local micro-daemon + YAML/TS declarative adapters.Browser-level extension (chrome.debugger) + Local Native Messaging Host + TCP/CLI bridge.
Control Object & IsolationStrongly Isolated Sandbox
Fresh, isolated virtual browser processes.
Independent Environment
Standalone sandboxed Chrome or remote cloud browser containers.
Debugging Environment
Starts or attaches to a Chrome browser in a Debug Session.
Real/Local Browser
Attaches to the user's local, logged-in Chrome instance and some Electron desktop apps.
Real/Local Browser
Directly attaches to your active real browser window, inheriting all logins, cookies, and extensions.
Multi-instance & ParallelismStrong Native Support
High-concurrency management via browser_context.
Moderate
Concurrency usually relies on controlling multiple sandboxed browsers, resulting in higher overhead.
Weak
Each MCP Server usually maps to one browser process; requires manual planning of CDP ports.
Weak
Focused on unified global CLI pipelines and site adapters, not multi-instance concurrency.
Extremely Strong (Multi-port parallelism)
Explicitly allocates and returns dynamic port numbers upon startup, supporting multiple independent relay instances via different ports to solve scheduling conflicts.
Remote Agent DebuggingSupported
Via connect_over_cdp to remote WebSockets, but requires exposed CDP ports or proxies.
Supported
Via cdp_url to cloud VMs.
Supported
Via specified remote CDP addresses/ports, typically for cloud build/test environments.
Weak
Focused on local CLI/Agent execution, lacking secure remote reverse debugging proxies.
Native Support (Remote Agent Debugging)
Uses encrypted tunnels to securely bridge the local logged-in Chrome debug channel to Agents on remote servers (e.g., Claude Code on GPU cloud) without exposing CDP ports.
Unique Features1. Cross-browser engine support (WebKit/Firefox/Chromium);
2. Mature auto-waiting mechanisms and recorder.
1. Deep integration with LLM Vision for semantic judgment;
2. Built-in LLM-based autonomous planning and error correction.
1. Official Lighthouse support;
2. Native network and CPU throttling simulation;
3. Native Trace performance analysis.
1. 80+ built-in site adapters, out-of-the-box;
2. AI-powered webpage-to-CLI command conversion;
3. Supports controlling desktop Electron apps.
1. Dynamic port allocation for script scheduling;
2. Multi-instance parallelism via different ports;
3. AXTree + DOM injection dual-semantic extraction;
4. Heap Snapshot export for deep debugging.
FocusSoftware Quality & Regression Testing.
Ensuring 100% reliability of code and webpage features.
Automated Business Process Replacement.
Enabling AI to navigate unknown websites like a human using vision and reasoning.
Frontend Dev Optimization & Diagnosis.
Giving AI assistants the power to detect performance bottlenecks and analyze request errors.
Global Web CLI & Deterministic Operation.
Eliminating LLM randomness by reducing complex tasks to stable CLI commands.
Real-device Remote Control & Deep Analysis.
Bridging the gap between local and cloud, allowing remote AI Agents to operate and diagnose local Chrome with high precision.
Usage CostNear zero (local execution, no API costs).Extremely high (requires frequent screenshot transfers to LLM).Low (included with IDE AI quotas or official MCP licensing).Low (mainly consumes lightweight structured text tokens).Low (mainly consumes text tokens; screenshots are triggered only when needed).

🛠️ Tool Reference

The AI Agent can autonomously invoke the following capabilities:

CategoryTools
Interactionclick, fillForm, hover, drag, uploadFile, typeText, pressKey, clickAt
NavigationnavigatePage, waitFor, selectPage, closePage, createTab
DebugginglistConsoleMessages, evaluateScript, listNetworkRequests, getNetworkResponseBody
Audit & SEOemulateDevice, resizePage, takeFullPageScreenshot, getCookies
MemorytakeHeapSnapshot, getHeapSnapshotSummary, getHeapSnapshotDetails, getHeapSnapshotRetainers

📦 Installation

1. Extension Setup

  1. Open chrome://extensions/.
  2. Enable "Developer mode".
  3. Click "Load unpacked" and select the dist folder.

2. Native Host Setup

The Native Host is required for secure communication between the Agent and Chrome.

  • Windows: Run host/install.bat
  • macOS/Linux: Run host/install.sh

🤖 Usage with AI Agents

More than just a physical relay, this project provides out-of-the-box "browser manipulation skills" for mainstream AI Agents. Empower your Agent with real execution capabilities:

  • Broad Agent Support: Native support for Claude Code, Codex, Openclaw/Hermes, and other mainstream AI frameworks. Provides human-like interaction, TCP instruction set integration, and real-time network/console monitoring.

Tip: The skills/browser-remote-control directory contains LLM-optimized Prompt Templates and tool definitions to help Agents better understand webpage context.

AI Agents can communicate with the browser by calling the local Command Line Interface (CLI) or connecting directly to the TCP port.

CLI Usage Examples:

  1. List all open tabs:
node skills/browser-remote-control/scripts/cli.js getTabs
  1. Open a new website:
node skills/browser-remote-control/scripts/cli.js createTab '{"url": "https://github.com"}'
  1. Read page content (Accessibility Tree):
# tabId can be obtained from getTabs
node skills/browser-remote-control/scripts/cli.js readPage '{"tabId": 12345, "filter": "interactive"}'
  1. Click an element (using UID from readPage):
node skills/browser-remote-control/scripts/cli.js click '{"tabId": 12345, "uid": "node-12"}'
  1. Type text into focused element:
node skills/browser-remote-control/scripts/cli.js typeText '{"tabId": 12345, "text": "Open Chrome CLI"}'
  1. Take a full-page screenshot:
node skills/browser-remote-control/scripts/cli.js takeFullPageScreenshot '{"tabId": 12345}'

Configuration Note: Developers can integrate the CLI tool into their Agent's toolbox, enabling the Agent to perform specific browser tasks.

🛡️ Security & Privacy

This extension operates via chrome.debugger. It can inspect and modify browser data. We recommend using an isolated profile for highly sensitive tasks, although the project is designed to help you with your daily workflows.


Part of the SEO Master Extension suite.