Browser Automation

  • Desktop app
  • All plans

What it does

Every project has a real Chromium browser embedded in the right panel, with its own persistent session (cookies, local storage) per project. Agents you grant the AgentsRoom browser capability to can drive that same browser through the AgentsRoom Browser MCP: navigate, click, type, take screenshots, run JavaScript, wait for an element, read the page state and the console. After every page-changing action the agent receives a screenshot inline, which is what makes agent-driven QA reliable. You watch the page the agent drives, and can take over at any moment. A dev agent can also hand a scenario to an ephemeral QA agent, which tests in the browser and returns a verdict.

Where to find it

  • Inside a project, the right panel has three tabs: Files, Changes, Tests. The Tests tab hosts the Browser (View and Trace) and the Electron testing surface. Opening a QA agent focuses it automatically.
  • Per agent: Edit agent > Capabilities > AgentsRoom browser (the neighbouring AgentsRoom in Chrome (extension) capability is a different feature, see AgentsRoom in Chrome (agents drive your own Chrome)). Shortcut: the identity chip at the top of the browser chrome ("Screen of X"), which opens the same grant / revoke list. "Restart the agent to apply."
  • Browser chrome: URL bar ("Enter URL or localhost:port"), back / forward, Reload without cache and Reload with cache, Copy screenshot to clipboard, Open in default browser, history (Recent, Clear), Point UI changes, Device frame, Screens loaded in the background, QA context: test data, URLs, instructions.

How to use it

  1. Open the Tests tab and type a URL, or let it pick one: a running preview tunnel, else a detected dev server, else https://localhost:3000. The browser runs with the HTTP cache off, so a rebuilding dev server is always fetched fresh.
  2. Grant AgentsRoom browser to the agent that should test (QA Engineers have it by design), then ask it in plain words: "sign up with a test email, verify the confirmation screen". Its navigate / click / evaluate calls appear in the Trace view.
  3. Point UI changes: click an element of the live page, write the change you want in the bubble, repeat, then Send to one or several agents or to the backlog. Works inside iframes and on sites that block in-page scripts (the tray tells you when a page refused the pointer; Try again).
  4. Device frame: pick a phone or tablet (rotate to landscape or portrait). The page is really emulated (viewport, touch, user agent, pixel ratio), not just shrunk; combine it with the pointer to point changes on the mobile layout.
  5. QA context (Datasets, URLs, Instructions): store logins, known URLs and project-wide test rules once; every browser-capable agent receives them, so "log in" needs no explanation. Stored on your account per project, not in the repo.
  6. Delegation: a dev agent calls the QA test runner; a temporary QA agent (cheaper model) runs the scenario in the browser and reports pass / fail / inconclusive before closing.

Settings

  • askBeforeTestTools (global, "Ask before test tools"): agents ask before driving the browser, the running app or the QA runner. A project can override it, and a single agent can be exempted.
  • AgentsRoom browser itself is a per-agent capability (browserAccess), saved with the agent (also editable through agents_save).

Agent tools (MCP)

  • browser_navigate, browser_click, browser_type, browser_screenshot, browser_evaluate, browser_wait_for, browser_get_state, browser_get_logs (console log / warn / error of the page), browser_set_viewport, browser_go_back, browser_go_forward, browser_reload. Page-changing calls return a PNG screenshot (capped at 1.6 MB, replaced by a text marker above that).
  • browser_set_viewport resizes the page viewport, so an agent can check a layout at a width other than the panel's: a device preset (iPhone SE / 15 / 15 Pro Max, Galaxy S24, Pixel 8, iPad mini, iPad Pro 11) for real mobile semantics (<meta name="viewport"> honoured, touch, devicePixelRatio, mobile user agent), a raw width (height defaults to 900) for a desktop breakpoint, or reset: true to go back. Widths and heights are CSS pixels, 200 to 4096. The panel resizes with it and the Device menu shows what is applied: a preset appears ticked, a free size appears as a Custom size entry with its dimensions, and Desktop puts the page back to the panel-sized view. The size sticks until it is reset.
  • Selectors (and browser_read_page) go through open shadow roots, like Playwright: #login_email finds a field inside a web component. host >>> inner scopes a selector to one host's shadow root; a pierce/ prefix is accepted.
  • browser_click / browser_type with trusted: true send real mouse and keyboard events (Chromium's own input pipeline over CDP) instead of synthetic ones: what canvas-rendered apps (Compose Multiplatform for Web, Flutter Web) need. browser_click also takes x / y (viewport CSS pixels) to click a point, and force: true to click an element's centre when something covers it (an accessibility overlay with pointer-events: none over the canvas). A trusted browser_type without a selector types into whatever has focus; clear selects all then deletes, submit presses Enter. Embedded browser only; Chrome control accepts the shadow DOM selectors but not trusted input yet.
  • run_qa_test (Test Runner): hands a scenario to an ephemeral QA agent. Exposed to every agent, but only to be called on an explicit request.
  • submit_verdict (QA Tester): how the QA agent returns pass / fail / inconclusive.

Providers

Works with every CLI that reads MCP servers: Claude Code, Codex CLI, GitHub Copilot CLI, Antigravity CLI, OpenCode, Cursor, Grok Build, Mistral Vibe, Kimi Code, Amp, oh-my-pi, Freebuff, Devin. Aider has no MCP support. An agent without the AgentsRoom browser capability has no browser_* tools and delegates testing to the QA agent; an agent with the capability drives the browser directly (restart it after granting).

Mobile

Not on the mobile companion. The phone can open a tunneled preview in its own browser, but cannot drive the embedded one.

The Allow a test tool? question (setting askBeforeTestTools) does appear on the phone since 2026-09-26, above the composer of the agent that asked: same tool, same arguments, same "Stop asking for this project / agent" boxes and countdown. Allowing there lets the agent drive the browser on the computer.

Limits

  • One screen is shown at a time; the chip names whose screen it is, and a banner warns when you look at another agent's page than the one you are talking to ("Screen of X. You are talking to Y").
  • Each agent screen keeps a Chromium process alive in the background; close them from Screens loaded in the background.
  • The browser bridge listens on the local loopback only, with a token regenerated at every launch; the project's .mcp.json entry is rewritten when the app starts.
  • Web only today. Driving Electron apps needs the open-source electron-mcp package inside your app; React Native is on the roadmap.
  • Smaller models are fragile for multi-step browser flows; a mid-tier model is recommended for QA runs.

Common questions

  • How is it different from Playwright MCP? Same browser you see, persistent login per project, visible in real time, and you can take over. No fresh headless instance per call.
  • Can the agent fill a login form and stay logged in? Yes, cookies persist per project; they never leak to another project.
  • Do I have to start the server myself? The embedded browser never starts one: localhost:3000 in the URL bar is only a default. Either start the app from Dev commands (or a terminal) before asking the agent to test, or let the agent do it: browser_navigate on a dead local port answers "Nothing is listening on localhost:PORT" and tells the agent to start the saved dev command (commands_list / commands_run). Save the command that serves your app (e.g. php -S localhost:8000) once in Dev commands, and give the right URL in QA context or in the team step's Browser URL field.
  • Why does my dev agent refuse to test itself? It has no browser access: tick AgentsRoom browser (Edit agent > Capabilities) and restart it, or let it delegate to the QA agent.
  • Can an agent test a Compose Multiplatform or Flutter Web app? Yes, since 2026-09-29: selectors reach the elements inside the app's shadow root, and the agent passes trusted: true (plus force: true over the canvas overlay) so the canvas receives real clicks and keystrokes. It logs in and walks the screens without a Playwright script.
  • browser_type typed but the form ignored it? Fixed: typing goes through the native setter so React state follows; the tool warns when the page state stayed empty.