# Browser Automation

> Every project has a real Chromium browser embedded in the right panel, with its own persistent session (cookies, local storage) per project.

- Area: Desktop app
- Plans: All plans
- Last checked against the product: 2026-09-29
- Web page: https://agentsroom.dev/docs/browser-automation

## What it does

Every project has a real Chromium browser embedded in the right panel, with its own persistent session (cookies, local storage) per project. Agents you grant the **AgentsRoom browser** capability to can drive that same browser through the AgentsRoom Browser MCP: navigate, click, type, take screenshots, run JavaScript, wait for an element, read the page state and the console. After every page-changing action the agent receives a screenshot inline, which is what makes agent-driven QA reliable. You watch the page the agent drives, and can take over at any moment. A dev agent can also hand a scenario to an ephemeral QA agent, which tests in the browser and returns a verdict.

## Where to find it

- Inside a project, the right panel has three tabs: **Files**, **Changes**, **Tests**. The Tests tab hosts the **Browser** (View and Trace) and the Electron testing surface. Opening a QA agent focuses it automatically.
- Per agent: Edit agent > Capabilities > **AgentsRoom browser** (the neighbouring **AgentsRoom in Chrome (extension)** capability is a different feature, see [AgentsRoom in Chrome (agents drive your own Chrome)](https://agentsroom.dev/docs/chrome-control.md)). Shortcut: the identity chip at the top of the browser chrome ("Screen of X"), which opens the same grant / revoke list. "Restart the agent to apply."
- Browser chrome: URL bar ("Enter URL or localhost:port"), back / forward, **Reload without cache** and **Reload with cache**, **Copy screenshot to clipboard**, **Open in default browser**, history (**Recent**, **Clear**), **Point UI changes**, **Device frame**, **Screens loaded in the background**, **QA context: test data, URLs, instructions**.

## How to use it

1. Open the Tests tab and type a URL, or let it pick one: a running preview tunnel, else a detected dev server, else `https://localhost:3000`. The browser runs with the HTTP cache off, so a rebuilding dev server is always fetched fresh.
2. Grant **AgentsRoom browser** to the agent that should test (QA Engineers have it by design), then ask it in plain words: "sign up with a test email, verify the confirmation screen". Its navigate / click / evaluate calls appear in the **Trace** view.
3. **Point UI changes**: click an element of the live page, write the change you want in the bubble, repeat, then **Send to** one or several agents or to the backlog. Works inside iframes and on sites that block in-page scripts (the tray tells you when a page refused the pointer; **Try again**).
4. **Device frame**: pick a phone or tablet (rotate to landscape or portrait). The page is really emulated (viewport, touch, user agent, pixel ratio), not just shrunk; combine it with the pointer to point changes on the mobile layout.
5. **QA context** (Datasets, URLs, Instructions): store logins, known URLs and project-wide test rules once; every browser-capable agent receives them, so "log in" needs no explanation. Stored on your account per project, not in the repo.
6. Delegation: a dev agent calls the QA test runner; a temporary QA agent (cheaper model) runs the scenario in the browser and reports pass / fail / inconclusive before closing.

## Settings

- `askBeforeTestTools` (global, "Ask before test tools"): agents ask before driving the browser, the running app or the QA runner. A project can override it, and a single agent can be exempted.
- **AgentsRoom browser** itself is a per-agent capability (`browserAccess`), saved with the agent (also editable through `agents_save`).

## Agent tools (MCP)

- `browser_navigate`, `browser_click`, `browser_type`, `browser_screenshot`, `browser_evaluate`, `browser_wait_for`, `browser_get_state`, `browser_get_logs` (console log / warn / error of the page), `browser_set_viewport`, `browser_go_back`, `browser_go_forward`, `browser_reload`. Page-changing calls return a PNG screenshot (capped at 1.6 MB, replaced by a text marker above that).
- `browser_set_viewport` resizes the page viewport, so an agent can check a layout at a width other than the panel's: a device preset (iPhone SE / 15 / 15 Pro Max, Galaxy S24, Pixel 8, iPad mini, iPad Pro 11) for real mobile semantics (`<meta name="viewport">` honoured, touch, devicePixelRatio, mobile user agent), a raw width (height defaults to 900) for a desktop breakpoint, or `reset: true` to go back. Widths and heights are CSS pixels, 200 to 4096. The panel resizes with it and the **Device** menu shows what is applied: a preset appears ticked, a free size appears as a **Custom size** entry with its dimensions, and **Desktop** puts the page back to the panel-sized view. The size sticks until it is reset.
- Selectors (and `browser_read_page`) go through open shadow roots, like Playwright: `#login_email` finds a field inside a web component. `host >>> inner` scopes a selector to one host's shadow root; a `pierce/` prefix is accepted.
- `browser_click` / `browser_type` with `trusted: true` send real mouse and keyboard events (Chromium's own input pipeline over CDP) instead of synthetic ones: what canvas-rendered apps (Compose Multiplatform for Web, Flutter Web) need. `browser_click` also takes `x` / `y` (viewport CSS pixels) to click a point, and `force: true` to click an element's centre when something covers it (an accessibility overlay with `pointer-events: none` over the canvas). A trusted `browser_type` without a selector types into whatever has focus; `clear` selects all then deletes, `submit` presses Enter. Embedded browser only; Chrome control accepts the shadow DOM selectors but not trusted input yet.
- `run_qa_test` (Test Runner): hands a scenario to an ephemeral QA agent. Exposed to every agent, but only to be called on an explicit request.
- `submit_verdict` (QA Tester): how the QA agent returns pass / fail / inconclusive.

## Providers

Works with every CLI that reads MCP servers: Claude Code, Codex CLI, GitHub Copilot CLI, Antigravity CLI, OpenCode, Cursor, Grok Build, Mistral Vibe, Kimi Code, Amp, oh-my-pi, Freebuff, Devin. Aider has no MCP support. An agent without the **AgentsRoom browser** capability has no `browser_*` tools and delegates testing to the QA agent; an agent with the capability drives the browser directly (restart it after granting).

## Mobile

Not on the mobile companion. The phone can open a tunneled preview in its own browser, but cannot drive the embedded one.

The **Allow a test tool?** question (setting `askBeforeTestTools`) does appear on the phone since 2026-09-26, above the composer of the agent that asked: same tool, same arguments, same "Stop asking for this project / agent" boxes and countdown. Allowing there lets the agent drive the browser on the computer.

## Limits

- One screen is shown at a time; the chip names whose screen it is, and a banner warns when you look at another agent's page than the one you are talking to ("Screen of X. You are talking to Y").
- Each agent screen keeps a Chromium process alive in the background; close them from **Screens loaded in the background**.
- The browser bridge listens on the local loopback only, with a token regenerated at every launch; the project's `.mcp.json` entry is rewritten when the app starts.
- Web only today. Driving Electron apps needs the open-source electron-mcp package inside your app; React Native is on the roadmap.
- Smaller models are fragile for multi-step browser flows; a mid-tier model is recommended for QA runs.

## Common questions

- **How is it different from Playwright MCP?** Same browser you see, persistent login per project, visible in real time, and you can take over. No fresh headless instance per call.
- **Can the agent fill a login form and stay logged in?** Yes, cookies persist per project; they never leak to another project.
- **Do I have to start the server myself?** The embedded browser never starts one: `localhost:3000` in the URL bar is only a default. Either start the app from **Dev commands** (or a terminal) before asking the agent to test, or let the agent do it: `browser_navigate` on a dead local port answers "Nothing is listening on localhost:PORT" and tells the agent to start the saved dev command (`commands_list` / `commands_run`). Save the command that serves your app (e.g. `php -S localhost:8000`) once in Dev commands, and give the right URL in **QA context** or in the team step's Browser URL field.
- **Why does my dev agent refuse to test itself?** It has no browser access: tick **AgentsRoom browser** (Edit agent > Capabilities) and restart it, or let it delegate to the QA agent.
- **Can an agent test a Compose Multiplatform or Flutter Web app?** Yes, since 2026-09-29: selectors reach the elements inside the app's shadow root, and the agent passes `trusted: true` (plus `force: true` over the canvas overlay) so the canvas receives real clicks and keystrokes. It logs in and walks the screens without a Playwright script.
- **`browser_type` typed but the form ignored it?** Fixed: typing goes through the native setter so React state follows; the tool warns when the page state stayed empty.

## Related

- [Agent Delegation](https://agentsroom.dev/docs/agent-delegation.md): the dev to QA handoff in detail.
- [Localhost Tunnel](https://agentsroom.dev/docs/localhost-tunnel.md): a public URL the browser targets automatically.
- [Dev Terminals](https://agentsroom.dev/docs/dev-terminals.md): start the server the agent tests.
- [Chrome Extension](https://agentsroom.dev/docs/chrome-extension.md): point changes from your own Chrome instead.
- [AgentsRoom in Chrome (agents drive your own Chrome)](https://agentsroom.dev/docs/chrome-control.md): drive your own Chrome (logins, profiles, another machine) instead of this embedded one.
- [Agent Teams](https://agentsroom.dev/docs/teams.md): a QA node gets browser access automatically.
