# Agent Delegation

> A developer agent that has just finished a feature does not test it itself: it calls one tool, run_qa_test, on the AgentsRoom Test Runner MCP server.

- Area: Desktop app
- Plans: All plans
- Last checked against the product: 2026-09-13
- Web page: https://agentsroom.dev/docs/agent-delegation

## What it does

A developer agent that has just finished a feature does not test it itself: it calls one tool, `run_qa_test`, on the AgentsRoom Test Runner MCP server. AgentsRoom then spawns a short-lived QA agent, usually on a cheaper model, which drives the embedded browser, runs the scenario, returns a verdict (pass, fail or inconclusive with a summary, an optional screenshot path and optional console logs) and is destroyed. The developer agent reads the short verdict and moves on; screenshots and DOM dumps never enter its context.

The constraint is mechanical: the browser tools are removed from the developer agent at spawn time, so the only way it can test in a browser is to delegate.

## Where to find it

There is no button. Delegation happens from inside an agent's conversation: the developer agent calls the tool when it needs a browser test. The QA agent appears as a temporary agent in the project's agent list for the duration of the test, then disappears. The project's QA context (datasets, URLs, instructions the QA agent reads) is edited from the QA agent's **QA context** modal (tabs **Datasets**, **URLs**, **Instructions**).

## How to use it

1. Give a developer agent a task that ends with a browser check ("implement the login form, then have it tested").
2. When the agent calls `run_qa_test`, AgentsRoom asks for your approval first unless you turned that prompt off (see Settings).
3. The QA agent starts, opens the Browser tab, runs the scenario and reports back. Watch it in its own tab if you want.
4. The developer agent receives the verdict and continues. On failure, the QA agent can also file a backlog ticket carrying the long-form details (scenario, screenshot, console logs).

To make QA runs reliable, fill the **QA context** once per project: named datasets ("admin user" with email and password), pinned URLs (login page, staging), and free-form instructions prepended to every QA session.

## Settings

- `askBeforeTestTools` (global, Settings; project override in Project settings): ask before an agent drives a test or automation MCP tool, including the QA test runner. A single agent can be exempted.

## Agent tools (MCP)

None on the `AgentsRoom-MCP` server. The delegation call is `run_qa_test` on the separate **AgentsRoom Test Runner** MCP server, and the QA agent answers through `submit_verdict` on the **AgentsRoom QA Tester** server. The QA agent also gets the `backlog_*` tools to file a ticket on failure.

## Providers

Cross-provider: the developer and the QA agent can run on different CLIs and models. A common split is a large model on the developer side (Claude Opus, Codex) and a small one on the QA side (Claude Haiku, Codex mini). The QA agent needs a provider that can use the AgentsRoom browser MCP tools.

## Mobile

Not on the mobile companion. The QA agent shows up in the agent list like any temporary agent while it runs, and its verdict lands in the developer agent's conversation, which the phone can read.

## Limits

- Browser tests only today. Electron app tests (via an AgentsRoom Electron MCP library) and React Native tests are on the roadmap.
- Everything runs on your machine: the developer agent, the QA agent, the MCP bridge and the browser. Only the model calls go to the provider.
- The QA agent tests the developer's checkout: when the developer works in a worktree, the QA agent inherits that worktree rather than getting a fresh one, so it sees uncommitted work.
- One run is capped at 10 minutes by default. The `timeoutMs` argument of `run_qa_test` can only raise that cap: a shorter value is ignored, because a QA agent needs the default just to boot its CLI, connect the browser and read the scenario. A run that hits the cap comes back with the status `timeout` and a summary that says whether the QA agent had started (it started but never submitted a verdict) or never finished starting (no verdict was ever possible).

## Common questions

- **Why not let the developer agent test itself?** Cost (a QA pass on a small model is roughly ten times cheaper than on the large one), context (the developer stays clean), reliability (a single-purpose tester clicks better than a multitasking developer).
- **Does this need a team graph?** No. It is the Dev to QA flavour of a team exposed as a single tool call. For multi-step pipelines with conditions and loops, use Agent Teams.
- **Can I stop the approval prompt?** Yes, turn off "ask before test tools" globally, per project, or exempt one agent.
- **The developer agent only gets "timed out" and never a verdict.** Read the summary of the timeout: "before the QA Bot finished starting" means the QA agent never came up (check that its CLI is installed and signed in), "started but never submitted a verdict" means the scenario, the page or the browser kept it busy past the cap. Asking for a cap below 10 minutes used to end every run this way; those values are now ignored.

## Related

- [Agent Teams](https://agentsroom.dev/docs/teams.md): the full graph-based orchestration.
- [Browser Automation](https://agentsroom.dev/docs/browser-automation.md): the embedded browser the QA agent drives.
- [AgentsRoom MCP](https://agentsroom.dev/docs/agentsroom-mcp.md): the MCP servers that give agents their tools.
- [Git Worktrees](https://agentsroom.dev/docs/worktrees.md): why the QA agent shares its master's checkout.
