Agent Delegation

  • Desktop app
  • All plans

What it does

A developer agent that has just finished a feature does not test it itself: it calls one tool, run_qa_test, on the AgentsRoom Test Runner MCP server. AgentsRoom then spawns a short-lived QA agent, usually on a cheaper model, which drives the embedded browser, runs the scenario, returns a verdict (pass, fail or inconclusive with a summary, an optional screenshot path and optional console logs) and is destroyed. The developer agent reads the short verdict and moves on; screenshots and DOM dumps never enter its context.

The constraint is mechanical: the browser tools are removed from the developer agent at spawn time, so the only way it can test in a browser is to delegate.

Where to find it

There is no button. Delegation happens from inside an agent's conversation: the developer agent calls the tool when it needs a browser test. The QA agent appears as a temporary agent in the project's agent list for the duration of the test, then disappears. The project's QA context (datasets, URLs, instructions the QA agent reads) is edited from the QA agent's QA context modal (tabs Datasets, URLs, Instructions).

How to use it

  1. Give a developer agent a task that ends with a browser check ("implement the login form, then have it tested").
  2. When the agent calls run_qa_test, AgentsRoom asks for your approval first unless you turned that prompt off (see Settings).
  3. The QA agent starts, opens the Browser tab, runs the scenario and reports back. Watch it in its own tab if you want.
  4. The developer agent receives the verdict and continues. On failure, the QA agent can also file a backlog ticket carrying the long-form details (scenario, screenshot, console logs).

To make QA runs reliable, fill the QA context once per project: named datasets ("admin user" with email and password), pinned URLs (login page, staging), and free-form instructions prepended to every QA session.

Settings

  • askBeforeTestTools (global, Settings; project override in Project settings): ask before an agent drives a test or automation MCP tool, including the QA test runner. A single agent can be exempted.

Agent tools (MCP)

None on the AgentsRoom-MCP server. The delegation call is run_qa_test on the separate AgentsRoom Test Runner MCP server, and the QA agent answers through submit_verdict on the AgentsRoom QA Tester server. The QA agent also gets the backlog_* tools to file a ticket on failure.

Providers

Cross-provider: the developer and the QA agent can run on different CLIs and models. A common split is a large model on the developer side (Claude Opus, Codex) and a small one on the QA side (Claude Haiku, Codex mini). The QA agent needs a provider that can use the AgentsRoom browser MCP tools.

Mobile

Not on the mobile companion. The QA agent shows up in the agent list like any temporary agent while it runs, and its verdict lands in the developer agent's conversation, which the phone can read.

Limits

  • Browser tests only today. Electron app tests (via an AgentsRoom Electron MCP library) and React Native tests are on the roadmap.
  • Everything runs on your machine: the developer agent, the QA agent, the MCP bridge and the browser. Only the model calls go to the provider.
  • The QA agent tests the developer's checkout: when the developer works in a worktree, the QA agent inherits that worktree rather than getting a fresh one, so it sees uncommitted work.
  • One run is capped at 10 minutes by default. The timeoutMs argument of run_qa_test can only raise that cap: a shorter value is ignored, because a QA agent needs the default just to boot its CLI, connect the browser and read the scenario. A run that hits the cap comes back with the status timeout and a summary that says whether the QA agent had started (it started but never submitted a verdict) or never finished starting (no verdict was ever possible).

Common questions

  • Why not let the developer agent test itself? Cost (a QA pass on a small model is roughly ten times cheaper than on the large one), context (the developer stays clean), reliability (a single-purpose tester clicks better than a multitasking developer).
  • Does this need a team graph? No. It is the Dev to QA flavour of a team exposed as a single tool call. For multi-step pipelines with conditions and loops, use Agent Teams.
  • Can I stop the approval prompt? Yes, turn off "ask before test tools" globally, per project, or exempt one agent.
  • The developer agent only gets "timed out" and never a verdict. Read the summary of the timeout: "before the QA Bot finished starting" means the QA agent never came up (check that its CLI is installed and signed in), "started but never submitted a verdict" means the scenario, the page or the browser kept it busy past the cap. Asking for a cap below 10 minutes used to end every run this way; those values are now ignored.