Voice Dictation

  • Desktop, mobile and web
  • All plans

What it does

Dictate your prompts instead of typing them. Click the microphone in the composer, speak, click again: the transcript is inserted in your draft at the cursor, ready to reread and send with Enter. Nothing is sent automatically. On desktop the audio is transcribed by the AgentsRoom voice backend (an accurate model for short clips, Whisper for recordings over 90 seconds, chosen automatically); on the phone recognition runs on-device and only the text is relayed, end-to-end encrypted, to the focused desktop agent. A live waveform shows the mic is picking you up, and silence is never transcribed, so a muted mic cannot produce invented text.

Where to find it

  • Desktop: the Dictate button (microphone) in the composer toolbar of every agent. Right-click it for Global dictation settings and Project dictation settings.
  • Settings > Composer: dictation language, sounds, music pause, Offline dictation.
  • Mobile companion: the microphone button on the agent screen (press and hold).

How to use it

  1. Put the cursor where the text should go and click Dictate. The first time, the OS asks for microphone access.
  2. Speak. The button pulses red with a live waveform. Click again to stop.
  3. The transcript is inserted at the cursor, split into one line per sentence. Edit if needed, press Enter to send.

Per-project tuning: Project dictation settings lets you edit the context sent with each dictation (agent and project names are filled in; add domain terms, library or product names) and force a spoken language for this project.

Mobile: hold the mic, speak, release. The recognized text is pasted into the focused desktop agent's prompt, without Enter; you confirm from the phone or the desktop.

Settings

  • sttLanguage (global, Settings > Composer > Dictation language; project override in Project dictation settings): auto by default, or one of 19 language codes. Force one only when short clips get mis-detected.
  • sttPromptContext (project, Project dictation settings): the context template sent with each dictation.
  • dictationSounds (global): start and stop cues.
  • stopMusicWhenRecording (global): mute the system audio output while recording, restored when you stop. On Windows it used to change the speaker channel volume instead of muting; fixed on 2026-09-24.
  • localSttEnabled, localSttMode (fallback or always), localSttModel (tiny or base) (global, Settings > Composer > Offline dictation): on-device Whisper, downloaded once (about 42 or 78 MB), used when credits or network run out, or always.
  • composerHiddenActions (global): hide the mic button.

Agent tools (MCP)

None. Dictation writes into your draft; agents only receive what you send.

Providers

All providers, no difference: the transcript is text in the composer. The transcription model is chosen by AgentsRoom, not by the agent's CLI.

Mobile

Present, with a different engine: the phone's own speech recognition (audio never leaves the device), text relayed encrypted to the focused desktop agent as a paste without Enter. If no desktop terminal is focused, the phone says so and sends nothing. The out-of-credits case falls back to the device engine.

Limits

  • Requires a signed-in account for server transcription. Included time per calendar month: 5 minutes of transcribed audio on Free, 30 minutes on Plus, 3 hours on Pro. These three values are compiled into the product: no back-office can widen them. Beyond the allowance each transcription spends voice credits (shared with Voice Mode and Read Aloud). Empty balance: dictation pauses and offers Keep dictating without credits (offline model) or a top-up.
  • Your own OpenAI key removes the credit spend, see Bring Your Own OpenAI Key.
  • macOS: without microphone permission the recording is silent; the composer shows "Microphone blocked or no audio captured".
  • Windows: when Let desktop apps access your microphone is off in the Windows privacy settings, the recording is refused up front and the same "Microphone blocked" notice appears (since 2026-09-23; before, the recording ran on silence and was dropped without a word).
  • Recordings over 90 seconds use Whisper, which is more robust on long, noisy audio; repeated sentences are collapsed as a safety net.
  • The offline model is less accurate than the online one.

Common questions

  • Does dictation send my prompt? No, it only fills the draft. You press Enter.
  • Can I mix typing and dictation? Yes, the transcript is inserted at the cursor, not in place of the draft.
  • Why does forcing a language give garbage? The model transcribes in the forced language whatever you speak. Keep Auto-detect unless detection fails on short clips.
  • The mic looks broken and nothing is transcribed. Most often the voice credit balance is empty or the OS blocked the microphone. Hover the mic button: its tooltip shows the reason ("Microphone blocked or no audio captured", "No speech detected, nothing was transcribed.") instead of the usual title.
  • The counter ran but no waveform and no text appeared (Windows). Fixed on 2026-09-23: when the audio analysis never started, the recording is now transcribed instead of being discarded as silence.
  • Which model is used? Chosen automatically by duration; there is no model picker any more.