Voice Mode

  • Desktop app
  • All plans

What it does

A two-way spoken conversation with one running agent. You speak, the turn is transcribed and sent, the agent works, and its reply is read back in a natural voice. In hands-free mode the microphone re-arms as soon as the reply ends, so the exchange flows like a phone call. The panel keeps a rolling log of the last turns so you can read what was understood, and the keyboard composer stays available: you choose message by message between voice and typing. The engine runs at the app level, so switching tabs never cuts the mic or the playback; a floating dock follows you when the agent's tab is not in front.

Where to find it

  • Composer toolbar of an agent: the Voice Mode button ("Talk to this agent and hear its replies"). The voice panel stacks above the composer; a floating dock appears elsewhere.
  • Panel gear: Voice settings (hands-free, language, reply voice, pause tolerance, waiting cue, barge-in, speak new answers).
  • Settings > Composer > Voice push-to-talk: the system-wide shortcut, bound only while Voice Mode is open.

How to use it

  1. Open Voice Mode on the agent you want to talk to and click Start talking (or press Space, the headset/media key, or the global shortcut).
  2. Speak. When you pause longer than the pause tolerance, the turn is sent. Clicking the button while listening also means Stop and send.
  3. The agent works (an optional soft beep plays every few seconds), then its answer is spoken. With Hands-free on, the mic reopens right after; off, it waits for you.
  4. Barge-in (on by default): the mic stays armed while the agent works or speaks. Start talking during playback to interrupt and open a new turn; talk while it works to add a correction mid-turn.
  5. Close the panel to release the mic and every shortcut.

Settings

  • voiceModeAutoSend (panel: Hands-free): re-arm the mic after each reply. Default on.
  • voiceModeVoice (panel: Reply voice): text-to-speech voice, with a preview. Shared with Read Aloud.
  • voiceModeLanguage (panel: Language): spoken language, auto by default.
  • voiceModeSilenceMs (panel: Pause tolerance): 1000 ms for fast dictation up to 4000 ms for thinking out loud.
  • voiceModeBargeIn (panel): keep the mic armed while the agent works or speaks. Default on.
  • voiceModeSpeakUnprompted (panel: Speak new answers): read out an answer you did not ask for (a peer message, a turn the agent takes on its own). Off by default.
  • voiceThinkingBeep (panel: Waiting cue): the soft beep while the agent works.
  • voiceModeShortcutEnabled, voiceModeShortcut (global, Settings > Composer > Voice push-to-talk): system-wide key, default Alt+Shift+V, active only while Voice Mode is open.
  • stopMusicWhenReading (global, Settings > Composer): pause the media player while a reply is spoken.

Agent tools (MCP)

None. The agent only sees text: a spoken preamble with the first message asks for short, prose replies, and a light reminder every few messages keeps it there.

Providers

All providers: turns are sent as text and the end of a turn is detected from the agent's status. What gets spoken comes, in order of preference, from the agent's conversation transcript (Claude Code and Codex today), from a dedicated reply field every CLI can write in its session file, or from the agent's last notification as a fallback.

Mobile

Not available on the mobile companion. The phone offers one-way dictation and Read Aloud of the last answer instead.

Limits

  • Needs a signed-in account and a running agent. Not available in the web demo.
  • Billed in voice credits, shared with dictation and Read Aloud: a small starting balance on every account, a monthly top-up for Pro, one-time packs otherwise. The balance is checked before each call; when it is empty the panel offers Top up. Your own OpenAI key removes the spend, see Bring Your Own OpenAI Key.
  • Spoken replies are trimmed of markdown, code, tables and paths and capped to roughly 900 characters.
  • Starting a Read Aloud closes Voice Mode (one voice at a time, and an open mic would hear the synthetic voice).
  • On loudspeakers, an echo of the synthetic voice can occasionally be picked up; a guard drops transcripts that match what was just read.

Common questions

  • How is it different from voice dictation? Dictation is one way (speech into the draft). Voice Mode is a live back-and-forth with spoken replies.
  • Is hands-free a push-to-talk? No. In both positions the end of your sentence is detected and sent on silence; hands-free only decides whether the mic reopens by itself after the reply.
  • Can I talk to the agent while I am in another app? Yes: the media/headset key and the global shortcut work while Voice Mode is open, without bringing the window forward.
  • The agent read a useless one-line status. Only the prose reply of the transcript is spoken once the turn settles; notifications are a last resort. If it happens with a provider without a readable transcript, the reply field fallback applies.
  • Why did the mic stop when I opened a document reader? Read Aloud takes the audio channel; reopen Voice Mode after.