openclaw/docs/tools/tts.md
Ayaan Zaidi 25fe1737a3
docs: correct ElevenLabs TTS model configuration key (#161917)
The TTS docs told people to set `tts.providers.elevenlabs.model`, but the ElevenLabs provider only reads `modelId`. Because the provider config accepts unknown keys, `model` passed validation and was silently ignored, so synthesis used the default model. The field reference, the index and all three JSON5 examples now document `modelId`, and the old anchor still resolves.

Related to #161914. Thanks @hartra344 for the report. Existing configs that set `model` still need separate handling, such as a Doctor migration or a warning, which is left for a maintainer to decide.

Proof: the real CLI on an isolated profile against a local fake ElevenLabs endpoint. `model: "eleven_v3"` sent `eleven_multilingual_v2`, and `modelId: "eleven_v3"` sent `eleven_v3`. The docs checks and `git diff --check` pass.

Co-authored-by: Ayaan Zaidi <hi@obviy.us>
2026-09-30 22:42:58 +08:00

20 KiB

summary title sidebarTitle read_when
Index of the OpenClaw text-to-speech documentation, one page per reader job Text-to-speech Text to speech (TTS)
Enabling text-to-speech for replies
Configuring a TTS provider, fallback chain, or persona
Using /tts commands or directives

OpenClaw converts outbound replies into native voice messages on Feishu, Matrix, Telegram, and WhatsApp. Every other channel receives an audio attachment. Telephony and Talk receive PCM or Ulaw streams.

TTS is the speech-output half of Talk's stt-tts mode (talk.speak calls this same synthesis path). Provider-native realtime Talk sessions synthesize speech inside the realtime provider instead. transcription sessions never synthesize an assistant voice reply.

This page is an index. Text-to-speech is documented on seven pages, one per reader job. Open the page that matches your task.

Page Read it when
Text-to-speech quickstart You are turning TTS on, choosing a provider, and testing it from chat.
Text-to-speech configuration You need the tts config block, a provider snippet, a local engine, or override precedence.
Text-to-speech personas You want one stable spoken identity, its provider bindings, and its fallback policy.
Commands and directives You need [[tts:...]] directives, the /tts commands, or where local preferences live.
Output and Auto-TTS behavior You need the audio format per channel, transcoding rules, or when Auto-TTS summarizes.
Text-to-speech field reference You need the type, default, env var, or legacy alias for one TTS field.
Agent tool and Gateway RPC You are calling TTS from an agent tool call or a Gateway RPC method.

Where each section moved

Every section heading from the previous single-page version keeps its anchor here, so an existing link such as /tools/tts#per-agent-voice-overrides still resolves. Each entry points at the page that now holds the content.

Component anchors

The previous single-page version also minted an anchor for every step, tab, accordion, and field. Those anchors are preserved here so that any deep link into the old page still resolves. Nine accordion anchors lost a -1 suffix when the tab that shared their slug moved to a different page. The stub below keeps the old id and points at the new one.

Quickstart

Configuration

Field reference