Files
openclaw/docs/providers/azure-speech.md
Peter Steinberger edecdbd05e refactor(config): config-surface reduction tranche 3 — product consolidations (review request) (#111527)
* refactor(config): consolidate media model lists

* refactor(config): unify memory configuration

* refactor(config): consolidate TTS ownership

* refactor(config): move typing policy to agents

* refactor(config): retire product-level config surfaces

* refactor(config): share scoped tool policy type

* chore(config): refresh generated baselines

* fix(config): honor agent typing overrides

* fix(config): migrate sibling config consumers

* refactor(infra): keep base64url decoder private

* fix(config): strip invalid legacy TTS values

* chore(config): refresh rebased baseline hash

* fix(doctor): route legacy messages.tts.realtime voice to talk during tts move

* refactor(config): polish final layout names

* refactor(config): freeze retired tuning defaults

* feat(config): add fast mode default symmetry

* refactor(config): key agent entries by id

* docs(config): update final layout reference

* test(config): cover final layout migrations

* chore(config): refresh final layout baselines

* fix(config): align final layout runtime readers

* fix(config): align remaining readers

* fix(config): stabilize final layout migrations

* fix(config): finalize config projection proof

* fix(config): address final layout review

* docs(release): preserve historical config names

* fix(config): complete keyed agent migration

* fix(config): close final migration gaps

* fix(config): finish full-branch review

* fix(config): complete runtime secret detection

* fix(config): close final review findings

* fix(config): finish canonical docs and heartbeat migration

* fix(config): integrate latest main after rebase

* refactor(env): isolate test-only controls

* refactor(env): isolate build and development controls

* refactor(env): collapse process identity indirection

* refactor(env): remove duplicate config and temp aliases

* docs(env): define the operator-facing allowlist

* ci(env): ratchet production variable count

* fix(env): remove stale provider helper import

* fix(env): make ratchet sorting explicit

* test(env): keep test seam in dead-code audit

* test(env): cover ratchet growth and boundary; document surface budgets

* docs(config): document tier-eval consolidations

* docs(config): clarify speech preference ownership

* test(memory): align retired tuning fixtures

* refactor(memory): freeze engine heuristics

* refactor(config): apply tier-eval tranche

* refactor(tts): move persona shaping to providers

* refactor(compaction): move prompt policy to providers

* test(config): align hookified prompt fixtures

* chore(deadcode): classify test-only exports

* chore(github): remove unused spawn helper

* chore(deadcode): classify queue diagnostics

* chore(deadcode): remove unused lane snapshot export

* chore(plugin-sdk): ratchet consolidated surface

* fix(config): integrate latest main after rebase
2026-07-21 20:28:43 -07:00

6.2 KiB

summary, read_when, title
summary read_when title
Azure AI Speech text-to-speech for OpenClaw replies
You want Azure Speech synthesis for outbound replies
You need native Ogg Opus voice-note output from Azure Speech
Azure Speech

Azure Speech is a bundled Azure AI Speech text-to-speech provider. OpenClaw calls the Azure Speech REST API directly with SSML, synthesizing MP3 for standard replies, native Ogg/Opus for voice notes, and 8 kHz mulaw for telephony channels such as Voice Call. The request sends the provider-owned output format through the X-Microsoft-OutputFormat header.

Detail Value
Provider ID azure-speech (alias: azure)
Website Azure AI Speech
Docs Speech REST text-to-speech
Auth AZURE_SPEECH_KEY plus AZURE_SPEECH_REGION
Default voice en-US-JennyNeural
Default file output audio-24khz-48kbitrate-mono-mp3
Default voice-note file ogg-24khz-16bit-mono-opus

Getting started

In the Azure portal, create a Speech resource. Copy **KEY 1** from Resource Management > Keys and Endpoint, and copy the resource location such as `eastus`.
```
AZURE_SPEECH_KEY=<speech-resource-key>
AZURE_SPEECH_REGION=eastus
```
```json5 { tts: { auto: "always", provider: "azure-speech", providers: { "azure-speech": { voice: "en-US-JennyNeural", lang: "en-US", }, }, }, } ``` Send a reply through any connected channel. OpenClaw synthesizes the audio with Azure Speech and delivers MP3 for standard audio, or Ogg/Opus when the channel expects a voice note.

Configuration options

All options live under tts.providers["azure-speech"].

Option Description
apiKey Azure Speech resource key. Falls back to AZURE_SPEECH_KEY, AZURE_SPEECH_API_KEY, or SPEECH_KEY.
region Azure Speech resource region. Falls back to AZURE_SPEECH_REGION or SPEECH_REGION.
endpoint Optional Azure Speech endpoint override. Falls back to trusted AZURE_SPEECH_ENDPOINT.
baseUrl Optional Azure Speech base URL override.
voice Azure voice ShortName (default en-US-JennyNeural). Legacy alias: voiceId.
lang SSML language code (default en-US).
outputFormat Audio-file output format (default audio-24khz-48kbitrate-mono-mp3).
voiceNoteOutputFormat Voice-note output format (default ogg-24khz-16bit-mono-opus).
timeoutMs Request timeout override in milliseconds. Falls back to the global tts.timeoutMs.

The provider is considered configured once apiKey is set plus one of region, endpoint, or baseUrl. Env vars are only checked as a fallback for config keys left unset. Workspace .env files cannot set AZURE_SPEECH_ENDPOINT; use the process environment, global runtime dotenv, or explicit config for endpoint routing.

Notes

Azure Speech uses a Speech resource key, not an Azure OpenAI key. The key is sent as `Ocp-Apim-Subscription-Key`; OpenClaw derives `https://.tts.speech.microsoft.com` from `region` unless you provide `endpoint` or `baseUrl`. Use the Azure Speech voice `ShortName` value, for example `en-US-JennyNeural`. The bundled provider can list voices through the same Speech resource and filters out voices marked deprecated, retired, or disabled. Azure accepts output formats such as `audio-24khz-48kbitrate-mono-mp3`, `ogg-24khz-16bit-mono-opus`, and `riff-24khz-16bit-mono-pcm`. OpenClaw requests Ogg/Opus for `voice-note` targets so channels can send native voice bubbles without an extra MP3 conversion, and forces `raw-8khz-8bit-mono-mulaw` for telephony targets. `azure` is accepted as a provider alias for existing config, but new config should use `azure-speech` to avoid confusion with Azure OpenAI model providers. TTS overview, providers, and `tts` config. Full config reference including `tts` settings. All bundled OpenClaw providers. Common issues and debugging steps.