mirror of
https://github.com/openclaw/openclaw.git
synced 2026-08-03 02:41:36 +00:00
* refactor(config): consolidate media model lists * refactor(config): unify memory configuration * refactor(config): consolidate TTS ownership * refactor(config): move typing policy to agents * refactor(config): retire product-level config surfaces * refactor(config): share scoped tool policy type * chore(config): refresh generated baselines * fix(config): honor agent typing overrides * fix(config): migrate sibling config consumers * refactor(infra): keep base64url decoder private * fix(config): strip invalid legacy TTS values * chore(config): refresh rebased baseline hash * fix(doctor): route legacy messages.tts.realtime voice to talk during tts move * refactor(config): polish final layout names * refactor(config): freeze retired tuning defaults * feat(config): add fast mode default symmetry * refactor(config): key agent entries by id * docs(config): update final layout reference * test(config): cover final layout migrations * chore(config): refresh final layout baselines * fix(config): align final layout runtime readers * fix(config): align remaining readers * fix(config): stabilize final layout migrations * fix(config): finalize config projection proof * fix(config): address final layout review * docs(release): preserve historical config names * fix(config): complete keyed agent migration * fix(config): close final migration gaps * fix(config): finish full-branch review * fix(config): complete runtime secret detection * fix(config): close final review findings * fix(config): finish canonical docs and heartbeat migration * fix(config): integrate latest main after rebase * refactor(env): isolate test-only controls * refactor(env): isolate build and development controls * refactor(env): collapse process identity indirection * refactor(env): remove duplicate config and temp aliases * docs(env): define the operator-facing allowlist * ci(env): ratchet production variable count * fix(env): remove stale provider helper import * fix(env): make ratchet sorting explicit * test(env): keep test seam in dead-code audit * test(env): cover ratchet growth and boundary; document surface budgets * docs(config): document tier-eval consolidations * docs(config): clarify speech preference ownership * test(memory): align retired tuning fixtures * refactor(memory): freeze engine heuristics * refactor(config): apply tier-eval tranche * refactor(tts): move persona shaping to providers * refactor(compaction): move prompt policy to providers * test(config): align hookified prompt fixtures * chore(deadcode): classify test-only exports * chore(github): remove unused spawn helper * chore(deadcode): classify queue diagnostics * chore(deadcode): remove unused lane snapshot export * chore(plugin-sdk): ratchet consolidated surface * fix(config): integrate latest main after rebase
6.2 KiB
6.2 KiB
summary, read_when, title
| summary | read_when | title | ||
|---|---|---|---|---|
| Azure AI Speech text-to-speech for OpenClaw replies |
|
Azure Speech |
Azure Speech is a bundled Azure AI Speech text-to-speech provider. OpenClaw
calls the Azure Speech REST API directly with SSML, synthesizing MP3 for
standard replies, native Ogg/Opus for voice notes, and 8 kHz mulaw for
telephony channels such as Voice Call. The request sends the provider-owned
output format through the X-Microsoft-OutputFormat header.
| Detail | Value |
|---|---|
| Provider ID | azure-speech (alias: azure) |
| Website | Azure AI Speech |
| Docs | Speech REST text-to-speech |
| Auth | AZURE_SPEECH_KEY plus AZURE_SPEECH_REGION |
| Default voice | en-US-JennyNeural |
| Default file output | audio-24khz-48kbitrate-mono-mp3 |
| Default voice-note file | ogg-24khz-16bit-mono-opus |
Getting started
In the Azure portal, create a Speech resource. Copy **KEY 1** from Resource Management > Keys and Endpoint, and copy the resource location such as `eastus`.```
AZURE_SPEECH_KEY=<speech-resource-key>
AZURE_SPEECH_REGION=eastus
```
```json5
{
tts: {
auto: "always",
provider: "azure-speech",
providers: {
"azure-speech": {
voice: "en-US-JennyNeural",
lang: "en-US",
},
},
},
}
```
Send a reply through any connected channel. OpenClaw synthesizes the audio
with Azure Speech and delivers MP3 for standard audio, or Ogg/Opus when
the channel expects a voice note.
Configuration options
All options live under tts.providers["azure-speech"].
| Option | Description |
|---|---|
apiKey |
Azure Speech resource key. Falls back to AZURE_SPEECH_KEY, AZURE_SPEECH_API_KEY, or SPEECH_KEY. |
region |
Azure Speech resource region. Falls back to AZURE_SPEECH_REGION or SPEECH_REGION. |
endpoint |
Optional Azure Speech endpoint override. Falls back to trusted AZURE_SPEECH_ENDPOINT. |
baseUrl |
Optional Azure Speech base URL override. |
voice |
Azure voice ShortName (default en-US-JennyNeural). Legacy alias: voiceId. |
lang |
SSML language code (default en-US). |
outputFormat |
Audio-file output format (default audio-24khz-48kbitrate-mono-mp3). |
voiceNoteOutputFormat |
Voice-note output format (default ogg-24khz-16bit-mono-opus). |
timeoutMs |
Request timeout override in milliseconds. Falls back to the global tts.timeoutMs. |
The provider is considered configured once apiKey is set plus one of
region, endpoint, or baseUrl. Env vars are only checked as a fallback
for config keys left unset. Workspace .env files cannot set
AZURE_SPEECH_ENDPOINT; use the process environment, global runtime dotenv,
or explicit config for endpoint routing.