Skip to navigation

Turn-based Speech (Flux)

Streaming, turn-based text-to-speech (Flux TTS) built for voice-agent pipelines. Stream LLM tokens in, speak them to the user, and report per-turn billing and timing.

HandshakeTry it

WSS
wss://api.deepgram.com/v2/speak

Authentication

AuthorizationToken

Use Authorization: Token <API_KEY> Example: Authorization: Token 12345abcdef

OR
AuthorizationBearer

Use Authorization: Bearer <JWT> Example: Authorization: Bearer eyJhbGciOiJ...

Headers

AuthorizationstringRequired

Use your API key or a temporary token for authentication via the Authorization header. In client-side environments where custom headers are not supported, use the Sec-WebSocket-Protocol header instead.

Example: Authorization: Token %DEEPGRAM_API_KEY% or Authorization: Bearer %DEEPGRAM_TOKEN%

Query parameters

modelstringRequired

The Flux TTS model used to synthesize speech. Required on every connection. Model strings follow the format flux-{voice}-{language} (e.g. flux-alexis-en). An Aura model string is rejected on /v2/speak; use /v1/speak for Aura voices.

encodingenumOptionalDefaults to linear16

Encoding of the raw output audio. The streaming WebSocket emits raw (non-containerized) audio, so only streaming-compatible encodings are supported. Compressed and containerized encodings (mp3, opus, flac, aac) are available on the batch REST transport only.

Allowed values:
sample_rateenumOptional

Output sample rate in Hz. With linear16, valid values are 8000, 16000, 24000, 32000, 44100, and 48000. With mulaw or alaw, valid values are 8000 and 16000. Defaults to the model's native sample rate.

speeddoubleOptional0.5-1.5Defaults to 1

Speech-rate multiplier. 1.0 is the model's nominal rate; lower is slower. Accepted values run 0.5 to 1.5 in 0.05 increments. A value outside that range is rejected with SPEED_OUT_OF_RANGE; a value inside it but off the 0.05 increment with SPEED_INCREMENT_INVALID. Models and languages without runtime speed control reject any value with SPEED_NOT_SUPPORTED. A speed other than 1.0 cannot be combined with inline pronunciation controls; see Speed, Pause, Pronunciation.

expressivityenumOptionalDefaults to 0

Expressive range of the generated speech, on a calm-to-animated axis. Accepted values: -2, -1, 0, 1, 2. 0 (the default) is the voice's tuned delivery and the production-validated setting, with -2 the calm end of the range and 2 the animated end. Supported on all Flux voices. Fixed for the connection — not settable via Configure. Beta: behavior may change in future model versions, and non-default values increase the risk of hallucinations and pronunciation errors; audition before shipping. An invalid value fails the connection with a 400 — EXPRESSIVITY_OUT_OF_RANGE for a value outside the range, EXPRESSIVITY_INCREMENT_INVALID for a fractional value. See Expressivity.

Allowed values:
mip_opt_outanyOptionalDefaults to false

Opts out requests from the Deepgram Model Improvement Program. Refer to our Docs for pricing impacts before setting this to true. https://dpgr.am/deepgram-mip

taganyOptional
Label your requests for the purpose of identification during usage reporting. Repeatable.

Send

SpeakV2SpeakobjectRequired
Send text to be synthesized into the active turn
OR
SpeakV2FlushobjectRequired
End the active turn and generate the remaining audio
OR
SpeakV2InterruptobjectRequired
Cancel the active turn because the user barged in
OR
SpeakV2ConfigureobjectRequired
Update synthesis configuration mid-session
OR
SpeakV2CloseobjectRequired
Gracefully close the connection, draining all remaining and queued audio

Receive

SpeakV2AudiostringRequiredformat: "binary"
Receive audio chunks as they are generated
OR
SpeakV2ConnectedobjectRequired
Receive a connected message on a successful connection
OR
SpeakV2SpeechStartedobjectRequired
Receive a message marking the start of a new turn, carrying the turn's unique identifier
OR
SpeakV2SpeechMetadataobjectRequired
Receive per-turn billing and timing after a manual Flush
OR
SpeakV2SpeechInterruptedobjectRequired
Receive what the user heard, and the interrupted turn's billing, after an Interrupt
OR
SpeakV2FlushedobjectRequired
Receive an echo confirming receipt of a manual Flush
OR
SpeakV2SessionMetadataobjectRequired
Receive cumulative session totals before the socket closes
OR
SpeakV2ConfigureSuccessobjectRequired
Receive confirmation that a Configure was accepted and applied, echoing the applied configuration
OR
SpeakV2ConfigureFailureobjectRequired
Receive notice that a Configure was rejected or failed to apply; the prior configuration is retained
OR
SpeakV2WarningobjectRequired
Receive a warning; synthesis continues and the connection is unaffected
OR
SpeakV2ErrorobjectRequired
Receive a fatal error message followed by a WebSocket close