Flux Text to Speech (batch)
Synthesize a complete block of text into a single audio response using Deepgram’s Flux TTS batch (REST) API. Use this for pre-rendering fixed audio (IVR prompts, notifications, narration) where the whole text is known up front and you don’t need incremental playback or interruption.
Authentication
Use Authorization: Token <API_KEY>
Example: Authorization: Token 12345abcdef
Use Authorization: Bearer <JWT>
Example: Authorization: Bearer eyJhbGciOiJ...
Query parameters
Opts out requests from the Deepgram Model Improvement Program. Refer to our Docs for pricing impacts before setting this to true. https://dpgr.am/deepgram-mip
Expressive range of the generated speech, on a calm-to-animated axis. Accepted values: -2, -1, 0, 1, 2. 0 (the default) is the voice's tuned delivery and the production-validated setting, with -2 the calm end of the range and 2 the animated end. Supported on all Flux voices; applies to the whole request. Beta: behavior may change in future model versions, and non-default values increase the risk of hallucinations and pronunciation errors; audition before shipping. An invalid value is rejected with a 400 — EXPRESSIVITY_OUT_OF_RANGE for a value outside the range, EXPRESSIVITY_INCREMENT_INVALID for a fractional value. See Expressivity.
Flux TTS model used to synthesize the submitted text, in the form flux-{voice}-{language} (for example, flux-alexis-en). Required; unlike the v1 (Aura) endpoint there is no default and only flux models are accepted. English-only at launch.
Speaking rate multiplier that adjusts the pace of generated speech while preserving natural prosody and voice quality. Accepted values run 0.5 to 1.5 in 0.05 increments. Not yet supported in all languages. When the text contains an inline pause marker, speed is capped at 1.15 (PAUSE_SPEED_CAP_EXCEEDED above that). A value other than 1.0 cannot be combined with inline pronunciation controls (CONTROL_COMBINATION_INVALID).
Processing priority for asynchronous (callback) requests. The only supported value is low.
Request
The text content to be converted to speech. The server normalizes and preprocesses the text before synthesis. May contain inline pause controls (\{pause:500ms\}, 500-3000 ms in 100 ms steps, at most 8 per request) and inline pronunciation controls (\{"word": "...", "pronounce": "<IPA>"\}, Early Access). Pronunciation cannot be combined with pause or with a speed other than 1.0, and speed is capped at 1.15 when a pause is present. See Speed, Pause, Pronunciation.
Response
Returns the synthesized audio in the requested encoding as a binary stream. When a callback URL is supplied, the request is processed asynchronously and the response body is instead a JSON acknowledgement (Content-Type application/json) of the form {"request_id": "..."}, with the audio delivered to the callback URL. Because this endpoint is typed as a binary audio stream, SDK callers that set callback receive this JSON acknowledgement through the audio byte iterator as raw bytes and must join the chunks and parse request_id themselves.