Speed, Pause, Pronunciation
Pronunciation control on Flux TTS is in Early Access; see Pronunciation control. Some control combinations are invalid on Flux TTS: pronunciation cannot be combined with speed or pause. See Combining controls.
Flux TTS (/v2/speak) voice controls let you adjust speech output: change the speaking rate, insert silence, and override the pronunciation of specific words. They are designed for use cases that need precise delivery of industry terminology, brand names, and complex content.
Aura-2 (/v1/speak) supports speed and pronunciation with the same syntax. Each section calls out where Aura-2 differs.
Availability
Flux TTS controls are available in English. Aura-2 speed and pronunciation are available in English and Spanish.
Speed control
Adjust the speaking rate of generated audio. Speed control modifies the pace of speech while maintaining natural prosody and voice quality.
Parameters
On a streaming session, set speed as a connection parameter or change it mid-session with a Configure message. The change applies at the next segment boundary.
Example request
Speed values
A value outside the range, or inside it but off the 0.05 increment, returns an error.
Aura-2: speed runs 0.7 - 1.5, so 0.5 and 0.6 are not supported. For Aura-2 Spanish voices, the recommended range is 0.9 - 1.5; values below 0.9 may introduce disfluencies.
Pause control
Insert silence at a specific point in the text. Pause control is available on Flux TTS batch requests.
Aura-2 does not support pause control. On Flux TTS streaming, a pause marker fails the connection with DATA-0002; use batch for text with pauses.
Syntax
Place an escaped pause marker where you want the silence:
Write the duration in milliseconds (\{pause:500\} or \{pause:500ms\}) or seconds (\{pause:1.5s\}). A number with no unit is read as milliseconds. A structured form, {pause:{duration_ms:500}}, is also accepted and is easier for LLMs to generate. Write the structured form without backslashes; an escaped structured marker, or a simple marker without backslashes, is rejected with BREAK_SYNTAX_INVALID.
Example request
Validation rules
Pronunciation control
Override the default pronunciation of specific words using International Phonetic Alphabet (IPA) notation.
Pronunciation control on Flux TTS is in Early Access. The same override can come out differently from one generation to the next; some generations may not follow the IPA. Generate each term several times, on each voice you use, before you rely on it in production.
Syntax
Pronunciation overrides are specified inline within the text using escaped JSON objects:
Where:
wordis the original text (used for billing and display)pronounceis the IPA phonetic transcription- Curly braces must be escaped with backslashes (
\{and\})
Writing IPA for Flux TTS
Flux TTS was trained on a specific IPA style. Overrides that follow it are applied more reliably:
- Use broad (phonemic) transcription, for example
kˈɑpiɹˌaɪt(copyright). Leave out narrow phonetic detail. - Use American pronunciations. The model was trained mostly on American English.
- Always mark primary stress with
ˈ. Nearly every word the model was trained on carries one. You can place it before the syllable (ˈkɑpi) or directly before the vowel (kˈɑpi). - Secondary stress is optional. Add
ˌwhen you want more control over a long word. - Use
ɹfor the English r sound, notr. - Add length markers where the model slips. If a vowel comes out short or swapped, mark it long:
-miːnrather than-min.
Example request
The curly braces must be escaped with \\{ and \\} in the cURL command.
Common use cases
Sourcing IPA transcriptions
A few rules of thumb for producing IPA for your own vocabulary:
- Short lists (<20 words): generate with an LLM and validate by ear.
- Longer lists: use authoritative dictionaries that publish IPA directly:
Best practices:
- Always validate by ear. IPA that looks correct on the page can still sound off when synthesized — listen to the output before shipping.
- Match the dialect. UK and US pronunciations differ (e.g., schedule, aluminum). Make sure the IPA you choose matches the voice and audience you’re targeting.
Validation rules
Aura-2 pronunciation control is generally available, with the same syntax and limits and a maximum input text length of 2000 characters. On Aura-2, place the stress mark directly before the vowel (duːpˈɪljuːmæb); a stress mark before a consonant returns a pronunciation warning.
Combining controls
A speed of exactly 1.0 does not count as a speed control, so it never triggers these rules. On a streaming session, a pronunciation control sent on a connection opened with a speed other than 1.0, or after a Configure that set one, fails the connection with DATA-0002. A mid-stream Configure that sets speed while a queued turn still carries a pronunciation control is refused with a ConfigureFailure (CONTROL_COMBINATION_INVALID), and the previous speed stays in force. Pronunciations in the turn that is already playing do not block it. See ConfigureFailure codes.
Aura-2 allows speed and pronunciation in the same request. Aura-2 has no pause control, so no other combinations apply.
Healthcare example
This example uses pronunciation control, which is in Early Access on Flux TTS.
Use raw string (r'...') with escaped braces \{ and \} for pronunciation control in Python.
Appointment reminder example
Speed and pause can be combined as long as speed stays at or below 1.15.
IPA reference
Vowels (American English)
Consonants
Stress markers
Billing
Example: Hello, \{"word": "Mr.", "pronounce": "ˈmɪstɚ"\} Bond. is billed as Hello, Mr. Bond. (16 characters)
Reporting applied controls
Batch response headers
Batch requests report applied controls in the response headers. Flux TTS (/v2/speak) and Aura-2 (/v1/speak) return the same headers.
The response does not echo the speed value; it is the value you sent in the request.
Streaming
On a streaming session, each turn’s SpeechMetadata reports controls_applied: pronunciations_applied, breaks_applied, and pronunciation_warnings.
A pronunciation override that triggers an IPA warning is still applied best-effort and counted in pronunciations_applied; the warning is reported separately. Listen to any term that produced a warning before you ship it.
Pronunciation warning codes
Batch requests list these codes in dg-warnings. On a streaming session, the same conditions produce a PRONUNCIATION_WARNINGS Warning.
Error handling
Batch requests return a 400 with one of these err_code values:
Invalid IPA and invalid speed values are also rejected. On a streaming session, see the warning, ConfigureFailure, and error codes.
Aura-2 returns these errors for speed and pronunciation: