Skip to navigation

Speed, Pause, Pronunciation

Adjust speaking speed, insert pauses, and override pronunciation for specific words with Flux TTS.

Pronunciation control on Flux TTS is in Early Access; see Pronunciation control. Some control combinations are invalid on Flux TTS: pronunciation cannot be combined with speed or pause. See Combining controls.

Flux TTS (/v2/speak) voice controls let you adjust speech output: change the speaking rate, insert silence, and override the pronunciation of specific words. They are designed for use cases that need precise delivery of industry terminology, brand names, and complex content.

Aura-2 (/v1/speak) supports speed and pronunciation with the same syntax. Each section calls out where Aura-2 differs.

Availability

ControlBatch (REST)Streaming (WebSocket)Aura-2
SpeedYesYes, including mid-stream ConfigureYes
PronunciationEarly AccessEarly AccessYes
PauseYesNoNot supported

Flux TTS controls are available in English. Aura-2 speed and pronunciation are available in English and Spanish.

Speed control

Adjust the speaking rate of generated audio. Speed control modifies the pace of speech while maintaining natural prosody and voice quality.

Parameters

ParameterLocationTypeDefaultRange
speedqueryfloat1.00.5 - 1.5, in 0.05 increments

On a streaming session, set speed as a connection parameter or change it mid-session with a Configure message. The change applies at the next segment boundary.

Example request

curl --request POST \
--header "Content-Type: application/json" \
--header "Authorization: Token DEEPGRAM_API_KEY" \
--output your_output_file.mp3 \
--data '{"text":"Hello, how can I help you today?"}' \
--url "https://api.deepgram.com/v2/speak?model=flux-haley-en&speed=0.9"

Speed values

ValueEffectUse Case
0.550% slowerSlowest supported rate
0.730% slowerLanguage learning, accessibility, legal compliance
0.820% slowerComplex instructions, elderly users
0.910% slowerClear explanations, training content
1.0Normal speedDefault conversational pace
1.110% fasterEfficient notifications
1.220% fasterQuick alerts, time-sensitive content
1.550% fasterRapid playback, content preview

A value outside the range, or inside it but off the 0.05 increment, returns an error.

Aura-2: speed runs 0.7 - 1.5, so 0.5 and 0.6 are not supported. For Aura-2 Spanish voices, the recommended range is 0.9 - 1.5; values below 0.9 may introduce disfluencies.

Pause control

Insert silence at a specific point in the text. Pause control is available on Flux TTS batch requests.

Aura-2 does not support pause control. On Flux TTS streaming, a pause marker fails the connection with DATA-0002; use batch for text with pauses.

Syntax

Place an escaped pause marker where you want the silence:

Your confirmation number is 4 7 2. \{pause:1s\} Is there anything else I can help with?

Write the duration in milliseconds (\{pause:500\} or \{pause:500ms\}) or seconds (\{pause:1.5s\}). A number with no unit is read as milliseconds. A structured form, {pause:{duration_ms:500}}, is also accepted and is easier for LLMs to generate. Write the structured form without backslashes; an escaped structured marker, or a simple marker without backslashes, is rejected with BREAK_SYNTAX_INVALID.

Example request

curl --request POST \
--header "Content-Type: application/json" \
--header "Authorization: Token DEEPGRAM_API_KEY" \
--output your_output_file.mp3 \
--data '{"text":"Your confirmation number is 4 7 2. \\{pause:1s\\} Is there anything else I can help with?"}' \
--url "https://api.deepgram.com/v2/speak?model=flux-haley-en"

Validation rules

RuleLimit
Duration range500 ms to 3000 ms
Increment100 ms. Off-grid values are rejected, never rounded.
Max pause markers per request8
Adjacent pausesTwo pauses need text between them
Delivery tolerancePauses land within about ±100 ms of the requested duration

Pronunciation control

Override the default pronunciation of specific words using International Phonetic Alphabet (IPA) notation.

Pronunciation control on Flux TTS is in Early Access. The same override can come out differently from one generation to the next; some generations may not follow the IPA. Generate each term several times, on each voice you use, before you rely on it in production.

Syntax

Pronunciation overrides are specified inline within the text using escaped JSON objects:

\{"word": "dupilumab", "pronounce": "duːˈpɪljuːmæb"\}

Where:

  • word is the original text (used for billing and display)
  • pronounce is the IPA phonetic transcription
  • Curly braces must be escaped with backslashes (\{ and \})

Writing IPA for Flux TTS

Flux TTS was trained on a specific IPA style. Overrides that follow it are applied more reliably:

  • Use broad (phonemic) transcription, for example kˈɑpiɹˌaɪt (copyright). Leave out narrow phonetic detail.
  • Use American pronunciations. The model was trained mostly on American English.
  • Always mark primary stress with ˈ. Nearly every word the model was trained on carries one. You can place it before the syllable (ˈkɑpi) or directly before the vowel (kˈɑpi).
  • Secondary stress is optional. Add ˌ when you want more control over a long word.
  • Use ɹ for the English r sound, not r.
  • Add length markers where the model slips. If a vowel comes out short or swapped, mark it long: -miːn rather than -min.

Example request

curl -X POST "https://api.deepgram.com/v2/speak?model=flux-haley-en" \
-H "Authorization: token DEEPGRAM_API_KEY" \
-H "Content-Type: application/json" \
--output your_output_file.mp3 \
-d '{"text": "Take \\{\"word\": \"Azathioprine\", \"pronounce\": \"æzəˈθaɪəpɹiːn\"\\} twice daily with \\{\"word\": \"dupilumab\", \"pronounce\": \"duːˈpɪljuːmæb\"\\}."}'

The curly braces must be escaped with \\{ and \\} in the cURL command.

Common use cases

CategoryWordIPASpoken As
Medicaldupilumabduːˈpɪljuːmæb”doo-PIL-yoo-mab”
Medicalazathioprineæzəˈθaɪəpɹiːn”az-uh-THIGH-oh-preen”
BrandHermèsɛəɹˈmɛz”air-MEZ”
Personal nameNguyenˈwɪn”win”
TechnicalSQLˈsiːkwəl”sequel”

Sourcing IPA transcriptions

A few rules of thumb for producing IPA for your own vocabulary:

Best practices:

  • Always validate by ear. IPA that looks correct on the page can still sound off when synthesized — listen to the output before shipping.
  • Match the dialect. UK and US pronunciations differ (e.g., schedule, aluminum). Make sure the IPA you choose matches the voice and audience you’re targeting.

Validation rules

RuleLimit
Max pronunciations per request500
Max IPA string length128 characters
IPA length ratioCannot exceed 10x the source word length (min floor = 15)

Aura-2 pronunciation control is generally available, with the same syntax and limits and a maximum input text length of 2000 characters. On Aura-2, place the stress mark directly before the vowel (duːpˈɪljuːmæb); a stress mark before a consonant returns a pronunciation warning.

Combining controls

CombinationFlux TTS
Speed + pauseAllowed, with speed capped at 1.15 when any pause marker is present. Speed on its own keeps the full 0.5 - 1.5 range.
Speed + pronunciationRejected with CONTROL_COMBINATION_INVALID (a speed of 1.0 is exempt)
Pause + pronunciationRejected with CONTROL_COMBINATION_INVALID
All threeRejected with CONTROL_COMBINATION_INVALID

A speed of exactly 1.0 does not count as a speed control, so it never triggers these rules. On a streaming session, a pronunciation control sent on a connection opened with a speed other than 1.0, or after a Configure that set one, fails the connection with DATA-0002. A mid-stream Configure that sets speed while a queued turn still carries a pronunciation control is refused with a ConfigureFailure (CONTROL_COMBINATION_INVALID), and the previous speed stays in force. Pronunciations in the turn that is already playing do not block it. See ConfigureFailure codes.

Aura-2 allows speed and pronunciation in the same request. Aura-2 has no pause control, so no other combinations apply.

curl -X POST "https://api.deepgram.com/v1/speak?model=aura-2-thalia-en&speed=0.8" \
-H "Authorization: token DEEPGRAM_API_KEY" \
-H "Content-Type: application/json" \
--output medical_instructions.mp3 \
-d '{"text": "Take \\{\"word\": \"Azathioprine\", \"pronounce\": \"æzəˈθaɪəpɹiːn\"\\} twice daily."}'

Healthcare example

This example uses pronunciation control, which is in Early Access on Flux TTS.

curl -X POST "https://api.deepgram.com/v2/speak?model=flux-haley-en" \
-H "Authorization: token DEEPGRAM_API_KEY" \
-H "Content-Type: application/json" \
--output medical_instructions.mp3 \
-d '{"text": "Take \\{\"word\": \"Azathioprine\", \"pronounce\": \"æzəˈθaɪəpɹiːn\"\\} twice daily with \\{\"word\": \"dupilumab\", \"pronounce\": \"duːˈpɪljuːmæb\"\\}."}'

Use raw string (r'...') with escaped braces \{ and \} for pronunciation control in Python.

Appointment reminder example

Speed and pause can be combined as long as speed stays at or below 1.15.

curl -X POST "https://api.deepgram.com/v2/speak?model=flux-haley-en&speed=0.9" \
-H "Authorization: token DEEPGRAM_API_KEY" \
-H "Content-Type: application/json" \
--output appointment_reminder.mp3 \
-d '{"text": "Your appointment is on Tuesday at 3 PM. \\{pause:800ms\\} Reply YES to confirm."}'

IPA reference

Vowels (American English)

SymbolExampleAs in
iː/biːt/beat
ɪ/bɪt/bit
eɪ/beɪt/bait
ɛ/bɛt/bet
æ/bæt/bat
ɑː/fɑːðɚ/father
ɔː/kɔːt/caught
oʊ/boʊt/boat
ʊ/pʊt/put
uː/buːt/boot
ʌ/kʌt/cut
ə/əˈbaʊt/about

Consonants

SymbolExampleAs in
p/pɪn/pin
b/bɪn/bin
t/tɪn/tin
d/dɪn/din
k/kæt/cat
ɡ/ɡɛt/get
f/fɪn/fin
v/væn/van
θ/θɪŋk/think
ð/ðæt/that
s/sɪt/sit
z/zɪp/zip
ʃ/ʃɪp/ship
ʒ/ˈvɪʒən/vision
h/hæt/hat
tʃ/tʃɪp/chip
dʒ/dʒʌmp/jump
m/mæn/man
n/nɛt/net
ŋ/sɪŋ/sing
l/lɛt/let
ɹ/ɹɛd/red
w/wɪn/win
j/jɛs/yes

Stress markers

SymbolMeaningExample
ˈPrimary stress/ˈæp.əl/ (apple)
ˌSecondary stress/ˌɪnfɚˈmeɪʃən/ (information)

Billing

ControlBilling behavior
SpeedNot billed - adjusting rate doesn’t affect billing
PauseNot billed - pause markers are removed before billing
PronunciationBilled by underlying word - IPA input is not billed

Example: Hello, \{"word": "Mr.", "pronounce": "ˈmɪstɚ"\} Bond. is billed as Hello, Mr. Bond. (16 characters)

Reporting applied controls

Batch response headers

Batch requests report applied controls in the response headers. Flux TTS (/v2/speak) and Aura-2 (/v1/speak) return the same headers.

HTTP/1.1 200 OK
content-type: audio/mpeg
dg-request-id: 3f2a9c1e-8b4d-4e2a-9f1c-7d6e5b4a3c21
dg-model-name: flux-haley-en
dg-char-count: 47
dg-pronunciations-applied: 2
dg-breaks-applied: 0
HeaderDescription
dg-pronunciations-appliedNumber of pronunciation overrides applied
dg-breaks-appliedNumber of pause markers applied. Always 0 on Aura-2, which does not support pause.
dg-warningsComma-separated PRON-NNN codes for pronunciation overrides that triggered an IPA warning. Present only when there are warnings. See Pronunciation warning codes.

The response does not echo the speed value; it is the value you sent in the request.

Streaming

On a streaming session, each turn’s SpeechMetadata reports controls_applied: pronunciations_applied, breaks_applied, and pronunciation_warnings.

A pronunciation override that triggers an IPA warning is still applied best-effort and counted in pronunciations_applied; the warning is reported separately. Listen to any term that produced a warning before you ship it.

Pronunciation warning codes

Batch requests list these codes in dg-warnings. On a streaming session, the same conditions produce a PRONUNCIATION_WARNINGS Warning.

CodeCondition
PRON-001Invalid IPA character
PRON-002Modifier (such as ː or ʰ) follows an invalid character
PRON-003IPA string starts with a modifier (such as ː)
PRON-004Tie bar (͡) follows a non-base character
PRON-005Tie bar precedes a non-base character
PRON-006IPA string ends with a tie bar
PRON-007IPA string starts with a tie bar
PRON-008Stress mark precedes a non-vowel (Aura-2 only; Flux TTS accepts it)
PRON-009IPA string ends with a stress mark

Error handling

Batch requests return a 400 with one of these err_code values:

err_codeTrigger
CONTROL_COMBINATION_INVALIDPronunciation combined with speed, pause, or both
PAUSE_SPEED_CAP_EXCEEDEDA pause marker with speed above 1.15
BREAK_OUT_OF_RANGEA pause shorter than 500 ms or longer than 3000 ms
BREAK_INCREMENT_INVALIDA pause duration off the 100 ms grid
BREAKS_LIMIT_EXCEEDEDMore than 8 pause markers, or two pauses with no text between them
BREAK_SYNTAX_INVALIDA malformed pause marker, such as {pause:800ms} without backslashes or an escaped structured marker

Invalid IPA and invalid speed values are also rejected. On a streaming session, see the warning, ConfigureFailure, and error codes.

Aura-2 returns these errors for speed and pronunciation:

{"err_code": "speed_out_of_range", "err_msg": "Speed must be between 0.7 and 1.5"}
{"err_code": "pronunciation_invalid", "err_msg": "Invalid IPA notation for 'azathioprine'"}

Limits

LimitFlux TTSAura-2
Speed range0.5 - 1.5 (0.05 increments; max 1.15 with a pause)0.7 - 1.5
Max pause markers per request8Not supported
Pause duration500 - 3000 ms (100 ms increments)Not supported
Max pronunciations per request500500
Max IPA string length128 characters128 characters
Max input text lengthSee Flux TTS batch2000 characters