Skip to navigation

Flux TTS Inline Controls: Pause and Pronunciation

Flux TTS (/v2/speak) now supports inline pause and pronunciation controls in the text you send, alongside speed.

  • Pause control. Insert silence with an escaped marker such as \{pause:1s\}: 500–3000 ms in 100 ms steps, up to 8 per request. Available on batch (REST) requests.
  • Pronunciation control (Early Access). Override how a word is said with inline IPA, \{"word": "dupilumab", "pronounce": "duːˈpɪljuːmæb"\}, on both batch and streaming. Results can vary between generations while the feature is in Early Access, so generate each term several times before relying on it in production.
  • Combination guardrails. Combinations that degrade audio are rejected with a clear error instead of producing bad audio: pronunciation cannot be combined with pause or with a speed other than 1.0, and speed is capped at 1.15 when a pause is present. Rejections return CONTROL_COMBINATION_INVALID or PAUSE_SPEED_CAP_EXCEEDED on batch, and ConfigureFailure or DATA-0002 on streaming.
  • Reporting. Applied controls are counted in controls_applied on streaming and in the dg-pronunciations-applied and dg-breaks-applied headers on batch.

See Speed, Pause, Pronunciation for syntax, limits, and error codes.