Turn-based Audio (Flux)
Real-time conversational speech recognition with contextual turn detection for natural voice conversations
HandshakeTry it
Authentication
Use Authorization: Token <API_KEY>
Example: Authorization: Token 12345abcdef
Use Authorization: Bearer <JWT>
Example: Authorization: Bearer eyJhbGciOiJ...
Headers
Use your API key or a temporary token for authentication via the Authorization header. In client-side environments where custom headers are not supported, use the Sec-WebSocket-Protocol header instead.
Example: Authorization: Token %DEEPGRAM_API_KEY% or Authorization: Bearer %DEEPGRAM_TOKEN%
Query parameters
Encoding of the audio stream. Required if sending non-containerized/raw audio. If sending containerized audio, this parameter should be omitted.
Sample rate of the audio stream in Hz. Required if sending non-containerized/raw audio. If sending containerized audio, this parameter should be omitted.
End-of-turn confidence required to fire an eager end-of-turn event.
When set, enables EagerEndOfTurn and TurnResumed events. Valid
Values 0.3 - 0.9.
End-of-turn confidence required to finish a turn. Valid Values 0.5 -
1.0. Set to 1.0 to suppress confidence-based end-of-turn detection.
eot_timeout_ms still ends idle turns; increase it when using
ForceEndTurn for full manual turn control.
A turn will be finished when this much time has passed after speech, regardless of EOT confidence. Valid Values 500 - 60000.
Keyterm prompting improves recognition of specialized terminology.
keyterm accepts plain terms only. Unlike the legacy keywords feature,
it does not support weights or intensifiers. Appending one
(for example, keyterm=term:0.15) is not rejected—the weight is
silently ignored and the entire value is treated as a literal keyterm.
To boost multiple separate keyterms, repeat the keyterm parameter
(for example, keyterm=term1&keyterm=term2). To boost one multi-word
phrase as a single keyterm, join the words with %20 or +
(for example, keyterm=customer%20service). Do not separate keyterms
with commas, semicolons, or line breaks.
Language hints constrain and prioritize language detection for the flux-general-multi model. Pass multiple language_hint query parameters to specify multiple language codes. Empty values are rejected. Only valid when model is flux-general-multi.
Profanity Filter looks for recognized profanity and converts it to the nearest recognized non-profane word or removes it from the transcript completely.
Redaction removes sensitive information from your transcripts. On Flux, only numbers and aggressive_numbers are supported.
Opts out requests from the Deepgram Model Improvement Program. Refer to our Docs for pricing impacts before setting this to true. https://dpgr.am/deepgram-mip