Correction: over-limit Configure keyterm updates on /v1/listen
The October 2 entry on mid-stream Nova-3 keyterm updates said that a Configure message whose keyterms exceed the 500-token limit returns an Error and leaves the stream’s keyterms unchanged. That is not the current behavior. An over-limit update stops transcription without an Error, and after about 30 seconds the server closes the stream with 1011 (NET-0000). Check the list size before sending the update. See Configure and Keyterm Prompting.
Nova-3 Medical: Improved English Models for Batch and Streaming
We have released improved Nova-3 Medical English models for batch and streaming, with fewer word errors and fewer errors on medical entities.
Nova-3 Medical Multilingual: Dutch, English, French, German, Hindi, Italian, Japanese, Portuguese, Russian, and Spanish
The nova-3-medical model now supports multilingual transcription with language=multi. Nova-3 Medical can now transcribe medical audio in which speakers switch between Dutch, English, French, German, Hindi, Italian, Japanese, Portuguese, Russian, and Spanish. Available for both batch and streaming.
Nova-3 Improved Models for Catalan, French, Indonesian, and Vietnamese
We have released improved Nova-3 monolingual models for four existing languages. These updates improve batch transcription for Vietnamese and Catalan and streaming transcription for French and Indonesian.
Update Nova-3 Keyterms Mid-Stream
Swap in the vocabulary for each step of a call without reconnecting. To replace a Nova-3 stream’s keyterms during a /v1/listen session, send a keyterms array in a Configure message:
Each keyterms array replaces the whole list, including keyterms set with the keyterm query parameter, and an empty array [] clears them. The 500-token keyterm limit that applies to the keyterm query parameter also applies to each update; an over-limit update returns an Error and leaves the stream’s keyterms unchanged. The same message can also turn formatting features on or off, for example "features": { "numerals": true }.
October 1, 2026
Deepgram Self-Hosted October 2026 Release (261001)
Container Images (release 261001)
-
quay.io/deepgram/self-hosted-api:release-261001- Equivalent image to:
quay.io/deepgram/self-hosted-api:1.192.51
- Equivalent image to:
-
quay.io/deepgram/self-hosted-engine:release-261001-
Equivalent image to:
quay.io/deepgram/self-hosted-engine:3.135.0
-
Minimum required NVIDIA driver version:
>=580, open kernel module flavor. See Drivers and Containerization Platforms for installation and verification steps.
-
-
quay.io/deepgram/self-hosted-license-proxy:release-261001- Equivalent image to:
quay.io/deepgram/self-hosted-license-proxy:1.11.1-1
- Equivalent image to:
-
quay.io/deepgram/self-hosted-billing:release-261001- Equivalent image to:
quay.io/deepgram/self-hosted-billing:1.14.0-3
- Equivalent image to:
Flux TTS Inline Controls: Pause and Pronunciation
Flux TTS (/v2/speak) now supports inline pause and pronunciation controls in the text you send, alongside speed.
- Pause control. Insert silence with an escaped marker such as
\{pause:1s\}: 500–3000 ms in 100 ms steps, up to 8 per request. Available on batch (REST) requests. - Pronunciation control (Early Access). Override how a word is said with inline IPA,
\{"word": "dupilumab", "pronounce": "duːˈpɪljuːmæb"\}, on both batch and streaming. Results can vary between generations while the feature is in Early Access, so generate each term several times before relying on it in production. - Combination guardrails. Combinations that degrade audio are rejected with a clear error instead of producing bad audio: pronunciation cannot be combined with pause or with a
speedother than1.0, andspeedis capped at1.15when a pause is present. Rejections returnCONTROL_COMBINATION_INVALIDorPAUSE_SPEED_CAP_EXCEEDEDon batch, andConfigureFailureorDATA-0002on streaming. - Reporting. Applied controls are counted in
controls_appliedon streaming and in thedg-pronunciations-appliedanddg-breaks-appliedheaders on batch.
Toggle Numerals Mid-Stream on Flux STT
Switch Flux STT to digits for a PIN, phone number, or order number, then back to words, without reconnecting. To turn Numerals on or off during a Flux STT stream, send numerals as a boolean in a Configure message:
The update applies to transcripts Flux STT sends after it processes the message, and ConfigureSuccess now includes numerals in the full active configuration it echoes. The numerals query parameter still sets the initial value when the stream opens. This replaces the connection-time-only behavior in the July 17 entry.
Nova-3 Improved Models for Flemish, German (Switzerland), Lithuanian, and Portuguese
We have released improved Nova-3 monolingual models for four existing languages. These updates enhance transcription quality for batch and streaming workloads.
Nova-3 Improved Models for Danish, Estonian, Flemish, Italian, Lithuanian, Macedonian, Polish, Urdu, and Vietnamese
We have released improved Nova-3 monolingual models for nine existing languages. These updates enhance transcription quality for batch and streaming workloads.
Correction: Browser Agent SDK has no client-side VAD
The Browser Agent SDK never shipped client-side Silero voice activity detection. Documentation published with the SDK described it as an available feature, and that documentation was wrong. The following did not exist in any released package:
- The
vadoption onMicrophoneOptionsin@deepgram/agents. - The
speechThresholdandsilenceThresholdVAD options. - The
speech-startandspeech-endevents onAgentMicrophone. - The
vadfield inWidgetConfigand the widget’s VAD configuration section. - The
@ricky0123/vad-webandonnxruntime-webpeer dependencies. Uninstall them if you added them for VAD.
Correction: token minting uses ttl_seconds
The server-side token-minting example on the Browser Agent SDK overview posted {"ttl": 30} to POST /v1/auth/grant. The field is named ttl_seconds:
/v1/auth/grant accepts unrecognized fields, returns 200, and issues a token with the 30-second default, so a request built from the old example raised no error while ignoring the requested lifetime. If you copied that example and increased the number, your tokens expired after 30 seconds.
Nova-3 Pharma: Speech-to-Text for Pharmaceutical Use Cases (English)
We have released nova-3-pharma, a new Nova-3 model purpose-built for pharmaceutical vocabulary, with a focus on accurate drug-name recognition. It is designed for pharmacy and healthcare voice-agent workflows where getting the medication right matters most. Available in English for both batch and streaming.
September 15, 2026
India Endpoint Now Generally Available
The Deepgram India endpoint (api.in.deepgram.com) is now generally available for customers requiring data processing within India.
Supported APIs
The India endpoint supports the following Deepgram APIs:
September 15, 2026
Deepgram Self-Hosted September 2026 Release (260915)
Container Images (release 260915)
-
quay.io/deepgram/self-hosted-api:release-260915- Equivalent image to:
quay.io/deepgram/self-hosted-api:1.192.37
- Equivalent image to:
-
quay.io/deepgram/self-hosted-engine:release-260915-
Equivalent image to:
quay.io/deepgram/self-hosted-engine:3.131.2
-
Minimum required NVIDIA driver version:
>=580, open kernel module flavor- The minimum driver version has increased this release, up from
>=570.172.08. The Engine image is built against CUDA 13, and NVIDIA’s official CUDA 13 support begins with the580driver branch. NVIDIA’s documentation states the CUDA 13 minimum as>=580rather than a specific version. - The driver must be the open kernel module flavor, for example
nvidia-driver-580-open. The proprietary build of the same version is not supported.- On Blackwell hardware the consequence is total: NVIDIA never added Blackwell support to the proprietary
580branch, so on that build a Blackwell GPU disappears from the system,nvidia-smireports no devices, and Engine will not start. See Drivers and Containerization Platforms for installation and verification steps.
- On Blackwell hardware the consequence is total: NVIDIA never added Blackwell support to the proprietary
- The minimum driver version has increased this release, up from
-
-
quay.io/deepgram/self-hosted-license-proxy:release-260915- Equivalent image to:
quay.io/deepgram/self-hosted-license-proxy:1.11.1
- Equivalent image to:
-
quay.io/deepgram/self-hosted-billing:release-260915- Equivalent image to:
quay.io/deepgram/self-hosted-billing:1.14.0-2
- Equivalent image to:
SDK releases
Deepgram Python SDK v7.9.0, JavaScript SDK v5.11.0, and Java SDK
v0.10.0 are now available. All three releases add Voice Agent
ForceEndTurn support, Flux TTS expressivity controls for Agent providers, and
the expanded Flux TTS speed range of 0.5 through 1.5 in 0.05 increments.
@deepgram/react 0.2.0
@deepgram/react 0.2.0 adds runtime Voice Agent controls and makes React session lifecycle handling safer. The release updates the React layer for @deepgram/agents 0.1.2 and @deepgram/sdk 5.9.0.
Hold a function call until the user's turn is confirmed
The agent begins building a reply before speech-to-text confirms the user has finished speaking. Agent audio is held until that confirmation, but function calls have always dispatched as soon as the LLM emitted them. For a function with an irreversible side effect, such as ending a call or booking an appointment, that meant the action could run for a turn the user then continued.
Nova-3 Adds Kazakh, Plus Improved Models for Estonian, Hebrew, Latvian, Lithuanian, Macedonian, Malay, and Polish
We have added Kazakh as a new Nova-3 language and released improved Nova-3 monolingual models for seven existing languages. These updates expand language coverage and enhance transcription quality for batch and streaming workloads.
Wider speed range for Flux TTS
Flux TTS speed now runs from 0.5 to 1.5 in 0.05 increments, up from 0.85–1.15. The default stays 1.0, and every previously accepted value still works.
The wider range applies everywhere speed is accepted on Flux TTS: the /v2/speak streaming WebSocket (at connection and mid-session with Configure), the /v2/speak batch REST endpoint, and agent.speak.provider.speed in Voice Agent sessions running provider.version v2. A value outside the range is still rejected with SPEED_OUT_OF_RANGE, and one off the 0.05 increment with SPEED_INCREMENT_INVALID.