Skip to navigation

Changelog

Update Nova-3 Keyterms Mid-Stream

Swap in the vocabulary for each step of a call without reconnecting. To replace a Nova-3 stream’s keyterms during a /v1/listen session, send a keyterms array in a Configure message:

{ "type": "Configure", "keyterms": ["Deepgram", "customer service"] }

Each keyterms array replaces the whole list, including keyterms set with the keyterm query parameter, and an empty array [] clears them. The 500-token keyterm limit that applies to the keyterm query parameter also applies to each update; an over-limit update returns an Error and leaves the stream’s keyterms unchanged. The same message can also turn formatting features on or off, for example "features": { "numerals": true }.

Speech-to-Text

October 1, 2026

Deepgram Self-Hosted October 2026 Release (261001)

Container Images (release 261001)

  • quay.io/deepgram/self-hosted-api:release-261001

    • Equivalent image to:
      • quay.io/deepgram/self-hosted-api:1.192.51
  • quay.io/deepgram/self-hosted-engine:release-261001

    • Equivalent image to:

      • quay.io/deepgram/self-hosted-engine:3.135.0
    • Minimum required NVIDIA driver version: >=580, open kernel module flavor. See Drivers and Containerization Platforms for installation and verification steps.

  • quay.io/deepgram/self-hosted-license-proxy:release-261001

    • Equivalent image to:
      • quay.io/deepgram/self-hosted-license-proxy:1.11.1-1
  • quay.io/deepgram/self-hosted-billing:release-261001

    • Equivalent image to:
      • quay.io/deepgram/self-hosted-billing:1.14.0-3
Self Hosted

Flux TTS Inline Controls: Pause and Pronunciation

Flux TTS (/v2/speak) now supports inline pause and pronunciation controls in the text you send, alongside speed.

  • Pause control. Insert silence with an escaped marker such as \{pause:1s\}: 500–3000 ms in 100 ms steps, up to 8 per request. Available on batch (REST) requests.
  • Pronunciation control (Early Access). Override how a word is said with inline IPA, \{"word": "dupilumab", "pronounce": "duːˈpɪljuːmæb"\}, on both batch and streaming. Results can vary between generations while the feature is in Early Access, so generate each term several times before relying on it in production.
  • Combination guardrails. Combinations that degrade audio are rejected with a clear error instead of producing bad audio: pronunciation cannot be combined with pause or with a speed other than 1.0, and speed is capped at 1.15 when a pause is present. Rejections return CONTROL_COMBINATION_INVALID or PAUSE_SPEED_CAP_EXCEEDED on batch, and ConfigureFailure or DATA-0002 on streaming.
  • Reporting. Applied controls are counted in controls_applied on streaming and in the dg-pronunciations-applied and dg-breaks-applied headers on batch.
Text-to-Speech

Toggle Numerals Mid-Stream on Flux STT

Switch Flux STT to digits for a PIN, phone number, or order number, then back to words, without reconnecting. To turn Numerals on or off during a Flux STT stream, send numerals as a boolean in a Configure message:

{ "type": "Configure", "numerals": true }

The update applies to transcripts Flux STT sends after it processes the message, and ConfigureSuccess now includes numerals in the full active configuration it echoes. The numerals query parameter still sets the initial value when the stream opens. This replaces the connection-time-only behavior in the July 17 entry.

Speech-to-Text

Nova-3 Improved Models for Flemish, German (Switzerland), Lithuanian, and Portuguese

We have released improved Nova-3 monolingual models for four existing languages. These updates enhance transcription quality for batch and streaming workloads.

Speech-to-Text

Nova-3 Improved Models for Danish, Estonian, Flemish, Italian, Lithuanian, Macedonian, Polish, Urdu, and Vietnamese

We have released improved Nova-3 monolingual models for nine existing languages. These updates enhance transcription quality for batch and streaming workloads.

Speech-to-Text

Correction: Browser Agent SDK has no client-side VAD

The Browser Agent SDK never shipped client-side Silero voice activity detection. Documentation published with the SDK described it as an available feature, and that documentation was wrong. The following did not exist in any released package:

  • The vad option on MicrophoneOptions in @deepgram/agents.
  • The speechThreshold and silenceThreshold VAD options.
  • The speech-start and speech-end events on AgentMicrophone.
  • The vad field in WidgetConfig and the widget’s VAD configuration section.
  • The @ricky0123/vad-web and onnxruntime-web peer dependencies. Uninstall them if you added them for VAD.
Voice Agent

Correction: token minting uses ttl_seconds

The server-side token-minting example on the Browser Agent SDK overview posted {"ttl": 30} to POST /v1/auth/grant. The field is named ttl_seconds:

body: JSON.stringify({ ttl_seconds: 30 }),

/v1/auth/grant accepts unrecognized fields, returns 200, and issues a token with the 30-second default, so a request built from the old example raised no error while ignoring the requested lifetime. If you copied that example and increased the number, your tokens expired after 30 seconds.

Voice Agent

Nova-3 Pharma: Speech-to-Text for Pharmaceutical Use Cases (English)

We have released nova-3-pharma, a new Nova-3 model purpose-built for pharmaceutical vocabulary, with a focus on accurate drug-name recognition. It is designed for pharmacy and healthcare voice-agent workflows where getting the medication right matters most. Available in English for both batch and streaming.

Speech-to-Text

September 15, 2026

India Endpoint Now Generally Available

The Deepgram India endpoint (api.in.deepgram.com) is now generally available for customers requiring data processing within India.

Supported APIs

The India endpoint supports the following Deepgram APIs:

Platform

September 15, 2026

Deepgram Self-Hosted September 2026 Release (260915)

Container Images (release 260915)

  • quay.io/deepgram/self-hosted-api:release-260915

    • Equivalent image to:
      • quay.io/deepgram/self-hosted-api:1.192.37
  • quay.io/deepgram/self-hosted-engine:release-260915

    • Equivalent image to:

      • quay.io/deepgram/self-hosted-engine:3.131.2
    • Minimum required NVIDIA driver version: >=580, open kernel module flavor

      • The minimum driver version has increased this release, up from >=570.172.08. The Engine image is built against CUDA 13, and NVIDIA’s official CUDA 13 support begins with the 580 driver branch. NVIDIA’s documentation states the CUDA 13 minimum as >=580 rather than a specific version.
      • The driver must be the open kernel module flavor, for example nvidia-driver-580-open. The proprietary build of the same version is not supported.
        • On Blackwell hardware the consequence is total: NVIDIA never added Blackwell support to the proprietary 580 branch, so on that build a Blackwell GPU disappears from the system, nvidia-smi reports no devices, and Engine will not start. See Drivers and Containerization Platforms for installation and verification steps.
  • quay.io/deepgram/self-hosted-license-proxy:release-260915

    • Equivalent image to:
      • quay.io/deepgram/self-hosted-license-proxy:1.11.1
  • quay.io/deepgram/self-hosted-billing:release-260915

    • Equivalent image to:
      • quay.io/deepgram/self-hosted-billing:1.14.0-2
Self Hosted

SDK releases

Deepgram Python SDK v7.9.0, JavaScript SDK v5.11.0, and Java SDK v0.10.0 are now available. All three releases add Voice Agent ForceEndTurn support, Flux TTS expressivity controls for Agent providers, and the expanded Flux TTS speed range of 0.5 through 1.5 in 0.05 increments.

Developer Tools

@deepgram/react 0.2.0

@deepgram/react 0.2.0 adds runtime Voice Agent controls and makes React session lifecycle handling safer. The release updates the React layer for @deepgram/agents 0.1.2 and @deepgram/sdk 5.9.0.

Voice Agent

Hold a function call until the user's turn is confirmed

The agent begins building a reply before speech-to-text confirms the user has finished speaking. Agent audio is held until that confirmation, but function calls have always dispatched as soon as the LLM emitted them. For a function with an irreversible side effect, such as ending a call or booking an appointment, that meant the action could run for a turn the user then continued.

Voice Agent

Nova-3 Adds Kazakh, Plus Improved Models for Estonian, Hebrew, Latvian, Lithuanian, Macedonian, Malay, and Polish

We have added Kazakh as a new Nova-3 language and released improved Nova-3 monolingual models for seven existing languages. These updates expand language coverage and enhance transcription quality for batch and streaming workloads.

Speech-to-Text

Wider speed range for Flux TTS

Flux TTS speed now runs from 0.5 to 1.5 in 0.05 increments, up from 0.85–1.15. The default stays 1.0, and every previously accepted value still works.

The wider range applies everywhere speed is accepted on Flux TTS: the /v2/speak streaming WebSocket (at connection and mid-session with Configure), the /v2/speak batch REST endpoint, and agent.speak.provider.speed in Voice Agent sessions running provider.version v2. A value outside the range is still rejected with SPEED_OUT_OF_RANGE, and one off the 0.05 increment with SPEED_INCREMENT_INVALID.

Text-to-SpeechVoice Agent

Flexible turn-taking control for Flux

Flux now gives you full control over how turns end. Beyond Flux’s native end-of-turn detection, you can override it, suppress it entirely, or blend the two — and switch between these modes on the fly, even for a single turn, with a Configure message. Turn-taking goes from a fixed behavior to something you shape around each moment of a conversation.

Speech-to-TextVoice Agent

Nova-3 Model Improvements for Bulgarian, Croatian, Estonian, Georgian, Italian, Latvian, Lithuanian, Malay, Marathi, and Telugu

We have released improved Nova-3 monolingual models for ten existing languages. These updates enhance transcription quality for batch and streaming workloads.

Speech-to-TextVoice Agent

Nova-3 Adds Assamese, Mongolian, and Pashto, Plus Improved Models for Czech, Danish, Swedish, Tagalog, and Turkish

We have added Assamese, Mongolian, and Pashto as new Nova-3 languages and released improved Nova-3 monolingual models for several existing languages. These updates expand language coverage and enhance transcription quality for batch and streaming workloads.

Speech-to-Text

August 26, 2026

Deepgram Self-Hosted August 2026 Release (260826)

Container Images (release 260826)

  • quay.io/deepgram/self-hosted-api:release-260826

    • Equivalent image to:
      • quay.io/deepgram/self-hosted-api:1.192.31
  • quay.io/deepgram/self-hosted-engine:release-260826

    • Equivalent image to:

      • quay.io/deepgram/self-hosted-engine:3.128.0
    • Minimum required NVIDIA driver version: >=570.172.08

  • quay.io/deepgram/self-hosted-license-proxy:release-260826

    • Equivalent image to:
      • quay.io/deepgram/self-hosted-license-proxy:1.11.0-1
  • quay.io/deepgram/self-hosted-billing:release-260826

    • Equivalent image to:
      • quay.io/deepgram/self-hosted-billing:1.14.0-1
Self Hosted