> For clean Markdown of any page, append .md to the page URL.
> For a complete documentation index, see https://developers.deepgram.com/llms.txt.
> For AI client integration (Claude Code, Cursor, etc.), connect to the MCP server at https://developers.deepgram.com/_mcp/server.

# December 29, 2025

> Deepgram Self-Hosted release 251229 adds Aura-2 TTS multilingual support (Dutch, German, French, Italian, Japanese), PHI redaction, and Flux Engine metrics for API 1.173.4 and Engine 3.107.0.

## Deepgram Self-Hosted December 2025 Release (251229)

### Container Images (release 251229)

* `quay.io/deepgram/self-hosted-api:release-251229`
  * Equivalent image to:
    * `quay.io/deepgram/self-hosted-api:1.173.4`

* `quay.io/deepgram/self-hosted-engine:release-251229`
  * Equivalent image to:
    * `quay.io/deepgram/self-hosted-engine:3.107.0`

  * Minimum required NVIDIA driver version: `>=570.172.08`

* `quay.io/deepgram/self-hosted-license-proxy:release-251229`
  * Equivalent image to:
    * `quay.io/deepgram/self-hosted-license-proxy:1.9.2`

* `quay.io/deepgram/self-hosted-billing:release-251229`
  * Equivalent image to:
    * `quay.io/deepgram/self-hosted-billing:1.12.1`

### This Release Contains The Following Changes

* **Expands Aura-2 TTS language support** - Adds TTS support for Dutch, German, French, Italian, and Japanese. See the [relevant changelog entry](https://developers.deepgram.com/changelog#aura-2-tts-language-expansion). Reach out to your Deepgram representative to obtain the new Aura-2 models.

* **Adds Engine metrics for Flux** - Adds `flux_max_streams`, `flux_used_streams`, `flux_fraction_streams`, and `flux_cursor_latency` metrics to the Engine container for Flux monitoring and auto-scaling.

* **Adds PHI redaction category** - Enables the use of `redact=phi` to redact six applicable sub-categories of PHI entities. See the [related changelog entry](https://developers.deepgram.com/changelog#phi-redaction-now-available-for-batch-and-streaming-speech-to-text) for details.

* **Allows optional blocking on model pre-loading before Engine becomes ready** - By default, [models pre-load](https://support.deepgram.com/en-US/deepgram/article/ART-528-how-can-i-preload-models-to-reduce-cold-start-latency) in the background, which can cause a delay on the first request. Setting `blocking = true` under `[preload_models]` in engine.toml makes the Engine wait until model pre-loading completes before accepting traffic. The tradeoff is longer startup time (potentially minutes), so orchestration and health checks should allow for a delayed readiness signal.

* **Includes General Improvements** — Keeps our software up-to-date.