Invoke a Deepgram SageMaker Endpoint
Once your endpoint is deployed and in service, you invoke it to transcribe audio. A real-time endpoint supports two invocation modes, depending on how you need the response returned.
Passing Deepgram parameters. For synchronous invocations, the Deepgram model and feature parameters are passed in the CustomAttributes field (the X-Amzn-SageMaker-Custom-Attributes header) as v1/listen?model=...&language=.... For streaming, the same values are split across ModelInvocationPath (v1/listen) and ModelQueryString. In all cases an API path such as v1/listen is required — without it the container returns a 404. The examples on this page use v1/listen (speech-to-text), but other routes are available (for example, v1/speak for text-to-speech).
Complete, runnable examples for both modes — in Python, TypeScript, and Java — are maintained in the deepgram-devs/dg-sagemaker repository. The sections below explain each mode and link to the corresponding example. See the repository’s README for setup and prerequisites.
Use the Deepgram SDKs with the SageMaker transport
You don’t have to call the AWS APIs directly. The Deepgram SDKs can target a SageMaker endpoint through a SageMaker transport, so you keep the same client-side request and response patterns whether you call the Deepgram-hosted API or your own SageMaker deployment. You swap the transport; your listen request and result-handling code stays the same.
For example, the Deepgram Java SDK pairs with the Deepgram SageMaker transport (com.deepgram:deepgram-sagemaker):
The remaining sections show the underlying AWS APIs directly, which apply to any language.
Streaming (real-time)
Use streaming for live, interactive transcription over a persistent bidirectional connection. You send audio chunks and receive transcription results as the audio is processed, up to 30 minutes per connection.
Streaming uses the HTTP/2 bidirectional streaming client (@aws-sdk/client-sagemaker-runtime-http2 in TypeScript, aws-sdk-sagemaker-runtime-http2 in Python) against the SageMaker bidirectional runtime endpoint (https://runtime.sagemaker.<region>.amazonaws.com:8443). The request Body is an async iterable of payload parts:
- Binary audio is sent as a
Bytespayload withDataType: "BINARY". - Control messages (for example,
KeepAliveandCloseStream) are sent as UTF-8 encoded JSON withDataType: "UTF8".
Always include :8443 in the endpoint URL. The bidirectional streaming runtime listens on port 8443, not 443. A streaming client that hangs with no error and never receives a response is almost always pointed at the endpoint without :8443.
Python
TypeScript
Requires aws-sdk-sagemaker-runtime-http2 0.11 or later with the awscrt extra (pip install "aws-sdk-sagemaker-runtime-http2[awscrt]>=0.11"). The client takes explicit credentials and an AWS CRT transport; payload events are typed.
For the complete examples — file and microphone capture, payload wrapping, keepalive handling, and stream processing — see:
- TypeScript:
js-stt/stt.file.tsandstt.microphone.ts - Python:
python-stt/stt_wav_stress.py(streamsubcommand)
Synchronous (real-time)
Use synchronous invocation to transcribe a single pre-recorded file and receive the full transcript in one immediate response. This is Deepgram’s “batch” transcription on a real-time endpoint — there is no streaming connection and no queue. The request body is capped at 25 MB; use streaming for larger audio.
You send the audio as the request body to InvokeEndpoint, pass the Deepgram parameters via CustomAttributes, and parse the transcript from the JSON response.
For the complete example, see python-stt/stt_wav_stress.py (batch subcommand) in the repository.
Asynchronous
Asynchronous invocation (InvokeEndpointAsync, files up to 1 GB) is temporarily not supported for Marketplace-hosted Deepgram. If your use case needs it, contact a Deepgram representative.