Troubleshooting
If you’re experiencing any issues with your Deepgram deployment on Amazon SageMaker AI, start with the Deepgram container logs in Amazon CloudWatch, then work through the common causes below.
View container logs
If you open the SageMaker AI Endpoint resource details, there will be a link to open the Amazon CloudWatch Log Group for that endpoint. Within the CloudWatch Log Group, there should be a Log Stream that contains the Deepgram logs for all components. You can use the Amazon CloudWatch Logs Live Tail feature to watch logs in near-real-time while you are sending requests to the Deepgram API, via the SageMaker AI APIs.
To use the CloudWatch Logs Live Tail feature locally, from the AWS CLI tool, you can use the following command.
Endpoint fails to start (CUDA / driver preflight)
Current Deepgram model packages run a CUDA 13 runtime and require NVIDIA driver 580 or later on the host. If the endpoint boots on an older default inference AMI — which is what happens when the Endpoint Configuration was created from the SageMaker AI console, or without InferenceAmiVersion set — the container fails its CUDA preflight check and the endpoint never reaches InService. The CloudWatch log stream for the endpoint contains a line similar to:
To fix this, create a new Endpoint Configuration with the AWS CLI, Boto3, or Terraform that sets InferenceAmiVersion to al2023-ami-sagemaker-inference-gpu-4-1 on the production variant, then create or update the endpoint with it. See Inference AMI Versions for the available versions and Deepgram’s recommendation.
Endpoint is InService but every request returns 400
A 400 on every request usually means the request does not match the product you deployed, not that the endpoint is unhealthy.
- Multilingual Nova-3 listings require
language=multi. Sendinglanguage=en(or any single language code) to a multilingual Nova-3 endpoint returns400, which looks like a dead endpoint. Passlanguage=multiin the query string. - Flux multilingual is selected by model name, not a language parameter. Use
model=flux-general-multi; there is nolanguageparameter for Flux multilingual. - Streaming-mode bundles reject synchronous invocation. A product listing published for streaming returns
400 No such model/language/tierwhen called through the synchronous/invocationspath (InvokeEndpoint). UseInvokeEndpointWithBidirectionalStreamfor streaming listings, or deploy the batch listing for synchronous and asynchronous invocation. See Invoke a Deepgram SageMaker Endpoint.
Endpoint stuck in Creating
If the endpoint stays in Creating well beyond the usual several minutes, or moves to Failed, check FailureReason:
Common causes:
ModelDataDownloadTimeoutInSecondsis too low. Large multilingual bundles can take longer than the default download window. Recreate the Endpoint Configuration with a higher value on the production variant (Deepgram’s CLI steps use600; large multilingual Nova-3 bundles may need more).- Capacity or quota. The instance type is unavailable in the Availability Zone, or your account quota for it is
0. AResourceLimitExceededfailure reason points to quota — see Requesting SageMaker Quota.
Checklist
If you experience any issues using Deepgram services running on the Amazon SageMaker AI platform, please review this checklist before contacting Deepgram support.
- Ensure that your application’s AWS IAM User or IAM Role has permission to call the
InvokeEndpointWithBidirectionalStreamSageMaker AI action. - Ensure your application is targeting the correct AWS account and region, where your SageMaker Endpoint exists.
- Ensure the Deepgram product you’ve deployed (eg. streaming Speech-to-Text), from the AWS Marketplace, corresponds to the Deepgram API you’re calling.
- Ensure the Endpoint Configuration pins
InferenceAmiVersiontoal2023-ami-sagemaker-inference-gpu-4-1. The SageMaker AI console does not expose this setting; create the Endpoint Configuration with the AWS CLI or API, or with Terraform. See Inference AMI Versions. - If you subscribed through a private offer in an AWS organization, ensure the linked account deploying the endpoint has accepted the offer or holds a License Manager entitlement. See Private offers.