Skip to main content

AWS Startups

Deploy voice agents with Pipecat, Deepgram and Amazon SageMaker AI

by Victor Wang, Dustin Liu, Kwindla Hultman Kramer, Susmitha Marupaka, Daniel Wirjo and Andre Gomes | 1 October 2026

Deploy voice agents with Pipecat, Deepgram and Amazon SageMaker AI

This post is written in collaboration with Pipecat and Deepgram.

Imagine calling a support line and having a natural, back-and-forth conversation with an AI agent that understands you, responds instantly, and you can interrupt mid-sentence just like a human would. That experience is compelling, but building it is hard. A production voice agent means orchestrating a lot of moving parts: voice activity detection (VAD), speech-to-text (STT), a large language model (LLM), text-to-speech (TTS), telephony, auto scaling infrastructure, and observability. All of these must be wired into a single streaming pipeline. The hard part has always been the distance between “I want to build a voice agent” and “it’s running in production.” There are too many services to coordinate, latency compounds at every boundary, and there’s too much infrastructure to get right before you can even test whether the idea works.

In previous posts, we showed how to deploy speech models with vLLM on Amazon SageMaker AI and how to use bidirectional streaming for real-time inference on Amazon SageMaker AI, focusing on individual model deployment and the streaming protocol. This post goes a step further: you deploy a complete voice agent that orchestrates every pipeline stage using Pipecat, an open-source Python framework for voice AI pipelines. Deepgram STT and TTS models run on Amazon SageMaker AI with bidirectional streaming, and the LLM runs on Amazon Bedrock. AI-assisted deployment skills guide you through setup, so you can go from cloning a repository to a phone number taking live voice calls with minimal manual configuration.

Solution overview

The architecture directly addresses the coordination and latency challenges. Instead of stitching services together yourself, Pipecat orchestrates every stage in one streaming pipeline, and each boundary is optimized to keep latency low. A caller connects through WebRTC or the public switched telephone network (PSTN) to a Pipecat pipeline running on Amazon Elastic Container Service (Amazon ECS) with AWS Fargate. The pipeline processes audio through multiple stages: VAD detects when the caller is speaking, STT transcribes the speech, the LLM generates a response, and TTS synthesizes the audio output.

Amazon SageMaker AI bidirectional streaming uses HTTP/2 on port 8443 to bridge client connections to WebSocket-enabled containers. Your client connects to runtime.sagemaker.<region>.amazonaws.com:8443. Amazon SageMaker AI converts each HTTP/2 stream into WebSocket frames and forwards them to your model container at ws://localhost:8080/invocations-bidirectional-stream. Your STT and TTS models receive continuous audio streams without request/response round-trips. A virtual private cloud (VPC) endpoint for Amazon SageMaker Runtime handles both standard invoke (port 443) and bidirectional streaming (port 8443), so audio never leaves your VPC.

Pipecat manages frame-based data flow between processors, handles buffering, and supports barge-in (interruption handling) when the caller speaks while the agent is responding.

The deployment includes production infrastructure by default:

  • Session-aware auto scaling that scales on active calls per container, not CPU.
  • An Amazon CloudWatch dashboard tracking end-to-end latency, agent response time, turn count, interruption count, and audio quality metrics.
  • Graceful shutdown with drain mode and ECS Task Scale-in Protection designed to avoid dropped calls during scale-in.
  • An agent-to-agent architecture for extending capabilities through AWS Cloud Map discovery.
  • Topic and tool scoping that constrains what the agent can act on and discuss.

Because callers interact with the agent in real time, keep safety controls out of the turn-by-turn latency budget rather than adding a synchronous filtering call to the critical path. Scope the LLM’s tool registry to only the actions your use case needs. Add an explicit denied-topics instruction to the system prompt so the agent declines and hands off gracefully instead of engaging. Log every transcript for asynchronous review. This constrains what the agent can say and do without adding a blocking round-trip to the pipeline.

For the full implementation, see the Guidance for voice agents on AWS repository on GitHub.

Prerequisites

Before deploying, confirm the following:

  • An AWS account with Amazon Bedrock model access enabled for Claude in your target Region.
  • Service quota for ml.g6.2xlarge and ml.g6.12xlarge endpoint instances (request through Service Quotas).
  • Node.js 18+, Python 3.12+, AWS Command Line Interface (AWS CLI) v2, and Finch (or Docker) installed locally.
  • AWS Cloud Development Kit (AWS CDK) bootstrapped in your account and Region (npx cdk bootstrap aws://<your-account-id>/<your-region>).
  • A Daily.co account for WebRTC transport.
  • Deepgram model subscriptions for STT and TTS on AWS Marketplace.

Amazon Bedrock model access for Claude, the ml.g6 instance types, and the Deepgram models on AWS Marketplace vary by AWS Region. Confirm availability in your target Region before deploying.

Deploy speech models on Amazon SageMaker AI

Start by deploying the infrastructure that hosts your speech models.

Clone the repository and install dependencies

Clone the repository and install the infrastructure dependencies:

git clone https://github.com/aws-solutions-library-samples/sample-voice-agent.git
cd sample-voice-agent/infrastructure
npm install

Create a new .env file based on .env.example and update it accordingly:

cp .env.example .env

AI-guided deployment (recommended)

For an interactive deployment experience:

  • Open the repository in Claude Code (or your preferred AI-assisted IDE). The repository contains skills under .agents/.
  • Run /deploy-sagemaker. The skill checks prerequisites, gathers your API keys, deploys CDK stacks in dependency order, pushes secrets to AWS Secrets Manager, and verifies all components are healthy.
  • Run /configure-daily to set up your phone number and /verify-deployment for a full infrastructure health check.

Manual deployment

Run the deployment script:

./deploy.sh deploy

This deploys all stacks in dependency order. Deployment takes approximately 15 minutes.

The CDK deployment creates a VPC with private subnets, VPC endpoints (including port 8443 for bidirectional streaming), secrets management, and two Amazon SageMaker AI endpoints:

  • STT: Deepgram Nova-3 (Deepgram Voice AI Nova-3 Monolingual Speech-to-Text (STT) Streaming tested) on ml.g6.2xlarge (1x L4 GPU).
  • TTS: Deepgram Aura-2 (Deepgram Voice AI Aura-2 Text-to-Speech (TTS) tested) on ml.g6.12xlarge (4x L4 GPU).

Verify endpoints

Confirm both endpoints are in service:

aws sagemaker describe-endpoint \
  --endpoint-name <your-stt-endpoint-name> \
  --query "EndpointStatus" --output text

aws sagemaker describe-endpoint \
  --endpoint-name <your-tts-endpoint-name> \
  --query "EndpointStatus" --output text

Both commands return InService when the endpoints are ready. Find your endpoint names in the CDK stack outputs or in the Amazon SageMaker AI console.

Build the Pipecat voice agent pipeline

With the speech model endpoints running, the voice pipeline is already deployed as part of the ./deploy.sh deploy command from the previous section. The Amazon ECS Fargate service (always-on, auto scaling 1 to 10 tasks) and the webhook handler (AWS Lambda behind Amazon API Gateway) are running. This section explains how the pipeline works.

Pipeline structure

The Pipecat pipeline in backend/voice-agent/app/pipeline_ecs.py defines the processing chain:

pipeline = Pipeline([
    transport.input(),           # WebRTC audio from caller
    stt,                         # Deepgram Nova-3 via SageMaker BiDi
    context_aggregator.user(),   # Accumulate user message context
    llm,                         # Bedrock Claude: generate response
    tts,                         # Deepgram Aura-2 via SageMaker BiDi
    transport.output(),          # WebRTC audio to caller
    context_aggregator.assistant(),
])

task = PipelineTask(
    pipeline,
    params=PipelineParams(allow_interruptions=True),
)

Each processor handles a specific stage:

  • Silero VAD: Detects speech boundaries with a 0.3-second silence threshold.
  • STT (Deepgram Nova-3): Streams audio to the Amazon SageMaker AI endpoint through HTTP/2 bidirectional streaming on port 8443, receiving transcriptions in real time.
  • LLM (Amazon Bedrock): Generates responses using Claude Haiku 4.5 with ConverseStream for token-by-token output.
  • TTS (Deepgram Aura-2): Converts each text chunk to audio as it arrives through bidirectional streaming, producing speech output with minimal buffering delay.

Setting allow_interruptions=True activates barge-in support. When the caller speaks mid-response, VAD detects the new speech, the current TTS output is canceled, and the pipeline generates a fresh response.

Verify the deployment

Confirm the Amazon ECS service is running:

aws ecs describe-services \
  --cluster voice-agent-{environment}-cluster \
  --services voice-agent-{environment}-service \
  --query "services[0].{status:status,running:runningCount}"

# The default ENVIRONMENT value in .env is poc,
# so the cluster name would be voice-agent-poc-cluster

Test locally

Test the full pipeline from your browser without a phone number:

cd backend/voice-agent
pip install -r requirements.txt
cp .env.example .env

# Point the pipeline at the endpoints you deployed earlier, and set AWS_REGION
# STT_PROVIDER=sagemaker
# TTS_PROVIDER=sagemaker
# STT_ENDPOINT_NAME=<your-stt-endpoint-name>
# TTS_ENDPOINT_NAME=<your-tts-endpoint-name>

python -m app.local_main
# Open http://localhost:7860 and click Connect

Speak into your microphone to hear the agent respond. All pipeline features, including tool calling and interruption handling, work identically to the production deployment.

Clean up

To avoid ongoing charges, remove all resources when you no longer need them. Amazon SageMaker AI GPU endpoints are the most expensive component (billed per hour regardless of usage).

Destroy all stacks:

cd infrastructure
npx cdk destroy --all --force

After stack deletion, verify no resources remain:

aws cloudformation list-stacks \
  --stack-status-filter CREATE_COMPLETE UPDATE_COMPLETE \
  --query "StackSummaries[?starts_with(StackName, 'VoiceAgent')].StackName"

If you purchased a Daily.co phone number, release it from the Daily.co dashboard under Phone > Numbers.

How Pipecat wraps the SageMaker bidirectional streaming interface

Pipecat uses a three-layer architecture to integrate speech models running on Amazon SageMaker AI. Understanding this pattern helps you wrap additional models beyond the Deepgram services included in the sample. The source code for all layers lives in the Pipecat repository under pipecat/services/aws/sagemaker/.

Layer 1: SageMakerBidiClient

At the base, Pipecat provides SageMakerBidiClient — a reusable HTTP/2 client that handles the SageMaker bidirectional streaming protocol. It manages SigV4 authentication, session lifecycle, and binary/text message framing:

from pipecat.services.aws.sagemaker.bidi_client import SageMakerBidiClient

client = SageMakerBidiClient(
    endpoint_name="my-endpoint",
    region="us-east-2",
    model_invocation_path="v1/speak",       # Container route
    model_query_string="model=aura-2&encoding=linear16&sample_rate=8000",
)

await client.start_session()
await client.send_json({"type": "Speak", "text": "Hello"})
response = await client.receive_response()   # Binary audio or JSON control
await client.close_session()

The client connects to runtime.sagemaker.<region>.amazonaws.com:8443 and uses InvokeEndpointWithBidirectionalStream from the AWS SDK. The model_invocation_path and model_query_string parameters map to the container’s internal routing, so you can target different model APIs on the same endpoint.

Layer 2: Service wrapper (TTSService or STTService)

The service wrapper extends Pipecat’s base TTSService or STTService class and implements the model-specific protocol. For example, DeepgramSageMakerTTSService translates Pipecat’s run_tts(text) interface into Deepgram’s WebSocket protocol messages:

class DeepgramSageMakerTTSService(TTSService):
    async def run_tts(self, text: str) -> AsyncGenerator[Frame, None]:
        # Send Deepgram protocol messages over BiDi
        await self._client.send_json({"type": "Speak", "text": text})
        await self._client.send_json({"type": "Flush"})

        # Yield audio frames as they arrive
        yield TTSStartedFrame()
        while True:
            chunk = await self._audio_queue.get()
            if chunk is None:  # Flushed sentinel
                break
            yield TTSAudioRawFrame(audio=chunk)
        yield TTSStoppedFrame()

    async def handle_interruption(self):
        # Barge-in: discard queued speech
        await self._client.send_json({"type": "Clear"})

A background task (_process_responses) continuously reads from the BiDi stream, distinguishing binary audio payloads from JSON control messages (Flushed, Warning, Error, Close) and routing them to the appropriate handler.

The STT wrapper follows the same pattern with model_invocation_path="v1/listen", sending raw audio by using send_audio_chunk() and parsing Deepgram’s transcription JSON responses. It also sends KeepAlive messages every 5 seconds to maintain the connection during silence, and Finalize when VAD detects end-of-speech to flush partial results.

Layer 3: Pipeline integration

The service factory (backend/voice-agent/app/services/factory.py) selects the provider at runtime based on environment variables. Switch between cloud APIs and SageMaker endpoints without changing pipeline code:

# Environment: STT_PROVIDER=sagemaker, TTS_PROVIDER=sagemaker
stt = create_stt_service(config)  # Returns DeepgramSageMakerSTTService
tts = create_tts_service(config)  # Returns DeepgramSageMakerTTSService

Wrap your own model

To add a new speech model deployed on Amazon SageMaker AI with bidirectional streaming, follow this pattern:

  • Identify the container’s protocol. Determine the route path (for example, /v1/synthesize), query parameters, and message format your model container expects (JSON commands, binary audio, or both). Your container must accept WebSocket connections on port 8080 at the /invocations-bidirectional-stream path, which is where Amazon SageMaker AI forwards the bridged stream.
  • Extend the base service class. Subclass TTSService or STTService and implement the required methods: For TTS, run_tts(text) sends text, yields TTSAudioRawFrame chunks, and handles flush/completion signals. For STT, run_stt(audio) sends audio chunks, and a response processor emits TranscriptionFrame when transcriptions arrive.
  • Configure the BiDi client. Set model_invocation_path and model_query_string to match your container’s routing. The SageMakerBidiClient handles authentication and HTTP/2 framing.
  • Handle lifecycle events. Implement _connect() to start the session and launch background tasks, _disconnect() for graceful teardown, and handle_interruption() if your model supports mid-stream cancellation.

Any model container that exposes a bidirectional streaming interface on Amazon SageMaker AI can be wrapped using this pattern. This includes third-party marketplace models and custom models you have deployed with vLLM or Triton. To contribute a new service wrapper, fork the Pipecat repository, implement your service class following the preceding pattern, and open a pull request. You can use an AI coding agent such as Claude Code or Kiro to scaffold the implementation from the existing Deepgram wrappers as a reference.

Optimize for latency

Each stage of the voice agent pipeline contributes to end-to-end latency. The following strategies help you achieve sub-second voice-in to voice-out response time.

  • VAD and turn detection: The default 0.3-second silence threshold works well for most conversations. For smarter endpointing that reduces false turn-end detections, switch to Deepgram Nova-3 with Flux which has built-in turn detection, or deploy the Pipecat smart-turn model on Amazon SageMaker AI to distinguish mid-sentence pauses from actual turn endings using conversation context.
  • LLM time-to-first-token (TTFT): Claude Haiku 4.5 on Amazon Bedrock with cross-Region inference profiles proactively routes requests across Regions within a defined geographic boundary, providing fast streaming responses. For even lower TTFT, deploy a model optimized for speed such as NVIDIA Nemotron Ultra on Amazon SageMaker JumpStart, where you control instance sizing and batching configuration.
  • STT and TTS streaming: Bidirectional streaming eliminates request/response overhead. Audio flows continuously through the VPC endpoint without waiting for complete utterances. For fine-tuning custom voice models or running specialized speech models, deploy publicly available alternatives with vLLM on Amazon SageMaker AI as described in Build real-time voice applications with Amazon SageMaker AI and vLLM.
  • Always-on architecture: New ECS tasks take approximately 90 seconds to reach traffic-ready state. Keep minimum capacity at 1 to avoid cold-start delays on the first call.

Conclusion

In this post, we showed how to deploy a voice agent that orchestrates STT, LLM, and TTS into a real-time conversational pipeline using Pipecat and Amazon SageMaker AI bidirectional streaming. Audio stays within your VPC, streaming at every boundary keeps latency low, and Pipecat handles interruption and turn management.

You’ve deployed the infrastructure, and the pipeline is running. From here, you can:

  • Connect a knowledge base agent for retrieval augmented generation (RAG) over your documents using the built-in agent-to-agent architecture.
  • Add custom tools (CRM lookups, appointment scheduling, order status) using the pluggable tool registry, without modifying pipeline code.
  • Swap STT, TTS, or LLM providers using the service factories to test latency and cost tradeoffs for your workload.

Clone the Guidance for voice agents on AWS repository and deploy with ./deploy.sh deploy to get started.

For more information, refer to:

Acknowledgement

The authors thank Vivek Gangasani and Andrew Smith for their review and contributions to this post.

About the authors

Victor Wang

Victor Wang is a Staff Software Engineer at Deepgram and Technical Advisor to the VP of Engineering based in San Francisco, CA. Before joining Deepgram, Victor held multiple roles at AWS including Sr. Solutions Architect, Technical Program Manager, Proserve Consultant, and Software Developer.

Kwindla Hultman Kramer

Kwindla Hultman Kramer is the Co-founder and CEO at Daily, pioneering low-latency real-time voice, video, and multimodal AI infrastructure. A leading voice AI thought leader, he created the open-source Pipecat framework for production voice agents and shares insights at voice AI meetups and his X account (@kwindla).

Daniel Wirjo

Daniel Wirjo is a Solutions Architect at AWS, focused on AI and SaaS startups. As a former startup CTO, he enjoys collaborating with founders and engineering leaders to drive growth and innovation on AWS. Outside of work, Daniel enjoys taking walks with a coffee in hand, appreciating nature, and learning new ideas.

Dustin Liu

Dustin Liu is a Solutions Architect at AWS, focused on supporting financial services and insurance (FSI) startups and SaaS companies. He has a diverse background spanning data engineering, data science, and machine learning, and he is passionate about using AI/ML to drive innovation and business transformation.

Susmitha Marupaka

Susmitha Marupaka leads go-to-market strategy for AWS AI Inference, helping enterprises, ISVs, and startups deploy, manage, and scale their generative AI models and agents on AWS. Outside of work, Susmitha is an accomplished dancer and a community builder.

Andre Gomes

Andre Gomes is a Frontier AI Solutions Architect at AWS, working with startups on large-scale training and real-time inference. As a former startup co-founder and CEO, he partners with founders to scale their workloads on AWS and reach customers worldwide. Andre earned his PhD at UFMG, with research at the Institute for Manufacturing, University of Cambridge, applying large language models and neural topic modeling to technology roadmapping.

Did you find what you were looking for today?

Let us know so we can improve the quality of the content on our pages