> For clean Markdown of any page, append .md to the page URL.
> For a complete documentation index, see https://developers.deepgram.com/llms.txt.
> For AI client integration (Claude Code, Cursor, etc.), connect to the MCP server at https://developers.deepgram.com/_mcp/server.

# Build a Multi-Agent Architecture

> **Info**
>
> For more information and to use our reference implementation, visit the [Deepgram Multi-Agent repo](https://github.com/deepgram-devs/deepgram-voice-agent-multi-agent).

## Why Multi-Agent Architecture?

Traditional single-agent voice systems face fundamental limitations as complexity grows. This implementation demonstrates how to overcome these challenges by treating the conversation as a sequence of specialized phases/states, where each phase has its own:

* **System instructions** (the agent's prompt)
* **Transformed context** (summarized from previous conversation)
* **Specific tools** (2-4 focused functions per agent)

This approach solves critical problems:

* **Context Management**: Each agent starts fresh rather than accumulating entire conversation history
* **Focused Responsibility**: Agents excel at their specific task instead of juggling everything
* **Better Reliability**: Fewer functions per agent means clearer decision-making for the LLM
* **Easier Debugging**: Issues are isolated to specific agents and transitions

## Architecture Overview

```
┌─────────────────────────────────────────┐
│      Twilio Phone Call (Persistent)     │
│         WebSocket Connection            │
└──────────────────┬──────────────────────┘
                   │ Audio Stream
         ┌─────────▼──────────┐
         │  CallOrchestrator  │ ← Central coordinator
         │ ┌────────────────┐ │
         │ │ Audio Forward  │ │ ← Persistent task
         │ │     Task       │ │
         │ └────────────────┘ │
         └─────────┬──────────┘
                   │ Manages lifecycle
    ┌──────────────┼──────────────┐
    │              │              │
    ▼              ▼              ▼
┌─────────┐   ┌─────────┐   ┌─────────┐
│Qualifier│──►│ Advisor │──►│ Closer  │
│  Agent  │   │  Agent  │   │  Agent  │  ← Ephemeral Voice Agent sessions
└─────────┘   └─────────┘   └─────────┘
    │              │              │
    └──────────────┼──────────────┘
                   │
            Groq API (Llama 3.3 70B by default)
          Context Summarization
```

**Key Architecture Points**:

* **Persistent**: Twilio WebSocket and audio forwarding task remain active throughout
* **Ephemeral**: Voice Agent sessions are created/destroyed per agent
* **Orchestration**: CallOrchestrator manages all transitions and state
* **Context Transfer**: A small and fast LLM (via Groq API) summarizes conversation between agents

## Key Technical Concepts

### Session-Based Agent Switching

This implementation creates separate Voice Agent sessions for each specialized agent. The `CallOrchestrator` (orchestrator/call\_orchestrator.py) manages these transitions while maintaining the Twilio connection.

### Audio Task Persistence Pattern

A critical implementation detail: the audio forwarding task starts once and persists throughout all agent transitions. This ensures the Twilio WebSocket remains active while agents switch:

**`Python`**

```python Python
# Audio task starts ONCE and persists
if not self.audio_task or self.audio_task.done():
    self.audio_task = asyncio.create_task(self.forward_twilio_audio())
```

**`Java`**

```java Java
// Audio task starts ONCE and persists across agent transitions
if (audioTask == null || audioTask.isDone()) {
    audioTask = CompletableFuture.runAsync(this::forwardTwilioAudio);
}
```

### Context Manager Lifecycle

Voice Agent connections are async context managers that require manual lifecycle management:

**`Python`**

```python Python
# Creating connection
self.current_agent_context = self.current_client.agent.v1.connect()
self.current_agent_connection = await self.current_agent_context.__aenter__()

# Closing connection (keep audio task alive during transitions)
await self.close_current_agent(keep_audio_task=True)
```

**`Java`**

```java Java
// Creating connection
DeepgramClient currentClient = DeepgramClient.builder().build();
V1WebSocketClient wsClient = currentClient.agent().v1().v1WebSocket();
wsClient.connect().get(10, TimeUnit.SECONDS);

// Closing connection (keep audio task alive during transitions)
closeCurrentAgent(/* keepAudioTask */ true);
```

### Function Call Response Pattern

Functions must be called with proper timing to maintain conversation flow:

**`Python`**

```python Python
# 1. Send response to current agent
response = AgentV1FunctionCallResponseMessage(...)
await self.current_agent_connection.send_function_call_response(response)

# 2. Brief pause for processing
await asyncio.sleep(0.5)

# 3. Then perform the action (e.g., transition)
await self.transition_to_agent(next_agent)
```

**`Java`**

```java Java
// 1. Send response to current agent
wsClient.sendFunctionCallResponse(
    AgentV1SendFunctionCallResponse.builder()
        .id(request.getId())
        .output("{\"status\": \"success\"}")
        .build());

// 2. Brief pause for processing
Thread.sleep(500);

// 3. Then perform the action (e.g., transition)
transitionToAgent(nextAgent);
```

## The Three Agents

Each agent represents a distinct conversation phase with focused responsibilities:

### 1. Qualifier Agent (`agents/qualifier/config.py`)

* **Purpose**: Initial contact and lead qualification
* **Functions**: `handoff_to_next_agent`, `end_conversation`
* **Collects**: Name, location, specific needs

### 2. Advisor Agent (`agents/advisor/config.py`)

* **Purpose**: Provide consultation and recommendations
* **Functions**: `handoff_to_next_agent`, `end_conversation`
* **Context**: Receives summary from qualifier

### 3. Closer Agent (`agents/closer/config.py`)

* **Purpose**: Schedule follow-up and gather feedback
* **Functions**: `schedule_followup`, `record_satisfaction`, `end_conversation`
* **Context**: Receives summary from advisor

## Implementation Details

### How Agent Transitions Work

When an agent calls `handoff_to_next_agent`, the orchestrator:

1. **Summarizes** the current conversation using Groq AI
2. **Closes** the current Voice Agent session (but keeps audio task running)
3. **Starts** a new Voice Agent session with the summarized context
4. **Continues** audio forwarding to the new agent seamlessly

### Creating an Agent Configuration

Each agent is configured with settings for STT, LLM, TTS, and functions:

**`Python`**

```python Python
from deepgram.agent.v1.types import AgentV1Settings

def get_qualifier_config(context: str = "") -> AgentV1Settings:
    return AgentV1Settings(
        audio=AgentV1SettingsAudio(...),  # Audio encoding settings
        agent=AgentV1SettingsAgent(
            listen=AgentV1SettingsAgentListen(...),   # Deepgram Flux STT
            think=ThinkSettingsV1(           # LLM configuration
                model="gpt-4o-mini",
                prompt=QUALIFIER_PROMPT,
                functions=QUALIFIER_FUNCTIONS
            ),
            speak=SpeakSettingsV1(...),  # Deepgram Aura TTS
            greeting="Hi, this is Alex..."
        )
    )
```

**`Java`**

```java Java
import com.deepgram.resources.agent.v1.types.*;

public static AgentV1Settings getQualifierConfig(String context) {
    return AgentV1Settings.builder()
        .audio(AgentV1SettingsAudio.builder()
            .input(AgentV1SettingsAudioInput.builder()
                .encoding("linear16").sampleRate(48000).build())
            .output(AgentV1SettingsAudioOutput.builder()
                .encoding("linear16").sampleRate(16000).container("none").build())
            .build())
        .agent(AgentV1SettingsAgent.builder()
            .listen(AgentV1SettingsAgentListen.builder()
                .provider(AgentV1SettingsAgentListenProvider.builder()
                    .type("deepgram").model("nova-3").build())
                .build())
            .think(AgentV1SettingsAgentThink.builder()
                .provider(AgentV1SettingsAgentThinkProvider.builder()
                    .type("open_ai").model("gpt-4o-mini").build())
                .prompt(QUALIFIER_PROMPT)
                .functions(QUALIFIER_FUNCTIONS)
                .build())
            .speak(AgentV1SettingsAgentSpeak.builder()
                .provider(AgentV1SettingsAgentSpeakProvider.builder()
                    .type("deepgram").model("aura-2-thalia-en").build())
                .build())
            .greeting("Hi, this is Alex...")
            .build())
        .build();
}
```

### Function Definitions

Functions include detailed descriptions to guide the LLM's behavior. See `agents/shared/functions.py` for examples. For comprehensive function definition best practices, refer to [docs/FUNCTION\_GUIDE.md](https://github.com/deepgram-devs/deepgram-voice-agent-multi-agent/blob/main/docs/FUNCTION_GUIDE.md).

Key pattern: Functions should wait for customer confirmation:

**`Python`**

```python Python
HANDOFF_FUNCTION = AgentV1Function(
    name="handoff_to_next_agent",
    description="""...
    CORRECT PATTERN:
    1. You ask: "I can connect you with [next agent]. Would that work?"
    2. WAIT for customer response
    3. Customer says: "Yes"
    4. You IMMEDIATELY call this function WITHOUT additional text
    """
)
```

**`Java`**

```java Java
AgentV1Function handoffFunction = AgentV1Function.builder()
    .name("handoff_to_next_agent")
    .description("""
        CORRECT PATTERN:
        1. You ask: "I can connect you with [next agent]. Would that work?"
        2. WAIT for customer response
        3. Customer says: "Yes"
        4. You IMMEDIATELY call this function WITHOUT additional text
        """)
    .parameters(Map.of(
        "type", "object",
        "properties", Map.of()))
    .build();
```

## Quick Start

### Prerequisites

* Python 3.8+
* Twilio account with phone number (with outbound calling enabled)
* Deepgram API key
* Groq API key (free tier available at [https://console.groq.com/](https://console.groq.com/))
* Public tunnel (ngrok, zrok, etc.)

### 1. Install Dependencies

This implementation uses the Deepgram Python SDK:

```bash
# Create virtual environment
python -m venv venv
source venv/bin/activate  # On Windows: venv\Scripts\activate

# Install packages
pip install -r requirements.txt
```

### 2. Configure Environment

Copy `.env.example` to `.env` and add your API keys:

```bash
cp .env.example .env
```

The key environment variables (`config.py` manages these):

* `DEEPGRAM_API_KEY` - Your Deepgram API key
* `GROQ_API_KEY` - Groq API key for conversation summarization
* `TWILIO_ACCOUNT_SID`, `TWILIO_AUTH_TOKEN`, `TWILIO_PHONE_NUMBER` - Twilio credentials
* `LEAD_SERVER_EXTERNAL_URL` - Your public tunnel URL
* `LEAD_PHONE_NUMBER` - Phone number to call

### 3. Start Public Tunnel

```bash
# Using zrok (free, easy)
zrok share public localhost:8000

# Or ngrok
ngrok http 8000 --scheme=ws
```

Copy the public URL and update `LEAD_SERVER_EXTERNAL_URL` in your `.env` file.

### 4. Run the System

```bash
python main.py
```

The system will:

1. Start a WebSocket server on port 8000
2. Initiate an outbound call to `LEAD_PHONE_NUMBER`
3. Begin the conversation with the Qualifier agent

## Example Conversation Flow

Here's a simplified flow showing agent transitions and function calls:

### Phase 1: Qualifier Agent

```
Agent: "Hi, this is Alex calling from our advisory services.
        Is now a good time to chat briefly?"
Customer: "Sure, I have a few minutes"
Agent: "Great! May I get your name?"
Customer: "John Smith"
Agent: "Where are you located?"
Customer: "Seattle"
Agent: "What brings you to call us today?"
Customer: "I need help with retirement planning"
Agent: "I can connect you with one of our advisors. Would that work?"
Customer: "Yes, that'd be great"
[Agent calls handoff_to_next_agent function]
```

### Transition

* Groq summarizes: "John Smith from Seattle needs retirement planning advice"
* Qualifier session closes
* Advisor session starts with context

### Phase 2: Advisor Agent

```
Agent: "Hi John! I understand you're interested in retirement planning.
        How can I help?"
Customer: "I'm 45 and want to retire by 60"
Agent: "Based on your timeline, I recommend a formal consultation.
        Can I connect you with our team to schedule that?"
Customer: "Yes please"
[Agent calls handoff_to_next_agent function]
```

### Transition

* Groq summarizes: "John Smith from Seattle would like to schedule a formal consultation"
* Advisor session closes
* Closer session starts with context

### Phase 3: Closer Agent

```
Agent: "Thanks for speaking with our advisor.
        When would work best for your consultation?"
Customer: "Next Wednesday afternoon"
[Agent calls schedule_followup function]
Agent: "Perfect! How would you rate your experience today from 1 to 5?"
Customer: "5"
[Agent calls record_satisfaction function]
Agent: "Thank you! Have a great day!"
[Agent calls end_conversation function]
```

## Project Structure

```
deepgram-voice-agent-multi-agent-/
├── main.py                      # Entry point - starts server, initiates calls
├── config.py                    # Environment variable management
│
├── orchestrator/
│   └── call_orchestrator.py    # Core orchestration logic
│
├── agents/
│   ├── shared/
│   │   └── functions.py        # Shared function definitions
│   ├── qualifier/               # Qualifier agent configuration
│   ├── advisor/                 # Advisor agent configuration
│   └── closer/                  # Closer agent configuration
│
├── utils/
│   └── context_summarizer.py   # Groq LLM summarization
│
├── call_handling/
│   └── twilio_client.py        # Twilio API wrapper
│
└── docs/
    ├── PROMPT_GUIDE.md          # Voice agent prompt best practices
    └── FUNCTION_GUIDE.md        # Function definition best practices
```

## Customization

### Modify Agent Behavior

Edit the prompts in each agent's config file:

**Qualifier** (`agents/qualifier/config.py`):

* Change greeting message
* Adjust information gathering flow
* Modify qualification criteria

**Advisor** (`agents/advisor/config.py`):

* Customize consultation approach
* Change expertise area (retirement, investments, etc.)
* Adjust handoff triggers

**Closer** (`agents/closer/config.py`):

* Modify scheduling questions
* Change satisfaction survey format
* Customize closing message

**Important**: See `docs/PROMPT_GUIDE.md` for voice-specific prompt engineering best practices.

### Add New Agents

Quick overview:

1. Create `agents/your_agent/config.py` with agent configuration
2. Update `orchestrator/call_orchestrator.py` transition logic
3. Add agent-specific function handlers if needed

### Change Summarization

Edit `utils/context_summarizer.py` to:

* Switch LLM models (currently Llama 3.3 70B on Groq). See [other available Groq models](https://console.groq.com/docs/models).
* Modify summarization prompts for better context extraction
* Change which data points are captured

## Additional Resources

* **[docs/PROMPT\_GUIDE.md](https://github.com/deepgram-devs/deepgram-voice-agent-multi-agent/blob/main/docs/PROMPT_GUIDE.md)** - Voice agent prompt engineering best practices
* **[docs/FUNCTION\_GUIDE.md](https://github.com/deepgram-devs/deepgram-voice-agent-multi-agent/blob/main/docs/FUNCTION_GUIDE.md)** - Function definition best practices