Hoppa till navigering

WebSocket med flera kontexter

Den här guiden visar hur du bygger röstagenter i realtid med WebSocket API:t för flera kontexter.

Avancerat

Att orkestrera röstagenter med detta WebSocket API för flera kontexter är en komplex uppgift som rekommenderas för avancerade utvecklare. För en mer hanterad lösning kan du utforska vår produktplattform för agenter, som förenklar många av dessa utmaningar.

Översikt

Att bygga responsiva röstagenter kräver att du kan hantera ljudströmmar dynamiskt, hantera avbrott smidigt och behålla naturligt tal mellan samtalets turer. Vårt WebSocket API för flera kontexter för Text to Speech (TTS) är särskilt utformat för dessa scenarier.

Detta API utökar vår vanliga TTS WebSocket-funktionalitet genom att introducera konceptet ”kontexter”. Varje kontext fungerar som en oberoende ljudgenereringsström inom en och samma WebSocket-anslutning. Det gör att du kan:

  • Hantera flera talrader samtidigt (t.ex. en agent som talar medan den förbereder ett svar på ett avbrott från användaren).
  • Smidigt hantera avbrott från användare genom att stänga en befintlig talkontext och starta en ny.
  • Bibehålla prosodisk konsekvens för yttranden inom samma logiska kontext.
  • Optimera resursanvändningen genom att selektivt stänga kontexter som inte längre behövs.

WebSocket API:t för flera kontexter är optimerat för röstapplikationer och är inte avsett för att generera flera orelaterade ljudströmmar samtidigt. Varje anslutning är därför begränsad till 5 samtidiga kontexter.

Den här guiden går igenom hur du ansluter till WebSocket med flera kontexter, hanterar kontexter och tillämpar bästa praxis för att bygga engagerande röstagenter.

Bästa praxis

Den här bästa praxisen är viktig för att bygga responsiva och effektiva röstagenter med vårt WebSocket API för flera kontexter.

1

Använd en enda WebSocket-anslutning

Upprätta en WebSocket-anslutning för varje slutanvändarsession. Det minskar overhead och latens jämfört med att skapa flera anslutningar. I denna enda anslutning kan du hantera flera kontexter för olika delar av samtalet.

2

Streama svar i delar och generera meningar

När du genererar långa svar ska du streama texten i mindre delar och använda flaggan flush: true i slutet av fullständiga meningar. Det förbättrar kvaliteten på det genererade ljudet och ökar responsiviteten.

3

Hantera avbrott smidigt

Streama text till en kontext tills ett avbrott sker, skapa sedan en ny kontext och stäng den befintliga. Detta säkerställer smidiga övergångar när samtalets flöde ändras.

4

Hantera kontextens livscykel

Stäng oanvända kontexter direkt. Servern kan ha upp till 5 samtidiga kontexter per anslutning, men du bör stänga kontexter när de inte längre behövs.

5

Förhindra tidsgränser för kontexter

Kontexter får som standard en tidsgräns efter 20 sekunder och stängs automatiskt. Tidsgränsen för inaktivitet är en parameter på WebSocket-nivå som gäller alla kontexter och kan vara upp till 180 sekunder vid behov. Skicka ett tomt textmeddelande till en kontext för att återställa tidsgränsen.

Hantera avbrott

När en användare avbryter din agent bör du stänga den aktuella kontexten och skapa en ny:

async def handle_interruption(websocket, old_context_id, new_context_id, new_response):
# Close the existing context that was interrupted
await websocket.send(json.dumps({
"context_id": old_context_id,
"close_context": True
}))
print(f"Closed interrupted context '{old_context_id}'")
# Create a new context for the new response
await send_text_in_context(websocket, new_response, new_context_id)

Håll en kontext aktiv

Kontexter får automatiskt en tidsgräns efter 20 sekunders inaktivitet som standard. Om du behöver hålla en kontext aktiv utan att generera text (till exempel under en bearbetningsfördröjning) kan du skicka ett tomt textmeddelande för att återställa tidsgränsen.

async def keep_context_alive(websocket, context_id):
await websocket.send(json.dumps({
"context_id": context_id,
"text": ""
}))

Stäng WebSocket-anslutningen

När samtalet är slut kan du rensa alla kontexter genom att stänga socketen:

async def end_conversation(websocket):
# This will close all contexts and close the connection
await websocket.send(json.dumps({
"close_socket": True
}))
print("Ending conversation and closing WebSocket")`

Komplett exempel på en samtalsagent

Krav

  • Ett ElevenLabs-konto med en API-nyckel (lär dig hur du hittar din API-nyckel).
  • Python eller Node.js (eller en annan JavaScript-runtime) installerat på din dator.
  • Grundläggande kunskap om WebSocket-kommunikation. Vi rekommenderar att du läser vår guide om vanlig WebSocket-streaming för grundläggande koncept.

Konfiguration

Installera de nödvändiga beroendena för det språk du valt:

pip install python-dotenv websockets

Skapa en .env-fil i din projektkatalog för att lagra din API-nyckel:

.env
ELEVENLABS_API_KEY=your_elevenlabs_api_key_here

Exempel på röstagent

Den här koden är ett exempel och är inte avsedd för användning i produktion
import os
import json
import asyncio
import websockets
from dotenv import load_dotenv
load_dotenv()
ELEVENLABS_API_KEY = os.getenv("ELEVENLABS_API_KEY")
VOICE_ID = "your_voice_id"
MODEL_ID = "eleven_flash_v2_5"
WEBSOCKET_URI = f"wss://api.elevenlabs.io/v1/text-to-speech/{VOICE_ID}/multi-stream-input?model_id={MODEL_ID}"
async def send_text_in_context(websocket, text, context_id, voice_settings=None):
"""Send text to be synthesized in the specified context."""
message = {
"text": text,
"context_id": context_id,
}
# Only include voice_settings for the first message in a context
if voice_settings:
message["voice_settings"] = voice_settings
await websocket.send(json.dumps(message))
async def continue_context(websocket, text, context_id):
"""Add more text to an existing context."""
await websocket.send(json.dumps({
"text": text,
"context_id": context_id
}))
async def flush_context(websocket, context_id):
"""Force generation of any buffered audio in the context."""
await websocket.send(json.dumps({
"context_id": context_id,
"flush": True
}))
async def handle_interruption(websocket, old_context_id, new_context_id, new_response):
"""Handle user interruption by closing current context and starting a new one."""
# Close the existing context that was interrupted
await websocket.send(json.dumps({
"context_id": old_context_id,
"close_context": True
}))
# Create a new context for the new response
await send_text_in_context(websocket, new_response, new_context_id)
async def end_conversation(websocket):
"""End the conversation and close the WebSocket connection."""
await websocket.send(json.dumps({
"close_socket": True
}))
async def receive_messages(websocket):
"""Process incoming WebSocket messages."""
context_audio = {}
try:
async for message in websocket:
data = json.loads(message)
context_id = data.get("contextId", "default")
if data.get("audio"):
print(f"Received audio for context '{context_id}'")
if data.get("is_final"):
print(f"Context '{context_id}' completed")
except (websockets.exceptions.ConnectionClosed, asyncio.CancelledError):
print("Message receiving stopped")
async def conversation_agent_demo():
"""Run a complete conversational agent demo."""
# Connect with API key in headers
async with websockets.connect(
WEBSOCKET_URI,
max_size=16 * 1024 * 1024,
additional_headers={"xi-api-key": ELEVENLABS_API_KEY}
) as websocket:
# Start receiving messages in background
receive_task = asyncio.create_task(receive_messages(websocket))
# Initial agent response
await send_text_in_context(
websocket,
"Hello! I'm your virtual assistant. I can help you with a wide range of topics. What would you like to know about today?",
"greeting"
)
# Wait a bit (simulating user listening)
await asyncio.sleep(2)
# Simulate user interruption
print("USER INTERRUPTS: 'Can you tell me about the weather?'")
# Handle the interruption by closing current context and starting new one
await handle_interruption(
websocket,
"greeting",
"weather_response",
"I'd be happy to tell you about the weather. Currently in your area, it's 72 degrees and sunny with a slight chance of rain later this afternoon."
)
# Add more to the weather context
await continue_context(
websocket,
" If you're planning to go outside, you might want to bring a light jacket just in case.",
"weather_response"
)
# Flush at the end of this turn to ensure all audio is generated
await flush_context(websocket, "weather_response")
# Wait a bit (simulating user listening)
await asyncio.sleep(3)
# Simulate user asking another question
print("USER: 'What about tomorrow?'")
# Create a new context for this response
await send_text_in_context(
websocket,
"Tomorrow's forecast shows temperatures around 75 degrees with partly cloudy skies. It should be a beautiful day overall!",
"tomorrow_weather"
)
# Flush and close this context
await flush_context(websocket, "tomorrow_weather")
await websocket.send(json.dumps({
"context_id": "tomorrow_weather",
"close_context": True
}))
# End the conversation
await asyncio.sleep(2)
await end_conversation(websocket)
# Cancel the receive task
receive_task.cancel()
try:
await receive_task
except asyncio.CancelledError:
pass
if __name__ == "__main__":
asyncio.run(conversation_agent_demo())

Nästa steg