멀티 컨텍스트 WebSocket

이 가이드에서는 멀티 컨텍스트 WebSocket API를 사용해 실시간 음성 에이전트를 구축하는 방법을 안내합니다.

고급

이 멀티 컨텍스트 WebSocket API를 사용한 음성 에이전트 오케스트레이션은 고급 개발자에게 권장되는 복잡한 작업입니다. 더 관리하기 쉬운 솔루션이 필요하다면 이러한 과제를 상당수 간소화하는 Agents Platform 제품을 살펴보세요.

개요

반응성이 뛰어난 음성 에이전트를 구축하려면 오디오 스트림을 동적으로 관리하고, 중단을 자연스럽게 처리하며, 대화 차례 전반에 걸쳐 자연스러운 음성을 유지할 수 있어야 합니다. 텍스트 음성 변환(TTS)을 위한 멀티 컨텍스트 WebSocket API는 이러한 시나리오를 위해 특별히 설계되었습니다.

이 API는 “컨텍스트” 개념을 도입하여 표준 TTS WebSocket 기능을 확장합니다. 각 컨텍스트는 단일 WebSocket 연결 내에서 독립적인 오디오 생성 스트림으로 작동합니다. 이를 통해 다음을 수행할 수 있습니다.

  • 여러 음성 라인을 동시에 관리합니다(예: 사용자 중단에 대한 응답을 준비하면서 에이전트가 말하는 경우).
  • 기존 음성 컨텍스트를 닫고 새 컨텍스트를 시작하여 사용자 끼어들기를 원활하게 처리합니다.
  • 동일한 논리적 컨텍스트 내 발화의 운율적 일관성을 유지합니다.
  • 더 이상 필요하지 않은 컨텍스트를 선택적으로 닫아 리소스 사용을 최적화합니다.

멀티 컨텍스트 WebSocket API는 음성 애플리케이션에 최적화되어 있으며, 서로 관련 없는 여러 오디오 스트림을 동시에 생성하는 용도가 아닙니다. 이를 반영해 각 연결은 동시 컨텍스트 5개로 제한됩니다.

이 가이드에서는 멀티 컨텍스트 WebSocket 연결, 컨텍스트 관리, 그리고 매력적인 음성 에이전트 구축을 위한 모범 사례를 안내합니다.

모범 사례

이러한 모범 사례는 멀티 컨텍스트 WebSocket API로 반응성이 뛰어나고 효율적인 음성 에이전트를 구축하는 데 필수적입니다.

1

단일 WebSocket 연결 사용

최종 사용자 세션마다 하나의 WebSocket 연결을 설정하세요. 여러 연결을 만드는 것보다 오버헤드와 지연 시간을 줄일 수 있습니다. 이 단일 연결 안에서 대화의 서로 다른 부분에 여러 컨텍스트를 관리할 수 있습니다.

2

응답을 청크로 스트리밍하고 문장 단위로 생성

긴 응답을 생성할 때는 텍스트를 더 작은 청크로 스트리밍하고 완전한 문장이 끝날 때 flush: true 플래그를 사용하세요. 이렇게 하면 생성된 오디오의 품질과 응답성이 향상됩니다.

3

중단을 자연스럽게 처리

중단이 발생할 때까지 한 컨텍스트에 텍스트를 스트리밍한 다음 새 컨텍스트를 만들고 기존 컨텍스트를 닫으세요. 이 방식은 대화 흐름이 바뀔 때 부드러운 전환을 보장합니다.

4

컨텍스트 수명 주기 관리

사용하지 않는 컨텍스트는 즉시 닫으세요. 서버는 연결당 최대 5개의 동시 컨텍스트를 유지할 수 있지만, 더 이상 필요하지 않은 컨텍스트는 닫아야 합니다.

5

컨텍스트 타임아웃 방지

컨텍스트는 기본적으로 20초 후 타임아웃되어 자동으로 닫힙니다. 비활성 타임아웃은 모든 컨텍스트에 적용되는 WebSocket 수준 매개변수이며, 필요한 경우 최대 180초까지 설정할 수 있습니다. 타임아웃 시간을 재설정하려면 컨텍스트에 빈 텍스트 메시지를 전송하세요.

중단 처리

사용자가 에이전트를 중단하면 현재 컨텍스트를 닫고 새 컨텍스트를 만들어야 합니다.

async def handle_interruption(websocket, old_context_id, new_context_id, new_response):
# Close the existing context that was interrupted
await websocket.send(json.dumps({
"context_id": old_context_id,
"close_context": True
}))
print(f"Closed interrupted context '{old_context_id}'")
# Create a new context for the new response
await send_text_in_context(websocket, new_response, new_context_id)

컨텍스트 유지하기

컨텍스트는 기본 비활성 시간 20초 후 자동으로 타임아웃됩니다. 텍스트를 생성하지 않고 컨텍스트를 유지해야 하는 경우(예: 처리 지연 중)에는 빈 텍스트 메시지를 전송하여 타임아웃 시간을 재설정할 수 있습니다.

async def keep_context_alive(websocket, context_id):
await websocket.send(json.dumps({
"context_id": context_id,
"text": ""
}))

WebSocket 연결 닫기

대화가 끝나면 소켓을 닫아 모든 컨텍스트를 정리할 수 있습니다.

async def end_conversation(websocket):
# This will close all contexts and close the connection
await websocket.send(json.dumps({
"close_socket": True
}))
print("Ending conversation and closing WebSocket")`

전체 대화형 에이전트 예제

요구 사항

설정

선택한 언어에 필요한 종속성을 설치하세요.

pip install python-dotenv websockets

API 키를 저장할 .env 파일을 프로젝트 디렉터리에 만드세요.

.env
ELEVENLABS_API_KEY=your_elevenlabs_api_key_here

음성 에이전트 예제

이 코드는 예시로 제공되며 프로덕션 용도로 사용하도록 설계되지 않았습니다.
import os
import json
import asyncio
import websockets
from dotenv import load_dotenv
load_dotenv()
ELEVENLABS_API_KEY = os.getenv("ELEVENLABS_API_KEY")
VOICE_ID = "your_voice_id"
MODEL_ID = "eleven_flash_v2_5"
WEBSOCKET_URI = f"wss://api.elevenlabs.io/v1/text-to-speech/{VOICE_ID}/multi-stream-input?model_id={MODEL_ID}"
async def send_text_in_context(websocket, text, context_id, voice_settings=None):
"""Send text to be synthesized in the specified context."""
message = {
"text": text,
"context_id": context_id,
}
# Only include voice_settings for the first message in a context
if voice_settings:
message["voice_settings"] = voice_settings
await websocket.send(json.dumps(message))
async def continue_context(websocket, text, context_id):
"""Add more text to an existing context."""
await websocket.send(json.dumps({
"text": text,
"context_id": context_id
}))
async def flush_context(websocket, context_id):
"""Force generation of any buffered audio in the context."""
await websocket.send(json.dumps({
"context_id": context_id,
"flush": True
}))
async def handle_interruption(websocket, old_context_id, new_context_id, new_response):
"""Handle user interruption by closing current context and starting a new one."""
# Close the existing context that was interrupted
await websocket.send(json.dumps({
"context_id": old_context_id,
"close_context": True
}))
# Create a new context for the new response
await send_text_in_context(websocket, new_response, new_context_id)
async def end_conversation(websocket):
"""End the conversation and close the WebSocket connection."""
await websocket.send(json.dumps({
"close_socket": True
}))
async def receive_messages(websocket):
"""Process incoming WebSocket messages."""
context_audio = {}
try:
async for message in websocket:
data = json.loads(message)
context_id = data.get("contextId", "default")
if data.get("audio"):
print(f"Received audio for context '{context_id}'")
if data.get("is_final"):
print(f"Context '{context_id}' completed")
except (websockets.exceptions.ConnectionClosed, asyncio.CancelledError):
print("Message receiving stopped")
async def conversation_agent_demo():
"""Run a complete conversational agent demo."""
# Connect with API key in headers
async with websockets.connect(
WEBSOCKET_URI,
max_size=16 * 1024 * 1024,
additional_headers={"xi-api-key": ELEVENLABS_API_KEY}
) as websocket:
# Start receiving messages in background
receive_task = asyncio.create_task(receive_messages(websocket))
# Initial agent response
await send_text_in_context(
websocket,
"Hello! I'm your virtual assistant. I can help you with a wide range of topics. What would you like to know about today?",
"greeting"
)
# Wait a bit (simulating user listening)
await asyncio.sleep(2)
# Simulate user interruption
print("USER INTERRUPTS: 'Can you tell me about the weather?'")
# Handle the interruption by closing current context and starting new one
await handle_interruption(
websocket,
"greeting",
"weather_response",
"I'd be happy to tell you about the weather. Currently in your area, it's 72 degrees and sunny with a slight chance of rain later this afternoon."
)
# Add more to the weather context
await continue_context(
websocket,
" If you're planning to go outside, you might want to bring a light jacket just in case.",
"weather_response"
)
# Flush at the end of this turn to ensure all audio is generated
await flush_context(websocket, "weather_response")
# Wait a bit (simulating user listening)
await asyncio.sleep(3)
# Simulate user asking another question
print("USER: 'What about tomorrow?'")
# Create a new context for this response
await send_text_in_context(
websocket,
"Tomorrow's forecast shows temperatures around 75 degrees with partly cloudy skies. It should be a beautiful day overall!",
"tomorrow_weather"
)
# Flush and close this context
await flush_context(websocket, "tomorrow_weather")
await websocket.send(json.dumps({
"context_id": "tomorrow_weather",
"close_context": True
}))
# End the conversation
await asyncio.sleep(2)
await end_conversation(websocket)
# Cancel the receive task
receive_task.cancel()
try:
await receive_task
except asyncio.CancelledError:
pass
if __name__ == "__main__":
asyncio.run(conversation_agent_demo())

다음 단계