マルチコンテキストWebSocket

このガイドでは、マルチコンテキストWebSocket APIを使用してリアルタイム音声エージェントを構築する方法を紹介します。

上級者向け

このマルチコンテキストWebSocket APIを使った音声エージェントのオーケストレーションは複雑な作業であり、 上級デベロッパーに推奨されます。より管理しやすいソリューションとして、こうした課題の多くを簡素化する Agents Platformプロダクトをご検討ください。

概要

応答性の高い音声エージェントを構築するには、オーディオストリームを動的に管理し、中断を適切に処理し、会話のターンをまたいで自然な音声を維持する必要があります。テキスト読み上げ(TTS)向けのマルチコンテキストWebSocket APIは、こうしたシナリオ専用に設計されています。

このAPIは、「コンテキスト」の概念を導入することで、標準TTS WebSocket機能を拡張します。各コンテキストは、1つのWebSocket接続内で独立したオーディオ生成ストリームとして動作します。これにより、以下が可能になります。

  • 複数の発話行を同時に管理する(例:ユーザーの割り込みへの応答を準備しながらエージェントが発話する)。
  • 既存の発話コンテキストを閉じ、新しいコンテキストを開始することで、ユーザーの割り込みをシームレスに処理する。
  • 同じ論理コンテキスト内の発話でプロソディの一貫性を維持する。
  • 不要になったコンテキストを選択的に閉じて、リソース使用量を最適化する。

マルチコンテキストWebSocket APIは音声アプリケーション向けに最適化されており、無関係な複数のオーディオストリームを同時に生成する用途を想定していません。そのため、各接続は同時に5つのコンテキストまでに制限されています。

このガイドでは、マルチコンテキストWebSocketへの接続、コンテキストの管理、魅力的な音声エージェントを構築するためのベストプラクティスを説明します。

ベストプラクティス

これらのベストプラクティスは、マルチコンテキストWebSocket APIを使って応答性と効率性に優れた音声エージェントを構築するために不可欠です。

1

単一のWebSocket接続を使用する

エンドユーザーセッションごとに1つのWebSocket接続を確立します。複数の接続を作成する場合と比べて、オーバーヘッドとレイテンシーを削減できます。この単一接続内で、会話のさまざまな部分に対応する複数のコンテキストを管理できます。

2

応答をチャンクでストリーミングし、文単位で生成する

長い応答を生成する際は、テキストを小さなチャンクに分けてストリーミングし、完全な文の末尾でflush: trueフラグを使用します。これにより、生成されるオーディオの品質と応答性が向上します。

3

中断を適切に処理する

中断が発生するまで1つのコンテキストにテキストをストリーミングし、その後、新しいコンテキストを作成して既存のコンテキストを閉じます。この方法により、会話の流れが変わってもスムーズに移行できます。

4

コンテキストのライフサイクルを管理する

使用していないコンテキストは速やかに閉じてください。サーバーでは接続ごとに最大5つのコンテキストを同時に維持できますが、不要になったコンテキストは閉じる必要があります。

5

コンテキストのタイムアウトを防ぐ

コンテキストはデフォルトで20秒後にタイムアウトし、自動的に閉じられます。非アクティブタイムアウトはすべてのコンテキストに適用されるWebSocketレベルのパラメーターで、必要に応じて最大180秒まで設定できます。コンテキストに空のテキストメッセージを送信すると、タイムアウトの時計をリセットできます。

中断の処理

ユーザーがエージェントを中断した場合は、現在のコンテキストを閉じ、新しいコンテキストを作成してください。

async def handle_interruption(websocket, old_context_id, new_context_id, new_response):
# Close the existing context that was interrupted
await websocket.send(json.dumps({
"context_id": old_context_id,
"close_context": True
}))
print(f"Closed interrupted context '{old_context_id}'")
# Create a new context for the new response
await send_text_in_context(websocket, new_response, new_context_id)

コンテキストを維持する

コンテキストは、デフォルトで20秒間操作がないと自動的にタイムアウトします。テキストを生成せずにコンテキストを維持する必要がある場合(たとえば処理の遅延中)は、空のテキストメッセージを送信してタイムアウトの時計をリセットできます。

async def keep_context_alive(websocket, context_id):
await websocket.send(json.dumps({
"context_id": context_id,
"text": ""
}))

WebSocket接続を閉じる

会話が終了したら、ソケットを閉じることで、すべてのコンテキストをクリーンアップできます。

async def end_conversation(websocket):
# This will close all contexts and close the connection
await websocket.send(json.dumps({
"close_socket": True
}))
print("Ending conversation and closing WebSocket")`

会話型エージェントの完全な例

要件

セットアップ

選択した言語に必要な依存関係をインストールします。

pip install python-dotenv websockets

プロジェクトディレクトリに.envファイルを作成し、APIキーを保存します。

.env
ELEVENLABS_API_KEY=your_elevenlabs_api_key_here

音声エージェントの例

このコードは例として提供されており、本番環境での使用を目的としていません
import os
import json
import asyncio
import websockets
from dotenv import load_dotenv
load_dotenv()
ELEVENLABS_API_KEY = os.getenv("ELEVENLABS_API_KEY")
VOICE_ID = "your_voice_id"
MODEL_ID = "eleven_flash_v2_5"
WEBSOCKET_URI = f"wss://api.elevenlabs.io/v1/text-to-speech/{VOICE_ID}/multi-stream-input?model_id={MODEL_ID}"
async def send_text_in_context(websocket, text, context_id, voice_settings=None):
"""Send text to be synthesized in the specified context."""
message = {
"text": text,
"context_id": context_id,
}
# Only include voice_settings for the first message in a context
if voice_settings:
message["voice_settings"] = voice_settings
await websocket.send(json.dumps(message))
async def continue_context(websocket, text, context_id):
"""Add more text to an existing context."""
await websocket.send(json.dumps({
"text": text,
"context_id": context_id
}))
async def flush_context(websocket, context_id):
"""Force generation of any buffered audio in the context."""
await websocket.send(json.dumps({
"context_id": context_id,
"flush": True
}))
async def handle_interruption(websocket, old_context_id, new_context_id, new_response):
"""Handle user interruption by closing current context and starting a new one."""
# Close the existing context that was interrupted
await websocket.send(json.dumps({
"context_id": old_context_id,
"close_context": True
}))
# Create a new context for the new response
await send_text_in_context(websocket, new_response, new_context_id)
async def end_conversation(websocket):
"""End the conversation and close the WebSocket connection."""
await websocket.send(json.dumps({
"close_socket": True
}))
async def receive_messages(websocket):
"""Process incoming WebSocket messages."""
context_audio = {}
try:
async for message in websocket:
data = json.loads(message)
context_id = data.get("contextId", "default")
if data.get("audio"):
print(f"Received audio for context '{context_id}'")
if data.get("is_final"):
print(f"Context '{context_id}' completed")
except (websockets.exceptions.ConnectionClosed, asyncio.CancelledError):
print("Message receiving stopped")
async def conversation_agent_demo():
"""Run a complete conversational agent demo."""
# Connect with API key in headers
async with websockets.connect(
WEBSOCKET_URI,
max_size=16 * 1024 * 1024,
additional_headers={"xi-api-key": ELEVENLABS_API_KEY}
) as websocket:
# Start receiving messages in background
receive_task = asyncio.create_task(receive_messages(websocket))
# Initial agent response
await send_text_in_context(
websocket,
"Hello! I'm your virtual assistant. I can help you with a wide range of topics. What would you like to know about today?",
"greeting"
)
# Wait a bit (simulating user listening)
await asyncio.sleep(2)
# Simulate user interruption
print("USER INTERRUPTS: 'Can you tell me about the weather?'")
# Handle the interruption by closing current context and starting new one
await handle_interruption(
websocket,
"greeting",
"weather_response",
"I'd be happy to tell you about the weather. Currently in your area, it's 72 degrees and sunny with a slight chance of rain later this afternoon."
)
# Add more to the weather context
await continue_context(
websocket,
" If you're planning to go outside, you might want to bring a light jacket just in case.",
"weather_response"
)
# Flush at the end of this turn to ensure all audio is generated
await flush_context(websocket, "weather_response")
# Wait a bit (simulating user listening)
await asyncio.sleep(3)
# Simulate user asking another question
print("USER: 'What about tomorrow?'")
# Create a new context for this response
await send_text_in_context(
websocket,
"Tomorrow's forecast shows temperatures around 75 degrees with partly cloudy skies. It should be a beautiful day overall!",
"tomorrow_weather"
)
# Flush and close this context
await flush_context(websocket, "tomorrow_weather")
await websocket.send(json.dumps({
"context_id": "tomorrow_weather",
"close_context": True
}))
# End the conversation
await asyncio.sleep(2)
await end_conversation(websocket)
# Cancel the receive task
receive_task.cancel()
try:
await receive_task
except asyncio.CancelledError:
pass
if __name__ == "__main__":
asyncio.run(conversation_agent_demo())

次のステップ