多上下文 WebSocket

本指南介绍如何使用多上下文 WebSocket API 构建实时语音智能体。

高级功能

使用此多上下文 WebSocket API 编排语音智能体是一项复杂任务,建议由高级开发者使用。如需更易于管理的解决方案,建议了解我们的 Agents 平台产品,它可简化许多此类难题。

概述

构建响应迅速的语音智能体,需要能够动态管理音频流、妥善处理打断,并在多轮对话中保持自然的语音。面向文本转语音(TTS)的多上下文 WebSocket API 专为这些场景设计。

此 API 在标准 TTS WebSocket 功能的基础上引入了“上下文”概念。每个上下文都在单个 WebSocket 连接中作为独立的音频生成流运行。因此你可以:

  • 并行管理多段语音(例如,智能体说话时准备回应用户的打断)。
  • 通过关闭现有语音上下文并启动新上下文,无缝处理用户插话。
  • 保持同一逻辑上下文内话语的韵律一致性。
  • 通过选择性关闭不再需要的上下文,优化资源使用。

多上下文 WebSocket API 专为语音应用优化,并非用于同时生成多个无关的音频流。因此,每个连接最多支持 5 个并发上下文。

本指南将介绍如何连接多上下文 WebSocket、管理上下文,以及应用构建引人入胜的语音智能体的最佳实践。

最佳实践

这些最佳实践对于使用多上下文 WebSocket API 构建响应迅速、高效的语音智能体至关重要。

1

使用单个 WebSocket 连接

为每个最终用户会话建立一个 WebSocket 连接。与创建多个连接相比,这可减少开销和延迟。在此单一连接中,可以管理对话不同部分的多个上下文。

2

分块流式传输响应,按句生成

生成较长响应时,请将文本分成较小的块进行流式传输,并在完整句子结束时使用 flush: true 标记。这可提升生成音频的质量和响应速度。

3

妥善处理打断

将文本流式传输到一个上下文中,直到发生打断;然后创建新上下文并关闭现有上下文。对话流程变化时,此方法可确保平滑过渡。

4

管理上下文生命周期

及时关闭未使用的上下文。服务器每个连接最多可维持 5 个并发上下文,但不再需要时应关闭上下文。

5

防止上下文超时

上下文默认会在 20 秒后超时并自动关闭。非活动超时是适用于所有上下文的 WebSocket 级参数,如有需要最多可设为 180 秒。向某个上下文发送空文本消息即可重置超时计时。

处理打断

当用户打断智能体时,应关闭当前上下文,并创建新上下文:

async def handle_interruption(websocket, old_context_id, new_context_id, new_response):
# Close the existing context that was interrupted
await websocket.send(json.dumps({
"context_id": old_context_id,
"close_context": True
}))
print(f"Closed interrupted context '{old_context_id}'")
# Create a new context for the new response
await send_text_in_context(websocket, new_response, new_context_id)

保持上下文活跃

上下文会在默认 20 秒无活动后自动超时。如果需要在不生成文本的情况下保持上下文活跃(例如处理延迟期间),可以发送空文本消息来重置超时计时。

async def keep_context_alive(websocket, context_id):
await websocket.send(json.dumps({
"context_id": context_id,
"text": ""
}))

关闭 WebSocket 连接

对话结束后,可以通过关闭 socket清理所有上下文:

async def end_conversation(websocket):
# This will close all contexts and close the connection
await websocket.send(json.dumps({
"close_socket": True
}))
print("Ending conversation and closing WebSocket")`

完整对话式智能体示例

要求

设置

为所选语言安装必要的依赖项:

pip install python-dotenv websockets

在项目目录中创建 .env 文件来存储 API 密钥:

.env
ELEVENLABS_API_KEY=your_elevenlabs_api_key_here

语音智能体示例

此代码仅供示例,不适用于生产环境
import os
import json
import asyncio
import websockets
from dotenv import load_dotenv
load_dotenv()
ELEVENLABS_API_KEY = os.getenv("ELEVENLABS_API_KEY")
VOICE_ID = "your_voice_id"
MODEL_ID = "eleven_flash_v2_5"
WEBSOCKET_URI = f"wss://api.elevenlabs.io/v1/text-to-speech/{VOICE_ID}/multi-stream-input?model_id={MODEL_ID}"
async def send_text_in_context(websocket, text, context_id, voice_settings=None):
"""Send text to be synthesized in the specified context."""
message = {
"text": text,
"context_id": context_id,
}
# Only include voice_settings for the first message in a context
if voice_settings:
message["voice_settings"] = voice_settings
await websocket.send(json.dumps(message))
async def continue_context(websocket, text, context_id):
"""Add more text to an existing context."""
await websocket.send(json.dumps({
"text": text,
"context_id": context_id
}))
async def flush_context(websocket, context_id):
"""Force generation of any buffered audio in the context."""
await websocket.send(json.dumps({
"context_id": context_id,
"flush": True
}))
async def handle_interruption(websocket, old_context_id, new_context_id, new_response):
"""Handle user interruption by closing current context and starting a new one."""
# Close the existing context that was interrupted
await websocket.send(json.dumps({
"context_id": old_context_id,
"close_context": True
}))
# Create a new context for the new response
await send_text_in_context(websocket, new_response, new_context_id)
async def end_conversation(websocket):
"""End the conversation and close the WebSocket connection."""
await websocket.send(json.dumps({
"close_socket": True
}))
async def receive_messages(websocket):
"""Process incoming WebSocket messages."""
context_audio = {}
try:
async for message in websocket:
data = json.loads(message)
context_id = data.get("contextId", "default")
if data.get("audio"):
print(f"Received audio for context '{context_id}'")
if data.get("is_final"):
print(f"Context '{context_id}' completed")
except (websockets.exceptions.ConnectionClosed, asyncio.CancelledError):
print("Message receiving stopped")
async def conversation_agent_demo():
"""Run a complete conversational agent demo."""
# Connect with API key in headers
async with websockets.connect(
WEBSOCKET_URI,
max_size=16 * 1024 * 1024,
additional_headers={"xi-api-key": ELEVENLABS_API_KEY}
) as websocket:
# Start receiving messages in background
receive_task = asyncio.create_task(receive_messages(websocket))
# Initial agent response
await send_text_in_context(
websocket,
"Hello! I'm your virtual assistant. I can help you with a wide range of topics. What would you like to know about today?",
"greeting"
)
# Wait a bit (simulating user listening)
await asyncio.sleep(2)
# Simulate user interruption
print("USER INTERRUPTS: 'Can you tell me about the weather?'")
# Handle the interruption by closing current context and starting new one
await handle_interruption(
websocket,
"greeting",
"weather_response",
"I'd be happy to tell you about the weather. Currently in your area, it's 72 degrees and sunny with a slight chance of rain later this afternoon."
)
# Add more to the weather context
await continue_context(
websocket,
" If you're planning to go outside, you might want to bring a light jacket just in case.",
"weather_response"
)
# Flush at the end of this turn to ensure all audio is generated
await flush_context(websocket, "weather_response")
# Wait a bit (simulating user listening)
await asyncio.sleep(3)
# Simulate user asking another question
print("USER: 'What about tomorrow?'")
# Create a new context for this response
await send_text_in_context(
websocket,
"Tomorrow's forecast shows temperatures around 75 degrees with partly cloudy skies. It should be a beautiful day overall!",
"tomorrow_weather"
)
# Flush and close this context
await flush_context(websocket, "tomorrow_weather")
await websocket.send(json.dumps({
"context_id": "tomorrow_weather",
"close_context": True
}))
# End the conversation
await asyncio.sleep(2)
await end_conversation(websocket)
# Cancel the receive task
receive_task.cancel()
try:
await receive_task
except asyncio.CancelledError:
pass
if __name__ == "__main__":
asyncio.run(conversation_agent_demo())

后续步骤