WebSocket com múltiplos contextos

Este guia mostra como criar agentes de voz em tempo real usando a API WebSocket com múltiplos contextos.

Avançado

Orquestrar agentes de voz usando esta API WebSocket com múltiplos contextos é uma tarefa complexa recomendada para desenvolvedores avançados. Para uma solução mais gerenciada, considere conhecer nosso produto da plataforma de Agents, que simplifica muitos desses desafios.

Visão geral

Criar agentes de voz responsivos exige a capacidade de gerenciar fluxos de áudio dinamicamente, lidar bem com interrupções e manter uma fala com som natural ao longo dos turnos da conversa. Nossa API WebSocket com múltiplos contextos para Text to Speech (TTS) foi desenvolvida especificamente para esses cenários.

Esta API expande nossa funcionalidade padrão de WebSocket para TTS ao introduzir o conceito de “contextos”. Cada contexto funciona como um fluxo independente de geração de áudio em uma única conexão WebSocket. Isso permite que você:

  • Gerencie várias linhas de fala simultaneamente (por exemplo, o agente falando enquanto prepara uma resposta para uma interrupção do usuário).
  • Lide sem interrupções com intervenções do usuário fechando um contexto de fala existente e iniciando um novo.
  • Mantenha a consistência prosódica de enunciados dentro do mesmo contexto lógico.
  • Otimize o uso de recursos fechando seletivamente contextos que não são mais necessários.

A API WebSocket com múltiplos contextos é otimizada para aplicações de voz e não se destina à geração simultânea de vários fluxos de áudio não relacionados. Cada conexão é limitada a 5 contextos simultâneos para refletir isso.

Este guia mostrará como se conectar ao WebSocket com múltiplos contextos, gerenciar contextos e aplicar boas práticas para criar agentes de voz envolventes.

Boas práticas

Estas boas práticas são essenciais para criar agentes de voz responsivos e eficientes com nossa API WebSocket com múltiplos contextos.

1

Use uma única conexão WebSocket

Estabeleça uma conexão WebSocket para cada sessão de usuário final. Isso reduz a sobrecarga e a latência em comparação com a criação de várias conexões. Dentro dessa única conexão, você pode gerenciar vários contextos para diferentes partes da conversa.

2

Transmita respostas em partes, gere frases

Ao gerar respostas longas, transmita o texto em partes menores e use a flag flush: true ao final de frases completas. Isso melhora a qualidade do áudio gerado e a responsividade.

3

Lide bem com interrupções

Transmita texto para um contexto até ocorrer uma interrupção, depois crie um novo contexto e feche o existente. Essa abordagem garante transições suaves quando o fluxo da conversa muda.

4

Gerencie o ciclo de vida dos contextos

Feche contextos não utilizados rapidamente. O servidor pode manter até 5 contextos simultâneos por conexão, mas você deve fechar os contextos quando não forem mais necessários.

5

Evite tempos limite de contexto

Por padrão, os contextos expiram após 20 segundos e são fechados automaticamente. O tempo limite de inatividade é um parâmetro no nível do websocket que se aplica a todos os contextos e pode chegar a 180 segundos se necessário. Envie uma mensagem de texto vazia em um contexto para reiniciar o contador do tempo limite.

Como lidar com interrupções

Quando um usuário interromper seu agente, você deve fechar o contexto atual e criar um novo:

async def handle_interruption(websocket, old_context_id, new_context_id, new_response):
# Close the existing context that was interrupted
await websocket.send(json.dumps({
"context_id": old_context_id,
"close_context": True
}))
print(f"Closed interrupted context '{old_context_id}'")
# Create a new context for the new response
await send_text_in_context(websocket, new_response, new_context_id)

Como manter um contexto ativo

Os contextos expiram automaticamente após 20 segundos de inatividade por padrão. Se precisar manter um contexto ativo sem gerar texto (por exemplo, durante um atraso de processamento), você pode enviar uma mensagem de texto vazia para reiniciar o contador do tempo limite.

async def keep_context_alive(websocket, context_id):
await websocket.send(json.dumps({
"context_id": context_id,
"text": ""
}))

Como fechar a conexão WebSocket

Quando sua conversa terminar, você poderá limpar todos os contextos fechando o socket:

async def end_conversation(websocket):
# This will close all contexts and close the connection
await websocket.send(json.dumps({
"close_socket": True
}))
print("Ending conversation and closing WebSocket")`

Exemplo completo de agente conversacional

Requisitos

Configuração

Instale as dependências necessárias para a linguagem escolhida:

pip install python-dotenv websockets

Crie um arquivo .env no diretório do seu projeto para armazenar sua chave de API:

.env
ELEVENLABS_API_KEY=your_elevenlabs_api_key_here

Exemplo de agente de voz

Este código é fornecido como exemplo e não se destina ao uso em produção
import os
import json
import asyncio
import websockets
from dotenv import load_dotenv
load_dotenv()
ELEVENLABS_API_KEY = os.getenv("ELEVENLABS_API_KEY")
VOICE_ID = "your_voice_id"
MODEL_ID = "eleven_flash_v2_5"
WEBSOCKET_URI = f"wss://api.elevenlabs.io/v1/text-to-speech/{VOICE_ID}/multi-stream-input?model_id={MODEL_ID}"
async def send_text_in_context(websocket, text, context_id, voice_settings=None):
"""Send text to be synthesized in the specified context."""
message = {
"text": text,
"context_id": context_id,
}
# Only include voice_settings for the first message in a context
if voice_settings:
message["voice_settings"] = voice_settings
await websocket.send(json.dumps(message))
async def continue_context(websocket, text, context_id):
"""Add more text to an existing context."""
await websocket.send(json.dumps({
"text": text,
"context_id": context_id
}))
async def flush_context(websocket, context_id):
"""Force generation of any buffered audio in the context."""
await websocket.send(json.dumps({
"context_id": context_id,
"flush": True
}))
async def handle_interruption(websocket, old_context_id, new_context_id, new_response):
"""Handle user interruption by closing current context and starting a new one."""
# Close the existing context that was interrupted
await websocket.send(json.dumps({
"context_id": old_context_id,
"close_context": True
}))
# Create a new context for the new response
await send_text_in_context(websocket, new_response, new_context_id)
async def end_conversation(websocket):
"""End the conversation and close the WebSocket connection."""
await websocket.send(json.dumps({
"close_socket": True
}))
async def receive_messages(websocket):
"""Process incoming WebSocket messages."""
context_audio = {}
try:
async for message in websocket:
data = json.loads(message)
context_id = data.get("contextId", "default")
if data.get("audio"):
print(f"Received audio for context '{context_id}'")
if data.get("is_final"):
print(f"Context '{context_id}' completed")
except (websockets.exceptions.ConnectionClosed, asyncio.CancelledError):
print("Message receiving stopped")
async def conversation_agent_demo():
"""Run a complete conversational agent demo."""
# Connect with API key in headers
async with websockets.connect(
WEBSOCKET_URI,
max_size=16 * 1024 * 1024,
additional_headers={"xi-api-key": ELEVENLABS_API_KEY}
) as websocket:
# Start receiving messages in background
receive_task = asyncio.create_task(receive_messages(websocket))
# Initial agent response
await send_text_in_context(
websocket,
"Hello! I'm your virtual assistant. I can help you with a wide range of topics. What would you like to know about today?",
"greeting"
)
# Wait a bit (simulating user listening)
await asyncio.sleep(2)
# Simulate user interruption
print("USER INTERRUPTS: 'Can you tell me about the weather?'")
# Handle the interruption by closing current context and starting new one
await handle_interruption(
websocket,
"greeting",
"weather_response",
"I'd be happy to tell you about the weather. Currently in your area, it's 72 degrees and sunny with a slight chance of rain later this afternoon."
)
# Add more to the weather context
await continue_context(
websocket,
" If you're planning to go outside, you might want to bring a light jacket just in case.",
"weather_response"
)
# Flush at the end of this turn to ensure all audio is generated
await flush_context(websocket, "weather_response")
# Wait a bit (simulating user listening)
await asyncio.sleep(3)
# Simulate user asking another question
print("USER: 'What about tomorrow?'")
# Create a new context for this response
await send_text_in_context(
websocket,
"Tomorrow's forecast shows temperatures around 75 degrees with partly cloudy skies. It should be a beautiful day overall!",
"tomorrow_weather"
)
# Flush and close this context
await flush_context(websocket, "tomorrow_weather")
await websocket.send(json.dumps({
"context_id": "tomorrow_weather",
"close_context": True
}))
# End the conversation
await asyncio.sleep(2)
await end_conversation(websocket)
# Cancel the receive task
receive_task.cancel()
try:
await receive_task
except asyncio.CancelledError:
pass
if __name__ == "__main__":
asyncio.run(conversation_agent_demo())

Próximas etapas