WebSocket

通过 AI 智能体创建实时互动语音对话

本文档面向直接集成 ElevenLabs WebSocket API 的开发者。为方便起见,建议使用 ElevenLabs 提供的官方 SDK。

ElevenAgents WebSocket API 支持与 AI 智能体进行实时互动语音对话。建立 WebSocket 连接后,你可以实时发送音频输入并接收音频响应,打造逼真的对话体验。

端点:wss://api.elevenlabs.io/v1/convai/conversation?agent_id={agent_id}

身份验证

使用智能体 ID

对于公开智能体,无需额外身份验证,直接在 WebSocket URL 中使用 agent_id:

wss://api.elevenlabs.io/v1/convai/conversation?agent_id=<your-agent-id>

使用签名 URL

对于私有智能体或需要授权的对话,请从服务器获取签名 URL。服务器会使用 API 密钥与 ElevenLabs API 安全通信。

使用 cURL 的示例

请求:

curl -X GET "https://api.elevenlabs.io/v1/convai/conversation/get-signed-url?agent_id=<your-agent-id>" \
-H "xi-api-key: <your-api-key>"

响应:

{
"signed_url": "wss://api.elevenlabs.io/v1/convai/conversation?agent_id=<your-agent-id>&token=<token>"
}
切勿在客户端暴露 ElevenLabs API 密钥。

WebSocket 事件

客户端到服务器事件

可从客户端向服务器发送以下事件:

发送不会中断对话的上下文信息,以更新对话状态。这样可以在不中断当前对话流程的情况下提供额外上下文。

{
"type": "contextual_update",
"text": "User clicked on pricing page"
}

使用场景:

  • 更新用户状态或偏好
  • 提供环境上下文
  • 添加背景信息
  • 跟踪用户界面交互

要点:

  • 不会中断当前对话流程
  • 更新会作为工具调用纳入对话历史
  • 无需打断自然对话即可保持上下文

上下文更新会异步处理,无需服务器直接响应。

Next.js 实现示例

本示例演示如何使用 ElevenLabs WebSocket API 在 Next.js 中实现基于 WebSocket 的对话式智能体客户端。

本示例使用 voice-stream 包处理麦克风输入,但你也可以自行实现音频采集和编码方案。这里重点演示 如何通过 ElevenLabs API 建立 WebSocket 连接和处理事件。

1

安装所需依赖项

首先,安装所需软件包:

npm install voice-stream

voice-stream 包可处理麦克风访问和音频流,并会按 ElevenLabs API 要求自动将音频编码为 base64 格式。

本示例使用 Tailwind CSS 进行样式设计。要将 Tailwind 添加到 Next.js 项目:

npm install -D tailwindcss postcss autoprefixer
npx tailwindcss init -p

然后按照 Next.js 官方 Tailwind CSS 设置指南操作。

或者,你可以用自己的 CSS 样式替换 className 属性。

2

创建 WebSocket 类型

定义 WebSocket 事件类型:

app/types/websocket.ts
type BaseEvent = {
type: string;
};
type UserTranscriptEvent = BaseEvent & {
type: "user_transcript";
user_transcription_event: {
user_transcript: string;
};
};
type AgentResponseEvent = BaseEvent & {
type: "agent_response";
agent_response_event: {
agent_response: string;
};
};
type AgentResponseCorrectionEvent = BaseEvent & {
type: "agent_response_correction";
agent_response_correction_event: {
original_agent_response: string;
corrected_agent_response: string;
};
};
type AudioResponseEvent = BaseEvent & {
type: "audio";
audio_event: {
audio_base_64: string;
event_id: number;
alignment: {
chars: string[];
char_durations_ms: number[];
char_start_times_ms: number[];
};
};
};
type InterruptionEvent = BaseEvent & {
type: "interruption";
interruption_event: {
reason: string;
};
};
type PingEvent = BaseEvent & {
type: "ping";
ping_event: {
event_id: number;
ping_ms?: number;
};
};
type AgentChatResponsePartEvent = BaseEvent & {
type: "agent_chat_response_part";
text_response_part: {
type: "start" | "delta" | "stop";
text: string;
event_id: number;
response_id: string;
};
};
export type ElevenLabsWebSocketEvent =
| UserTranscriptEvent
| AgentResponseEvent
| AgentResponseCorrectionEvent
| AudioResponseEvent
| InterruptionEvent
| PingEvent
| AgentChatResponsePartEvent;
3

创建 WebSocket Hook

创建用于管理 WebSocket 连接的自定义 Hook:

app/hooks/useAgentConversation.ts
'use client';
import { useCallback, useEffect, useRef, useState } from 'react';
import { useVoiceStream } from 'voice-stream';
import type { ElevenLabsWebSocketEvent } from '../types/websocket';
const sendMessage = (websocket: WebSocket, request: object) => {
if (websocket.readyState !== WebSocket.OPEN) {
return;
}
websocket.send(JSON.stringify(request));
};
export const useAgentConversation = () => {
const websocketRef = useRef<WebSocket>(null);
const [isConnected, setIsConnected] = useState<boolean>(false);
const { startStreaming, stopStreaming } = useVoiceStream({
onAudioChunked: (audioData) => {
if (!websocketRef.current) return;
sendMessage(websocketRef.current, {
user_audio_chunk: audioData,
});
},
});
const startConversation = useCallback(async () => {
if (isConnected) return;
const websocket = new WebSocket("wss://api.elevenlabs.io/v1/convai/conversation");
websocket.onopen = async () => {
setIsConnected(true);
sendMessage(websocket, {
type: "conversation_initiation_client_data",
});
await startStreaming();
};
websocket.onmessage = async (event) => {
const data = JSON.parse(event.data) as ElevenLabsWebSocketEvent;
// Handle ping events to keep connection alive
if (data.type === "ping") {
setTimeout(() => {
sendMessage(websocket, {
type: "pong",
event_id: data.ping_event.event_id,
});
}, data.ping_event.ping_ms);
}
if (data.type === "user_transcript") {
const { user_transcription_event } = data;
console.log("User transcript", user_transcription_event.user_transcript);
}
if (data.type === "agent_response") {
const { agent_response_event } = data;
console.log("Agent response", agent_response_event.agent_response);
}
if (data.type === "agent_response_correction") {
const { agent_response_correction_event } = data;
console.log("Agent response correction", agent_response_correction_event.corrected_agent_response);
}
if (data.type === "interruption") {
// Handle interruption
}
if (data.type === "audio") {
const { audio_event } = data;
// Implement your own audio playback system here
// Note: You'll need to handle audio queuing to prevent overlapping
// as the WebSocket sends audio events in chunks
}
if (data.type === "agent_chat_response_part") {
const { text_response_part } = data;
const { type: partType, text, response_id } = text_response_part;
// Handle the agent's response text as it is generated. Enable
// agent_chat_response_part in the agent's client_events to receive
// this during voice conversations.
console.log("Chat response part:", partType, text, response_id);
}
};
websocketRef.current = websocket;
websocket.onclose = async () => {
websocketRef.current = null;
setIsConnected(false);
stopStreaming();
};
}, [startStreaming, isConnected, stopStreaming]);
const stopConversation = useCallback(async () => {
if (!websocketRef.current) return;
websocketRef.current.close();
}, []);
useEffect(() => {
return () => {
if (websocketRef.current) {
websocketRef.current.close();
}
};
}, []);
return {
startConversation,
stopConversation,
isConnected,
};
};
4

创建对话组件

创建使用 WebSocket Hook 的组件:

app/components/Conversation.tsx
'use client';
import { useCallback } from 'react';
import { useAgentConversation } from '../hooks/useAgentConversation';
export function Conversation() {
const { startConversation, stopConversation, isConnected } = useAgentConversation();
const handleStart = useCallback(async () => {
try {
await navigator.mediaDevices.getUserMedia({ audio: true });
await startConversation();
} catch (error) {
console.error('Failed to start conversation:', error);
}
}, [startConversation]);
return (
<div className="flex flex-col items-center gap-4">
<div className="flex gap-2">
<button
onClick={handleStart}
disabled={isConnected}
className="px-4 py-2 bg-blue-500 text-white rounded disabled:bg-gray-300"
>
Start Conversation
</button>
<button
onClick={stopConversation}
disabled={!isConnected}
className="px-4 py-2 bg-red-500 text-white rounded disabled:bg-gray-300"
>
Stop Conversation
</button>
</div>
<div className="flex flex-col items-center">
<p>Status: {isConnected ? 'Connected' : 'Disconnected'}</p>
</div>
</div>
);
}

后续步骤

  1. 音频播放:使用 Web Audio API 或库实现自己的音频播放系统。请记得处理音频队列,避免 WebSocket 分块发送音频事件时发生重叠。
  2. 错误处理:添加重试逻辑和错误恢复机制
  3. UI 反馈:添加语音活动和连接状态的视觉指示器

延迟管理

为确保对话流畅,请采用以下策略:

  • 自适应缓冲: 根据网络状况调整音频缓冲。
  • 抖动缓冲: 实现抖动缓冲区,平滑数据包到达时间的波动。
  • Ping-Pong 监控: 使用 ping 和 pong 事件测量往返时间,并据此调整。

安全最佳实践

  • 定期轮换 API 密钥,并使用环境变量存储。
  • 实施速率限制以防止滥用。
  • 请求麦克风访问权限时,清楚说明用途。
  • 优化分块:调整音频块时长,平衡延迟与效率。

其他资源