React SDK

ElevenAgents SDK:数分钟内部署定制化交互式语音智能体。

有关 ElevenAgents 工作原理的说明,请参阅 ElevenAgents 概览。

安装

通过包管理器在项目中安装此软件包。

npm install @elevenlabs/react
# or
yarn add @elevenlabs/react
# or
pnpm install @elevenlabs/react

正在从早期版本升级?运行 npx skills add elevenlabs/packages,为 AI 编程智能体安装 elevenlabs:sdk-migration skill,自动处理导入变更、ConversationProvider 包装及 API 更新。

@elevenlabs/react 会重新导出 @elevenlabs/client 的所有内容,因此无需同时安装 两个软件包。

使用

以下是一个可运行的最小示例,用于连接智能体并让用户开始和结束语音对话:

import {
ConversationProvider,
useConversationControls,
useConversationStatus,
} from "@elevenlabs/react";
function App() {
return (
<ConversationProvider>
<Agent />
</ConversationProvider>
);
}
function Agent() {
const { startSession, endSession } = useConversationControls();
const { status } = useConversationStatus();
if (status === "connected") {
return <button onClick={endSession}>End</button>;
}
return (
<button onClick={() => startSession({ agentId: "agent_7101k5zvyjhmfg983brhmhkd98n6" })}>
Start
</button>
);
}

以下章节将详细说明各部分。

ConversationProvider

所有对话 Hook 都必须在 ConversationProvider 内使用。请使用此 Provider 包装应用(或相关子树)。

import { ConversationProvider } from "@elevenlabs/react";
function App() {
return (
<ConversationProvider>
<YourComponents />
</ConversationProvider>
);
}

Provider 属性

Provider 接受与 useConversation 相同的选项,包括回调、客户端工具、覆盖设置和服务器位置,因此可在 Provider 层配置,而非在每个 Hook 使用方中配置。

<ConversationProvider
onConnect={() => console.log("Connected")}
onDisconnect={() => console.log("Disconnected")}
onError={(error) => console.error("Error:", error)}
clientTools={{
displayMessage: (parameters: { text: string }) => {
alert(parameters.text);
return "Message displayed";
},
}}
serverLocation="eu-residency"
>
<YourComponents />
</ConversationProvider>
受控静音状态

Provider 支持用于受控静音状态管理的 isMuted 和 onMutedChange 属性,让你可以在外部持久化静音状态(例如跨会话保存)。

const [muted, setMuted] = useState(false);
<ConversationProvider isMuted={muted} onMutedChange={setMuted}>
<YourComponents />
</ConversationProvider>;

useConversation

一个便捷的 React Hook,将所有细粒度 Hook 整合为单一返回值。需要有祖先 ConversationProvider。

为获得更好的渲染性能,建议改用细粒度 Hook。 useConversation 会在任何状态变化时重新渲染,而细粒度 Hook 仅在其对应状态片段变化时 重新渲染。

初始化对话

import { useConversation } from "@elevenlabs/react";
function MyComponent() {
const conversation = useConversation();
// ...
}

请注意,ElevenAgents 需要麦克风访问权限才能进行语音对话。建议在对话开始前,在应用 UI 中说明并允许访问。

// call after explaining to the user why the microphone access is needed
await navigator.mediaDevices.getUserMedia({ audio: true });

选项

可选择使用选项初始化 Hook。这些选项也可在 ConversationProvider 层传入。

const conversation = useConversation({
/* options object */
});

选项包括:

  • clientTools - 可由智能体调用的客户端工具对象定义。详情请参阅下文。
  • overrides - 对话设置覆盖项的对象定义。详情请参阅下文。
  • textOnly - 对话是否应以纯文本模式运行。详情请参阅下文。
  • serverLocation - 指定服务器位置("us"、"eu-residency"、"in-residency"、"global")。默认为 "us"。

回调概览

  • onConnect - 建立对话连接时调用的处理程序。
  • onDisconnect - 结束对话连接时调用的处理程序。
  • onMessage - 收到新消息时调用的处理程序。这些消息可以是用户语音的临时或最终转录、LLM 生成的回复,或启用调试选项时的调试消息。
  • onError - 遇到错误时调用的处理程序。
  • onAudio - 收到音频数据时调用的处理程序。
  • onModeChange - 对话模式变化时调用的处理程序(说话/聆听)。
  • onStatusChange - 连接状态变化时调用的处理程序。
  • onCanSendFeedbackChange - 反馈发送能力变化时调用的处理程序。
  • onDebug - 有调试信息可用时调用的处理程序。
  • onUnhandledClientToolCall - 遇到未处理的客户端工具调用时调用的处理程序。
  • onVadScore - 语音活动检测分数变化时调用的处理程序。
  • onAudioAlignment - 收到音频对齐数据时调用的处理程序,为智能体语音提供字符级时间信息。
  • onAgentChatResponsePart - 在生成过程中,以开始、增量和停止事件提供智能体回复文本时调用的处理程序。纯文本模式始终发送;对于语音对话,请在智能体的 client_events 配置中启用 agent_chat_response_part。
客户端工具

客户端工具让智能体能够调用客户端功能。可用于触发客户端操作,例如打开模态窗口,或代表用户调用 API。

客户端工具定义是一个函数对象,且必须与 ElevenLabs UI 中的配置完全一致;你可以在那里命名和描述不同工具,并设置由智能体传入的参数。

const conversation = useConversation({
clientTools: {
displayMessage: (parameters: { text: string }) => {
alert(parameters.text);
return "Message displayed";
},
},
});

如果函数返回值,该值会作为响应传回智能体。

必须在 ElevenLabs UI 中将工具明确设为阻塞对话,智能体才会等待并响应结果。否则,智能体会假定调用成功并继续对话。

如需以更符合 React 惯例的方式注册客户端工具,请参阅 useConversationClientTool。

对话覆盖设置

你可以选择覆盖多项对话设置,并根据其他用户交互动态设置它们。

支持覆盖多项设置。这些设置为可选项,可用于定制对话体验。

可用设置如下:

const conversation = useConversation({
overrides: {
agent: {
prompt: {
prompt: "My custom prompt",
},
firstMessage: "My custom first message",
language: "en",
},
tts: {
voiceId: "custom voice id",
},
conversation: {
textOnly: true,
},
},
});
纯文本

如果智能体配置为纯文本模式,即不发送或接收音频消息,可使用此标志来使用更轻量的对话版本。此时不会请求用户授予麦克风权限,也不会创建音频上下文。

const conversation = useConversation({
textOnly: true,
});
受控状态

可通过 Hook 选项直接控制对话状态的某些方面:

const [micMuted, setMicMuted] = useState(false);
const conversation = useConversation({
micMuted,
// ... other options
});
// Update controlled state
setMicMuted(true); // This will automatically mute the microphone
数据驻留

可指定要连接的 ElevenLabs 服务器区域。更多信息请参阅数据驻留指南。

const conversation = useConversation({
serverLocation: "eu-residency", // or "us", "in-residency", "global"
});

方法

startSession

startSession 方法建立连接,并开始使用麦克风与 ElevenLabs Agents 智能体通信。该方法接受一个选项对象,其中必须提供 signedUrl、conversationToken 或 agentId。

可通过 ElevenLabs UI 获取智能体 ID。

还建议传入自己的最终用户 ID,以将对话映射到用户。

系统会根据对话模式自动推断连接类型。语音对话使用 WebRTC,纯文本对话默认使用 WebSocket。 如有需要,仍可明确指定 connectionType。

const conversation = useConversation();
// For public agents, pass in the agent ID
const conversationId = await conversation.startSession({
agentId: "agent_7101k5zvyjhmfg983brhmhkd98n6",
userId: "user_9302xkm82nds93", // optional field
});

对于公共智能体(即未启用身份验证的智能体),仅需提供 agentId。

如果对话需要授权,请使用 REST API 为 WebSocket 连接生成签名链接,或为 WebRTC 连接生成对话令牌。

startSession 会返回一个解析为 conversationId 的 Promise。该值是全局唯一的对话 ID,可用于标识不同对话。

// Node.js server
app.get("/signed-url", yourAuthMiddleware, async (req, res) => {
const response = await fetch(
`https://api.elevenlabs.io/v1/convai/conversation/get-signed-url?agent_id=${process.env.AGENT_ID}`,
{
headers: {
// Requesting a signed url requires your ElevenLabs API key
// Do NOT expose your API key to the client!
"xi-api-key": process.env.ELEVENLABS_API_KEY,
},
}
);
if (!response.ok) {
return res.status(500).send("Failed to get signed URL");
}
const body = await response.json();
res.send(body.signed_url);
});
// Client
const response = await fetch("/signed-url", yourAuthHeaders);
const signedUrl = await response.text();
await conversation.startSession({
signedUrl,
});
endSession

手动结束对话的方法。此方法会断开连接并结束对话。

await conversation.endSession();
setVolume

设置对话输出音量。接受一个 volume 字段介于 0 和 1 之间的对象。

await conversation.setVolume({ volume: 0.5 });
sendUserMessage

向智能体发送文本消息。

可用于让用户输入消息,而非使用麦克风。与 sendContextualUpdate 不同,这会被视为用户消息,并提示智能体在对话中作出回应。

const { sendUserMessage, sendUserActivity } = useConversation();
const [value, setValue] = useState("");
return (
<>
<input
value={value}
onChange={e => {
setValue(e.target.value);
sendUserActivity();
}}
/>
<button
onClick={() => {
sendUserMessage(value);
setValue("");
}}
>
SEND
</button>
</>
);
sendContextualUpdate

向智能体发送不会触发回复的上下文信息。

const { sendContextualUpdate } = useConversation();
sendContextualUpdate(
"User navigated to another page. Consider it for next response, but don't react to this contextual update."
);
sendFeedback

提供有关对话质量的反馈。这有助于提升智能体的表现。

const { sendFeedback } = useConversation();
sendFeedback(true); // positive feedback
sendFeedback(false); // negative feedback
sendUserActivity

通知智能体用户活动,以防止被打断。适用于用户正在主动使用应用、智能体应暂停说话的情况,例如用户正在聊天中输入内容时。

智能体收到此信号后会暂停说话约 2 秒。

const { sendUserActivity } = useConversation();
// Call this when user is typing to prevent interruption
sendUserActivity();
changeInputDevice

在进行中的语音对话期间切换音频输入设备。此方法仅适用于语音对话。

// Change to a specific input device
conversation.changeInputDevice({
sampleRate: 16000,
format: "pcm",
preferHeadphonesForIosDevices: true,
inputDeviceId: "a1b2c3d4e5f6", // Optional: specific device ID
});
changeOutputDevice

在进行中的语音对话期间切换音频输出设备。此方法仅适用于语音对话。

// Change to a specific output device
conversation.changeOutputDevice({
sampleRate: 16000,
format: "pcm",
outputDeviceId: "a1b2c3d4e5f6", // Optional: specific device ID
});

设备切换仅适用于语音对话。如果未提供特定 deviceId,浏览器将使用默认设备选择。可使用 MediaDevices.enumerateDevices() API 枚举可用设备。

getId

返回当前对话 ID。

const { getId } = useConversation();
const conversationId = getId();
console.log(conversationId); // e.g., "conv_9001k1zph3fkeh5s8xg9z90swaqa"
getInputVolume / getOutputVolume

返回当前输入/输出音量级别(0-1 范围)的方法。

const { getInputVolume, getOutputVolume } = useConversation();
const inputLevel = getInputVolume();
const outputLevel = getOutputVolume();
getInputByteFrequencyData / getOutputByteFrequencyData

返回包含当前输入/输出频率数据的 Uint8Array 的方法。更多信息请参阅 AnalyserNode.getByteFrequencyData。

const { getInputByteFrequencyData, getOutputByteFrequencyData } = useConversation();
const inputFrequencyData = getInputByteFrequencyData();
const outputFrequencyData = getOutputByteFrequencyData();

这些方法仅适用于语音对话。在 WebRTC 模式下,音频被硬编码为使用 pcm_48000, 因此使用返回数据的任何可视化可能会显示与 WebSocket 连接不同的模式。

sendMCPToolApprovalResult

发送 MCP(模型上下文协议)工具调用的批准结果。

const { sendMCPToolApprovalResult } = useConversation();
// Approve a tool call
sendMCPToolApprovalResult("tc_8k2m4n6p8r0t", true);
// Reject a tool call
sendMCPToolApprovalResult("tc_8k2m4n6p8r0t", false);

返回值

除上述方法外,useConversation 还会返回以下响应式状态:

  • status - 当前连接状态("disconnected"、"connecting"、"connected")。
  • isSpeaking - 智能体当前是否正在说话。
  • isListening - 智能体当前是否正在聆听。
  • mode - 当前对话模式("speaking" 或 "listening")。
  • isMuted - 麦克风当前是否静音。
  • setMuted - 用于将麦克风静音/取消静音的函数。
  • canSendFeedback - 当前对话是否可以提交反馈。
  • message - 对话中的最新消息。
const { status, isSpeaking, isListening, isMuted, setMuted, canSendFeedback } = useConversation();
return (
<div>
<p>Status: {status}</p>
<p>Agent is {isSpeaking ? 'speaking' : 'listening'}</p>
<button onClick={() => setMuted(!isMuted)}>
{isMuted ? 'Unmute' : 'Mute'}
</button>
</div>
);

细粒度 Hook

为获得更好的渲染性能,请使用这些 Hook 替代 useConversation。每个 Hook 仅订阅其特定状态片段,因此组件仅在其使用的数据变化时重新渲染。

所有细粒度 Hook 都需要有祖先 ConversationProvider。

useConversationControls

返回用于控制对话的操作方法。此 Hook 不会导致重新渲染,因为它仅提供稳定的函数引用。

import { useConversationControls } from "@elevenlabs/react";
function Controls() {
const {
startSession,
endSession,
sendUserMessage,
sendContextualUpdate,
sendUserActivity,
setVolume,
changeInputDevice,
changeOutputDevice,
sendMCPToolApprovalResult,
getId,
getInputVolume,
getOutputVolume,
getInputByteFrequencyData,
getOutputByteFrequencyData,
} = useConversationControls();
return (
<button onClick={() => startSession({ agentId: "agent_7101k5zvyjhmfg983brhmhkd98n6" })}>
Start
</button>
);
}

useConversationStatus

返回当前连接状态和可选的状态消息。

import { useConversationStatus } from "@elevenlabs/react";
function StatusIndicator() {
const { status, message } = useConversationStatus();
return <p>Status: {status}</p>; // "disconnected" | "connecting" | "connected"
}

useConversationInput

返回静音状态和用于切换麦克风的设置函数。

import { useConversationInput } from "@elevenlabs/react";
function MuteToggle() {
const { isMuted, setMuted } = useConversationInput();
return <button onClick={() => setMuted(!isMuted)}>{isMuted ? "Unmute" : "Mute"}</button>;
}

useConversationMode

返回智能体的说话/聆听状态。

import { useConversationMode } from "@elevenlabs/react";
function ModeIndicator() {
const { mode, isSpeaking, isListening } = useConversationMode();
return <p>Agent is {isSpeaking ? "speaking" : "listening"}</p>;
}

useConversationFeedback

返回反馈可用性及提交反馈的方法。

import { useConversationFeedback } from "@elevenlabs/react";
function FeedbackButtons() {
const { canSendFeedback, sendFeedback } = useConversationFeedback();
if (!canSendFeedback) return null;
return (
<div>
<button onClick={() => sendFeedback(true)}>Like</button>
<button onClick={() => sendFeedback(false)}>Dislike</button>
</div>
);
}

useRawConversation

返回原始对话实例。这是适用于高级用例的逃生出口,可直接访问底层 VoiceConversation 或 TextConversation 对象。

import { useRawConversation } from "@elevenlabs/react";
function Advanced() {
const conversation = useRawConversation();
// Access the raw conversation instance directly
}

useConversationClientTool

用于从 React 组件动态注册客户端工具的 Hook。组件卸载时,工具会自动注销。

当工具处理程序需要访问 Provider 层无法获得的组件状态或属性时,此 Hook 很有用。

import { useConversationClientTool } from "@elevenlabs/react";
import { useState } from "react";
function MapComponent() {
const [location, setLocation] = useState({ lat: 0, lng: 0 });
useConversationClientTool("getLocation", () => {
return `${location.lat},${location.lng}`;
});
useConversationClientTool("setLocation", (params: { lat: number; lng: number }) => {
setLocation(params);
return "Location updated";
});
return <Map center={location} />;
}

该 Hook 始终使用处理程序最新的闭包值,因此无需担心状态过时。