JavaScript SDK

ElevenAgents SDK:几分钟内即可部署自定义交互式语音智能体。

另请参阅 ElevenAgents 概览

安装

通过包管理器在项目中安装此软件包。

npm install @elevenlabs/client
# or
yarn add @elevenlabs/client
# or
pnpm install @elevenlabs/client

从早期版本升级?运行 npx skills add elevenlabs/packages,为 AI 编程智能体安装 elevenlabs:sdk-migration skill,自动处理导入变更和 API 更新。

使用

此库主要适用于原生 JavaScript 项目开发,也可作为针对特定框架定制库的基础。 建议先确认所用框架是否有专属库。 不过,任何基于 JavaScript 的项目都可以使用此库。

初始化对话

首先,使用 Conversation.startSession 创建新的对话会话:

const conversation = await Conversation.startSession(options);

这会建立连接,并开始通过麦克风与 ElevenLabs Agents 智能体交流。建议在开始对话前,先在应用 UI 中说明并请求麦克风访问权限:

// call after explaining to the user why the microphone access is needed
await navigator.mediaDevices.getUserMedia({ audio: true });

会话配置

传给 startSession 的选项用于指定如何建立会话。可通过公开或私有智能体发起对话。

公开智能体

无需任何身份验证的智能体可通过智能体 ID 发起对话。可在 ElevenLabs UI 中获取智能体 ID。

对于公开智能体,可直接使用 ID:

const conversation = await Conversation.startSession({
agentId: "agent_7101k5zvyjhmfg983brhmhkd98n6",
});

系统会根据对话模式自动推断连接类型。语音对话默认使用 WebRTC,纯文本对话默认使用 WebSocket。必要时,仍可显式指定 connectionType: 'webrtc' 或 connectionType: 'websocket'。

私有智能体

如果对话需要授权,需要在服务器上添加专用端点:使用 ElevenLabs API 请求签名 URL(使用 WebSockets 连接类型时)或对话令牌(使用 WebRTC 时),然后将其返回给客户端。

以下是 WebSocket 连接示例:

// Node.js server
app.get("/signed-url", yourAuthMiddleware, async (req, res) => {
const response = await fetch(
`https://api.elevenlabs.io/v1/convai/conversation/get-signed-url?agent_id=${process.env.AGENT_ID}`,
{
method: "GET",
headers: {
// Requesting a signed url requires your ElevenLabs API key
// Do NOT expose your API key to the client!
"xi-api-key": process.env.XI_API_KEY,
},
}
);
if (!response.ok) {
return res.status(500).send("Failed to get signed URL");
}
const body = await response.json();
res.send(body.signed_url);
});
// Client
const response = await fetch("/signed-url", yourAuthHeaders);
const signedUrl = await response.text();
const conversation = await Conversation.startSession({
signedUrl,
});

以下是 WebRTC 示例:

// Node.js server
app.get("/conversation-token", yourAuthMiddleware, async (req, res) => {
const response = await fetch(
`https://api.elevenlabs.io/v1/convai/conversation/token?agent_id=${process.env.AGENT_ID}`,
{
headers: {
// Requesting a conversation token requires your ElevenLabs API key
// Do NOT expose your API key to the client!
"xi-api-key": process.env.ELEVENLABS_API_KEY,
},
}
);
if (!response.ok) {
return res.status(500).send("Failed to get conversation token");
}
const body = await response.json();
res.send(body.token);
});

获取令牌后,将其提供给 startSession 即可通过 WebRTC 发起对话。

// Client
const response = await fetch("/conversation-token", yourAuthHeaders);
const conversationToken = await response.text();
const conversation = await Conversation.startSession({
conversationToken,
});

可选回调

传给 startSession 的选项还可用于注册可选回调:

  • onConnect:建立对话 WebSocket 连接时调用的处理函数。
  • onDisconnect:结束对话 WebSocket 连接时调用的处理函数。
  • onMessage:收到新文本消息时调用的处理函数。可能是用户语音的临时或最终转录文本,也可能是 LLM 生成的回复。主要用于处理对话转录。
  • onError:遇到错误时调用的处理函数。
  • onStatusChange:连接状态每次变化时调用的处理函数。状态可以是 connected、connecting 和 disconnected(初始状态)。
  • onModeChange:状态变化时调用的处理函数,例如智能体从 speaking 切换到 listening,或反向切换。
  • onCanSendFeedbackChange:反馈发送变为可用或不可用时调用的处理函数。
  • onAudioAlignment:收到音频对齐数据时调用的处理函数,为智能体语音提供字符级时间信息。

并非所有客户端事件都会默认对智能体启用。如果已启用回调但未收到事件,请确认 ElevenLabs 智能体已启用相应事件。 可在 ElevenLabs 控制台中智能体设置的“Advanced”标签页进行配置。

返回值

startSession 会返回可用于控制会话的对话实例(根据模式为 VoiceConversation 或 TextConversation)。如果无法建立会话,此方法会抛出错误。例如用户拒绝麦克风访问权限,或连接失败时,就会出现这种情况。

endSession

用于手动结束对话的方法。该方法会结束对话并断开 WebSocket 连接。 之后,对话实例将无法使用,可以安全丢弃。

await conversation.endSession();

getId

返回对话 ID 的方法。

const id = conversation.getId();

setVolume

用于设置对话输出音量的方法。接受一个包含音量字段的对象,取值范围为 0 到 1。

await conversation.setVolume({ volume: 0.5 });

getInputVolume / getOutputVolume

返回当前输入/输出音量的方法,范围为 0 到 1,其中 0 为 -100 dB,1 为 -30 dB。

const inputVolume = await conversation.getInputVolume();
const outputVolume = await conversation.getOutputVolume();

sendFeedback

用于向智能体发送二元反馈的方法。该方法接受布尔值:true 表示正面反馈,false 表示负面反馈。

反馈始终关联到最近一次智能体回复,每次回复只能发送一次。

可监听 onCanSendFeedbackChange,了解当前是否可以发送反馈。

conversation.sendFeedback(true); // positive feedback
conversation.sendFeedback(false); // negative feedback

sendContextualUpdate

用于向智能体发送上下文更新的方法。可用来告知智能体与对话没有直接关系、但可能影响其回复的用户操作。

conversation.sendContextualUpdate(
"User navigated to another page. Consider it for next response, but don't react to this contextual update."
);

sendUserMessage

向智能体发送文本消息。

可让用户输入消息,而不是使用麦克风。与 sendContextualUpdate 不同,这会被视为用户消息,并提示智能体在对话中作出回应。

sendButton.addEventListener("click", (e) => {
conversation.sendUserMessage(textInput.value);
textInput.value = "";
});

sendUserActivity

通知智能体用户活动。

检测到用户活动后,智能体至少 2 秒内不会尝试说话。

可在用户输入文字时防止智能体打断用户。

textInput.addEventListener("input", () => {
conversation.sendUserActivity();
});

setMicMuted

用于静音/取消静音麦克风的方法。

// Mute the microphone
conversation.setMicMuted(true);
// Unmute the microphone
conversation.setMicMuted(false);

changeInputDevice

可在正在进行的语音对话中更改音频输入设备。此方法仅适用于语音对话。

在 WebRTC 模式下,输入格式和采样率分别硬编码为 pcm 和 48000。 更改输入设备时修改这些值不会产生任何效果。

const conversation = await Conversation.startSession({
agentId: "agent_7101k5zvyjhmfg983brhmhkd98n6",
// Alternatively you can provide a device ID when starting the session
// Useful if you want to start the conversation with a non-default device
inputDeviceId: "a1b2c3d4e5f6",
});
// Change to a specific input device
await conversation.changeInputDevice({
sampleRate: 16000,
format: "pcm",
preferHeadphonesForIosDevices: true,
inputDeviceId: "a1b2c3d4e5f6",
});

如果设备 ID 无效,将改用默认设备。

changeOutputDevice

可在正在进行的语音对话中更改音频输出设备。此方法仅适用于语音对话。

在 WebRTC 模式下,输出格式和采样率分别硬编码为 pcm 和 48000。 更改输出设备时修改这些值不会产生任何效果。

const conversation = await Conversation.startSession({
agentId: "agent_7101k5zvyjhmfg983brhmhkd98n6",
// Alternatively you can provide a device ID when starting the session
// Useful if you want to start the conversation with a non-default device
outputDeviceId: "a1b2c3d4e5f6",
});
// Change to a specific output device
await conversation.changeOutputDevice({
sampleRate: 16000,
format: "pcm",
outputDeviceId: "a1b2c3d4e5f6",
});

设备切换仅适用于语音对话。如果未提供特定 deviceId,浏览器将使用默认设备选择。 可通过 MediaDevices.enumerateDevices() API 枚举可用设备。

getInputByteFrequencyData / getOutputByteFrequencyData

返回包含当前输入/输出频率数据的 Uint8Array 的方法。更多信息请参阅 AnalyserNode.getByteFrequencyData。

这些方法仅适用于语音对话。在 WebRTC 模式下,音频被硬编码为使用 pcm_48000,因此使用返回数据进行的可视化可能会显示出与 WebSocket 连接不同的模式。