React SDK
useScribe:在 React 中进行实时语音转文本
安装
npm install @elevenlabs/react# oryarn add @elevenlabs/react# orpnpm install @elevenlabs/react
使用 ElevenLabs 语音转文本技能,通过 AI 编程助手转录音频:
npx skills add elevenlabs/skills --skill speech-to-text
@elevenlabs/react 会重新导出 @elevenlabs/client 中的所有内容,因此无需同时安装
两个软件包。
使用方法
以下是连接 Scribe 并显示实时转录的最小可用示例:
import { useScribe } from "@elevenlabs/react";import { useEffect } from "react";function MyComponent() {const scribe = useScribe({modelId: "scribe_v2_realtime",onPartialTranscript: (data) => {console.log("Partial:", data.text);},onCommittedTranscript: (data) => {console.log("Committed:", data.text);},});// Start recordingconst handleStart = async () => {try {const token = await fetchTokenFromServer();await scribe.connect({token,microphone: {echoCancellation: true,noiseSuppression: true,},});} catch (err) {console.error("Failed to start recording:", err);}};// Stop recordingconst handleDisconnect = () => {scribe.disconnect();};// Disconnect on unmountuseEffect(() => {return () => {if (scribe.isConnected) {scribe.disconnect();}};}, [scribe]);return (<div><button onClick={handleStart} disabled={scribe.isConnected}>Start Recording</button><button onClick={handleDisconnect} disabled={!scribe.isConnected}>Stop</button>{scribe.partialTranscript && <p>Live: {scribe.partialTranscript}</p>}<div>{scribe.committedTranscripts.map((t) => (<p key={t.id}>{t.text}</p>))}</div></div>);}
获取令牌
Scribe 需要使用一次性令牌进行身份验证。在服务器上创建一个 API 端点:
// Node.js serverapp.get("/scribe-token", yourAuthMiddleware, async (req, res) => {const response = await fetch("https://api.elevenlabs.io/v1/single-use-token/realtime_scribe", {method: "POST",headers: {"xi-api-key": process.env.ELEVENLABS_API_KEY,},});const data = await response.json();res.json({ token: data.token });});
ElevenLabs API 密钥属于敏感信息,切勿将其暴露给客户端。始终在 服务器上生成令牌。
// Clientconst fetchToken = async () => {const response = await fetch("/scribe-token");const { token } = await response.json();return token;};
Hook 选项
通过默认选项和回调配置 Hook:
const scribe = useScribe({// Connection options (can be overridden in connect())token: "optional-default-token",modelId: "scribe_v2_realtime",baseUri: "wss://api.elevenlabs.io",// VAD optionscommitStrategy: CommitStrategy.VAD,vadSilenceThresholdSecs: 0.5,vadThreshold: 0.5,minSpeechDurationMs: 100,minSilenceDurationMs: 500,languageCode: "en",// Microphone options (for automatic mode)microphone: {deviceId: "optional-device-id",echoCancellation: true,noiseSuppression: true,autoGainControl: true,},// Manual audio options (for file transcription)audioFormat: AudioFormat.PCM_16000,sampleRate: 16000,// Auto-connect on mountautoConnect: false,// Event callbacksonSessionStarted: () => console.log("Session started"),onPartialTranscript: (data) => console.log("Partial:", data.text),onCommittedTranscript: (data) => console.log("Committed:", data.text),onCommittedTranscriptWithTimestamps: (data) => console.log("With timestamps:", data),onError: (error) => console.error("Error:", error),onAuthError: (data) => console.error("Auth error:", data.error),onQuotaExceededError: (data) => console.error("Quota exceeded:", data.error),onConnect: () => console.log("Connected"),onDisconnect: () => console.log("Disconnected"),});
连接选项
| 属性 | 类型 | 说明 |
|---|---|---|
| token | string | 用于 WebSocket 身份验证的一次性令牌。 |
| modelId | string | 模型 ID(例如 "scribe_v2_realtime")。 |
| baseUri | string | 自定义 WebSocket 基础 URI。默认为 wss://api.elevenlabs.io。 |
VAD 选项
使用 VAD 提交策略时,这些选项控制何时自动提交转录内容。
| 属性 | 类型 | 默认值 | 说明 |
|---|---|---|---|
| commitStrategy | CommitStrategy | "manual" | "manual" 或 "vad"。 |
| vadSilenceThresholdSecs | number | 1.5 | VAD 提交前的静音秒数(0.3-3.0)。 |
| vadThreshold | number | 0.4 | VAD 灵敏度(0.1-0.9,数值越低越灵敏)。 |
| minSpeechDurationMs | number | 100 | 最短语音时长,单位为 ms(50-2000)。 |
| minSilenceDurationMs | number | 100 | 最短静音时长,单位为 ms(50-2000)。 |
音频选项
| 属性 | 类型 | 说明 |
|---|---|---|
| languageCode | string | ISO-639-1 或 ISO-639-3 语言代码。留空可自动检测。 |
| microphone | object | 麦克风模式的麦克风设置。详见下文。 |
| audioFormat | AudioFormat | 手动模式的音频编码格式(例如 AudioFormat.PCM_16000)。 |
| sampleRate | number | 手动模式的采样率。必须与 audioFormat 匹配。 |
microphone 对象支持:
| 属性 | 类型 | 说明 |
|---|---|---|
| deviceId | string | 指定麦克风设备 ID。 |
| echoCancellation | boolean | 启用回声消除。 |
| noiseSuppression | boolean | 启用降噪。 |
| autoGainControl | boolean | 启用自动增益控制。 |
行为选项
| 属性 | 类型 | 默认值 | 说明 |
|---|---|---|---|
| autoConnect | boolean | false | 组件挂载时自动连接。 |
| includeTimestamps | boolean | false | 接收词级时间戳。提供 onCommittedTranscriptWithTimestamps 时会自动启用。 |
回调
所有事件回调均为可选,可作为 Hook 选项提供:
- onConnect - WebSocket 连接建立时调用的处理函数。
- onDisconnect - WebSocket 连接关闭时调用的处理函数。
- onSessionStarted - Scribe 会话开始时调用的处理函数。
- onPartialTranscript - 接收临时转录结果时调用的处理函数。接收
{ text: string }。 - onCommittedTranscript - 接收最终转录结果时调用的处理函数。接收
{ text: string }。 - onCommittedTranscriptWithTimestamps - 接收包含词级时间信息的最终转录结果时调用的处理函数。接收
{ text: string; words?: { start: number; end: number }[] }。 - onError - 处理所有错误的通用错误处理函数。接收
Error | Event。 - onAuthError - 发生身份验证错误时调用的处理函数。接收
{ error: string }。
错误回调
通用 onError 回调会在发生任何错误时触发。此外,还提供专用错误回调以便进行精细处理。所有专用错误回调均接收 { error: string }。
| 回调 | 说明 |
|---|---|
| onError | 处理所有错误的通用错误处理函数。 |
| onAuthError | 身份验证错误。 |
| onQuotaExceededError | 已超出使用配额。 |
| onCommitThrottledError | 提交请求受到限流。 |
| onTranscriberError | 转录引擎错误。 |
| onUnacceptedTermsError | 未接受服务条款。 |
| onRateLimitedError | 已限流。 |
| onInputError | 输入格式无效。 |
| onQueueOverflowError | 处理队列已满。 |
| onResourceExhaustedError | 服务器资源已满。 |
| onSessionTimeLimitExceededError | 已达到最长会话时长。 |
| onChunkSizeExceededError | 音频块过大。 |
| onInsufficientAudioActivityError | 音频活动不足,无法维持连接。 |
麦克风模式
直接从用户的麦克风流式传输音频:
function MicrophoneTranscription() {const scribe = useScribe({modelId: "scribe_v2_realtime",});const startRecording = async () => {const token = await fetchToken();await scribe.connect({token,microphone: {echoCancellation: true,noiseSuppression: true,autoGainControl: true,},});};return (<div><button onClick={startRecording} disabled={scribe.isConnected}>{scribe.status === "connecting" ? "Connecting..." : "Start"}</button><button onClick={scribe.disconnect} disabled={!scribe.isConnected}>Stop</button>{scribe.partialTranscript && (<div><strong>Speaking:</strong> {scribe.partialTranscript}</div>)}{scribe.committedTranscripts.map((transcript) => (<div key={transcript.id}>{transcript.text}</div>))}</div>);}
手动音频模式(文件转录)
转录预先录制的音频文件:
import { useScribe, AudioFormat } from "@elevenlabs/react";import { useState } from "react";function FileTranscription() {const [file, setFile] = useState<File | null>(null);const scribe = useScribe({modelId: "scribe_v2_realtime",audioFormat: AudioFormat.PCM_16000,sampleRate: 16000,});const transcribeFile = async () => {if (!file) return;const token = await fetchToken();await scribe.connect({ token });// Decode audio fileconst arrayBuffer = await file.arrayBuffer();const audioContext = new AudioContext({ sampleRate: 16000 });const audioBuffer = await audioContext.decodeAudioData(arrayBuffer);// Convert to PCM16const channelData = audioBuffer.getChannelData(0);const pcmData = new Int16Array(channelData.length);for (let i = 0; i < channelData.length; i++) {const sample = Math.max(-1, Math.min(1, channelData[i]));pcmData[i] = sample < 0 ? sample * 32768 : sample * 32767;}// Send in chunksconst chunkSize = 4096;for (let offset = 0; offset < pcmData.length; offset += chunkSize) {const chunk = pcmData.slice(offset, offset + chunkSize);const bytes = new Uint8Array(chunk.buffer);const base64 = btoa(String.fromCharCode(...bytes));scribe.sendAudio(base64);await new Promise((resolve) => setTimeout(resolve, 50));}// Commit transcriptionscribe.commit();};return (<div><input type="file" accept="audio/*" onChange={(e) => setFile(e.target.files?.[0] || null)} /><button onClick={transcribeFile} disabled={!file || scribe.isConnected}>Transcribe</button>{scribe.committedTranscripts.map((transcript) => (<div key={transcript.id}>{transcript.text}</div>))}</div>);}
返回值
状态
- status - 当前连接状态:
"disconnected"、"connecting"、"connected"、"transcribing"或"error"。 - isConnected - 表示是否已连接的布尔值。
- isTranscribing - 表示是否正在转录的布尔值。
- partialTranscript - 当前部分(临时)转录文本字符串。
- committedTranscripts -
TranscriptSegment对象数组(见下文)。 - error - 当前错误消息,或
null。
const scribe = useScribe(/* options */);console.log(scribe.status); // "connected"console.log(scribe.isConnected); // trueconsole.log(scribe.partialTranscript); // "hello world"console.log(scribe.committedTranscripts); // [{ id: "...", text: "...", words: ..., isFinal: true }]console.log(scribe.error); // null or error string
每个已提交的转录片段具有以下结构:
interface TranscriptSegment {id: string; // Unique identifiertext: string; // Transcript texttimestamp: number; // Unix timestampisFinal: boolean; // Always true for committed transcripts}
方法
connect(options?)
连接到 Scribe。此处提供的选项会覆盖 Hook 默认值:
await scribe.connect({token: "your-token", // Requiredmicrophone: {/* ... */}, // For microphone mode// ORaudioFormat: AudioFormat.PCM_16000, // For manual modesampleRate: 16000,});
disconnect()
断开连接并清理资源:
scribe.disconnect();
sendAudio(audioBase64, options?)
发送音频数据(仅限手动模式):
scribe.sendAudio(base64AudioChunk, {commit: false, // Optional: commit immediatelysampleRate: 16000, // Optional: override sample ratepreviousText: "Previous transcription text", // Optional: context from a previous transcription. Can only be sent in the first audio chunk.});
previousText 字段只能在会话的第一个音频块中发送。在后续
音频块中发送会导致错误。
commit()
手动提交当前转录内容:
scribe.commit();
clearTranscripts()
清除状态中的所有转录内容:
scribe.clearTranscripts();
getConnection()
获取底层连接实例:
const connection = scribe.getConnection();// Returns RealtimeConnection | null
提交策略
控制何时提交转录内容:
import { CommitStrategy } from '@elevenlabs/react';// Manual (default) - you control when to commitconst scribe = useScribe({commitStrategy: CommitStrategy.MANUAL,});// Later...scribe.commit(); // Commit transcription// Voice Activity Detection - model detects silences and automatically commitsconst scribe = useScribe({commitStrategy: CommitStrategy.VAD,});
详情请参阅转录内容和提交策略。
完整示例
以下是使用 useScribe Hook 和基于 VAD 提交策略的完整 React 组件示例:
import { useScribe, CommitStrategy } from "@elevenlabs/react";import { useEffect } from "react";function ScribeDemo() {const scribe = useScribe({modelId: "scribe_v2_realtime",commitStrategy: CommitStrategy.VAD,onSessionStarted: () => console.log("Started"),onCommittedTranscript: (data) => console.log("Committed:", data.text),onError: (error) => console.error("Error:", error),});const startMicrophone = async () => {const token = await fetchToken();await scribe.connect({token,microphone: {echoCancellation: true,noiseSuppression: true,},});};const handleDisconnect = () => scribe.disconnect();const handleClearTranscripts = () => scribe.clearTranscripts();useEffect(() => {return () => {handleDisconnect();};}, []);return (<div><h1>Scribe Demo</h1>{/* Status */}<div>Status: {scribe.status}{scribe.error && <span>Error: {scribe.error}</span>}</div>{/* Controls */}<div>{!scribe.isConnected ? (<button onClick={startMicrophone}>Start Recording</button>) : (<button onClick={handleDisconnect}>Stop</button>)}<button onClick={handleClearTranscripts}>Clear</button></div>{/* Live Transcript */}{scribe.partialTranscript && (<div><strong>Live:</strong> {scribe.partialTranscript}</div>)}{/* Committed Transcripts */}<div><h2>Transcripts ({scribe.committedTranscripts.length})</h2>{scribe.committedTranscripts.map((t) => (<div key={t.id}><span>{new Date(t.timestamp).toLocaleTimeString()}</span><p>{t.text}</p></div>))}</div></div>);}