服务端流式传输

本指南介绍如何使用 ElevenLabs 在服务端实时转录音频。

操作指南 · 假设你已完成文本转语音 快速入门。

概述

ElevenLabs 实时文本转语音 API 使用 Scribe Realtime v2 模型,以极低延迟实时转录音频流。无论是构建语音助手、转录服务,还是任何需要实时语音识别的应用,这个基于 WebSocket 的 API 都会在你说话时提供部分转录,并在语音片段完成后提供已提交转录。

Scribe v2 Realtime 可在服务端实现实时音频转录,音频可来自 URL、文件或自有音频流。

服务端实现与客户端有以下区别:

  • 使用 ElevenLabs API 密钥,而非一次性令牌。
  • 支持直接从 URL 流式传输,无需手动分割音频块。

如需直接从麦克风流式传输音频,请参阅客户端流式传输指南。

快速入门

本指南假设你已设置 API 密钥和 SDK。如尚未完成,请先完成快速入门。

1

配置 SDK

SDK 提供两种实时转录音频的方式:从 URL 流式传输,或手动分割文件或自有音频流中的音频。

有关 API 支持的完整参数和选项列表,请参阅 API 参考文档。

此示例展示如何使用官方 SDK 从 URL 流式传输音频文件。

从 URL 流式传输时需要 ffmpeg 工具。请访问其网站获取安装说明。

根据所选语言,新建名为 example.py 或 example.mts 的文件,并添加以下代码:

from dotenv import load_dotenv
import os
import asyncio
from elevenlabs import ElevenLabs, RealtimeEvents, RealtimeUrlOptions
load_dotenv()
async def main():
elevenlabs = ElevenLabs(api_key=os.getenv("ELEVENLABS_API_KEY"))
# Create an event to signal when to stop
stop_event = asyncio.Event()
# Connect to a streaming audio URL
connection = await elevenlabs.speech_to_text.realtime.connect(RealtimeUrlOptions(
model_id="scribe_v2_realtime",
url="https://npr-ice.streamguys1.com/live.mp3",
include_timestamps=True,
))
# Set up event handlers
def on_session_started(data):
print(f"Session started: {data}")
def on_partial_transcript(data):
print(f"Partial: {data.get('text', '')}")
def on_committed_transcript(data):
print(f"Committed: {data.get('text', '')}")
# Committed transcripts with word-level timestamps. Only received when include_timestamps is set to True.
def on_committed_transcript_with_timestamps(data):
print(f"Committed with timestamps: {data.get('words', '')}")
# Errors - will catch all errors, both server and websocket specific errors
def on_error(error):
print(f"Error: {error}")
# Signal to stop on error
stop_event.set()
def on_close():
print("Connection closed")
# Register event handlers
connection.on(RealtimeEvents.SESSION_STARTED, on_session_started)
connection.on(RealtimeEvents.PARTIAL_TRANSCRIPT, on_partial_transcript)
connection.on(RealtimeEvents.COMMITTED_TRANSCRIPT, on_committed_transcript)
connection.on(RealtimeEvents.COMMITTED_TRANSCRIPT_WITH_TIMESTAMPS, on_committed_transcript_with_timestamps)
connection.on(RealtimeEvents.ERROR, on_error)
connection.on(RealtimeEvents.CLOSE, on_close)
print("Transcribing audio stream... (Press Ctrl+C to stop)")
try:
# Wait until error occurs or connection closes
await stop_event.wait()
except KeyboardInterrupt:
print("\nStopping transcription...")
finally:
await connection.close()
if __name__ == "__main__":
asyncio.run(main())
2

执行代码

python example.py

控制台将输出音频文件的部分转录和已提交转录。

后续步骤