サーバーサイドストリーミング

このガイドでは、ElevenLabsを使用してサーバーサイドで音声をリアルタイムに文字起こしする方法を紹介します。

ハウツーガイド · スピーチtoテキスト クイックスタートを完了していることを前提としています。

概要

ElevenLabs Realtime Speech to Text APIでは、Scribe Realtime v2モデルを使用して、超低レイテンシーで音声ストリームをリアルタイムに文字起こしできます。音声アシスタント、文字起こしサービス、ライブ音声認識が必要なあらゆるアプリケーションの構築において、このWebSocketベースのAPIは、話している間は部分的な文字起こしを、音声セグメントが完了すると確定した文字起こしを提供します。

Scribe v2 Realtimeは、URL、ファイル、または独自の音声ストリームを介して、サーバーサイドでリアルタイムに音声を文字起こしできます。

サーバーサイドの実装は、クライアントサイドとはいくつかの点で異なります。

  • 単回使用トークンではなくElevenLabs APIキーを使用します。
  • 音声を手動でチャンク化することなく、URLから直接ストリーミングできます。

マイクから直接音声をストリーミングする場合は、クライアントサイドストリーミングガイドを参照してください。

クイックスタート

このガイドでは、APIキーとSDKのセットアップが完了していることを前提としています。まだの場合は、先にクイックスタートを完了してください。

1

SDKを設定する

SDKでは、リアルタイムでオーディオを書き起こす方法を2つ提供しています。URLからストリーミングする方法と、ファイルまたは独自のオーディオストリームからオーディオを手動でチャンク化する方法です。

APIでサポートされているパラメーターとオプションの一覧は、APIリファレンスをご覧ください。

この例では、公式SDKを使ってURLからオーディオファイルをストリーミングする方法を示します。

URLからストリーミングするには、ffmpegツールが必要です。インストール手順は公式サイトをご覧ください。

使用する言語に応じて、example.pyまたはexample.mtsという新しいファイルを作成し、次のコードを追加します。

from dotenv import load_dotenv
import os
import asyncio
from elevenlabs import ElevenLabs, RealtimeEvents, RealtimeUrlOptions
load_dotenv()
async def main():
elevenlabs = ElevenLabs(api_key=os.getenv("ELEVENLABS_API_KEY"))
# Create an event to signal when to stop
stop_event = asyncio.Event()
# Connect to a streaming audio URL
connection = await elevenlabs.speech_to_text.realtime.connect(RealtimeUrlOptions(
model_id="scribe_v2_realtime",
url="https://npr-ice.streamguys1.com/live.mp3",
include_timestamps=True,
))
# Set up event handlers
def on_session_started(data):
print(f"Session started: {data}")
def on_partial_transcript(data):
print(f"Partial: {data.get('text', '')}")
def on_committed_transcript(data):
print(f"Committed: {data.get('text', '')}")
# Committed transcripts with word-level timestamps. Only received when include_timestamps is set to True.
def on_committed_transcript_with_timestamps(data):
print(f"Committed with timestamps: {data.get('words', '')}")
# Errors - will catch all errors, both server and websocket specific errors
def on_error(error):
print(f"Error: {error}")
# Signal to stop on error
stop_event.set()
def on_close():
print("Connection closed")
# Register event handlers
connection.on(RealtimeEvents.SESSION_STARTED, on_session_started)
connection.on(RealtimeEvents.PARTIAL_TRANSCRIPT, on_partial_transcript)
connection.on(RealtimeEvents.COMMITTED_TRANSCRIPT, on_committed_transcript)
connection.on(RealtimeEvents.COMMITTED_TRANSCRIPT_WITH_TIMESTAMPS, on_committed_transcript_with_timestamps)
connection.on(RealtimeEvents.ERROR, on_error)
connection.on(RealtimeEvents.CLOSE, on_close)
print("Transcribing audio stream... (Press Ctrl+C to stop)")
try:
# Wait until error occurs or connection closes
await stop_event.wait()
except KeyboardInterrupt:
print("\nStopping transcription...")
finally:
await connection.close()
if __name__ == "__main__":
asyncio.run(main())
2

コードを実行する

python example.py

オーディオファイルの書き起こしが、部分的な文字起こしと確定した文字起こしとしてコンソールに表示されます。

次のステップ