서버 측 스트리밍

이 가이드에서는 ElevenLabs를 사용해 서버 측에서 오디오를 실시간으로 전사하는 방법을 안내합니다.

방법 가이드 · 텍스트 음성 변환 퀵스타트를 완료했다고 가정합니다.

개요

ElevenLabs Realtime Speech to Text API를 사용하면 Scribe Realtime v2 모델로 매우 낮은 지연 시간의 실시간 오디오 스트림 전사를 구현할 수 있습니다. 음성 어시스턴트, 전사 서비스 또는 실시간 음성 인식이 필요한 모든 애플리케이션을 구축할 때, 이 WebSocket 기반 API는 말하는 동안 부분 전사를 제공하고 음성 세그먼트가 완료되면 커밋된 전사를 제공합니다.

Scribe v2 Realtime은 서버 측에서 구현하여 URL, 파일 또는 자체 오디오 스트림의 오디오를 실시간으로 전사할 수 있습니다.

서버 측 구현은 클라이언트 측과 몇 가지 차이가 있습니다.

  • 일회용 토큰 대신 ElevenLabs API 키를 사용합니다.
  • 오디오를 수동으로 청크로 나눌 필요 없이 URL에서 직접 스트리밍을 지원합니다.

마이크에서 직접 오디오를 스트리밍하려면 클라이언트 측 스트리밍 가이드를 참조하세요.

퀵스타트

이 가이드는 API 키와 SDK를 설정했다고 가정합니다. 아직 설정하지 않았다면 먼저 퀵스타트를 완료하세요.

1

SDK 구성

SDK는 URL에서 스트리밍하거나 파일 또는 자체 오디오 스트림의 오디오를 수동으로 청크로 나누는 두 가지 실시간 오디오 전사 방식을 제공합니다.

API에서 지원하는 매개변수와 옵션의 전체 목록은 API 레퍼런스를 참조하세요.

이 예제에서는 공식 SDK를 사용해 URL에서 오디오 파일을 스트리밍하는 방법을 보여줍니다.

URL에서 스트리밍할 때는 ffmpeg 도구가 필요합니다. 설치 방법은 웹사이트를 방문하세요.

선택한 언어에 따라 example.py 또는 example.mts라는 새 파일을 만들고 다음 코드를 추가하세요.

from dotenv import load_dotenv
import os
import asyncio
from elevenlabs import ElevenLabs, RealtimeEvents, RealtimeUrlOptions
load_dotenv()
async def main():
elevenlabs = ElevenLabs(api_key=os.getenv("ELEVENLABS_API_KEY"))
# Create an event to signal when to stop
stop_event = asyncio.Event()
# Connect to a streaming audio URL
connection = await elevenlabs.speech_to_text.realtime.connect(RealtimeUrlOptions(
model_id="scribe_v2_realtime",
url="https://npr-ice.streamguys1.com/live.mp3",
include_timestamps=True,
))
# Set up event handlers
def on_session_started(data):
print(f"Session started: {data}")
def on_partial_transcript(data):
print(f"Partial: {data.get('text', '')}")
def on_committed_transcript(data):
print(f"Committed: {data.get('text', '')}")
# Committed transcripts with word-level timestamps. Only received when include_timestamps is set to True.
def on_committed_transcript_with_timestamps(data):
print(f"Committed with timestamps: {data.get('words', '')}")
# Errors - will catch all errors, both server and websocket specific errors
def on_error(error):
print(f"Error: {error}")
# Signal to stop on error
stop_event.set()
def on_close():
print("Connection closed")
# Register event handlers
connection.on(RealtimeEvents.SESSION_STARTED, on_session_started)
connection.on(RealtimeEvents.PARTIAL_TRANSCRIPT, on_partial_transcript)
connection.on(RealtimeEvents.COMMITTED_TRANSCRIPT, on_committed_transcript)
connection.on(RealtimeEvents.COMMITTED_TRANSCRIPT_WITH_TIMESTAMPS, on_committed_transcript_with_timestamps)
connection.on(RealtimeEvents.ERROR, on_error)
connection.on(RealtimeEvents.CLOSE, on_close)
print("Transcribing audio stream... (Press Ctrl+C to stop)")
try:
# Wait until error occurs or connection closes
await stop_event.wait()
except KeyboardInterrupt:
print("\nStopping transcription...")
finally:
await connection.close()
if __name__ == "__main__":
asyncio.run(main())
2

코드 실행

python example.py

콘솔에서 부분 전사와 커밋된 전사로 출력된 오디오 파일의 전사를 확인할 수 있습니다.

다음 단계