스트리밍

청크 전송 인코딩을 사용하여 ElevenLabs API에서 실시간 오디오를 스트리밍하는 방법을 알아보세요

ElevenLabs API는 일부 엔드포인트에서 실시간 오디오 스트리밍을 지원하며, 청크 전송 인코딩을 사용해 원시 오디오 바이트(예: MP3 데이터)를 HTTP로 직접 반환합니다. 따라서 생성되는 오디오를 클라이언트에서 점진적으로 처리하거나 재생할 수 있습니다.

공식 Node 및 Python 라이브러리에는 이 연속적인 오디오 스트림을 쉽게 처리할 수 있는 유틸리티가 포함되어 있습니다.

텍스트 음성 변환 API, 보이스 체인저 API 및 오디오 아이솔레이션 API에서 스트리밍을 지원합니다. 이 섹션에서는 텍스트 음성 변환 API 요청에서 스트리밍이 작동하는 방식을 중점적으로 설명합니다.

Python에서 스트리밍 요청은 다음과 같습니다.

from elevenlabs import stream
from elevenlabs.client import ElevenLabs
elevenlabs = ElevenLabs()
audio_stream = elevenlabs.text_to_speech.stream(
text="This is a test",
voice_id="JBFqnCBsd6RMkjVDRZzb",
model_id="eleven_multilingual_v2"
)
# option 1: play the streamed audio locally
stream(audio_stream)
# option 2: process the audio bytes manually
for chunk in audio_stream:
if isinstance(chunk, bytes):
print(chunk)

Node / Typescript에서 스트리밍 요청은 다음과 같습니다.

import { ElevenLabsClient, stream } from "@elevenlabs/elevenlabs-js";
import { Readable } from "stream";
const elevenlabs = new ElevenLabsClient();
async function main() {
const audioStream = await elevenlabs.textToSpeech.stream("JBFqnCBsd6RMkjVDRZzb", {
text: "This is a test",
modelId: "eleven_v4",
});
// option 1: play the streamed audio locally
await stream(Readable.from(audioStream));
// option 2: process the audio manually
for await (const chunk of audioStream) {
console.log(chunk);
}
}
main();