拼接多个请求

本指南介绍如何在多个文本片段或生成任务之间保持语音韵律。

操作指南 · 假设你已完成 ElevenAPI 快速入门。

将大量文本转换为音频时,不同片段之间的韵律可能会突然变化。转换跨越多个段落或章节的文本时,这一点尤其明显。要在多个片段中保持语音韵律,可以使用请求拼接功能。

此功能可让你提供已生成内容及后续将生成内容的上下文,帮助整篇文本保持一致的音色和韵律。

请求拼接不适用于 eleven_v3 模型。

以下是未使用请求拼接的示例:

以下是使用请求拼接的相同示例:

如何使用请求拼接

使用 ElevenLabs SDK 时,请求拼接最简单。

本指南假设你已设置 API 密钥和 SDK。如果尚未完成,请先完成 快速入门。

1

拼接多个请求

根据所用语言,新建名为 example.py 或 example.mts 的文件,并添加以下代码:

import os
from io import BytesIO
from elevenlabs.client import ElevenLabs
from elevenlabs.play import play
from dotenv import load_dotenv
load_dotenv()
ELEVENLABS_API_KEY = os.getenv("ELEVENLABS_API_KEY")
elevenlabs = ElevenLabs(
api_key=ELEVENLABS_API_KEY,
)
paragraphs = [
"The advent of technology has transformed countless sectors, with education ",
"standing out as one of the most significantly impacted fields.",
"In recent years, educational technology, or EdTech, has revolutionized the way ",
"teachers deliver instruction and students absorb information.",
"From interactive whiteboards to individual tablets loaded with educational software, ",
"technology has opened up new avenues for learning that were previously unimaginable.",
"One of the primary benefits of technology in education is the accessibility it provides.",
]
request_ids = []
audio_buffers = []
for paragraph in paragraphs:
# Usually we get back a stream from the convert function, but with_raw_response is
# used to get the headers from the response
with elevenlabs.text_to_speech.with_raw_response.convert(
text=paragraph,
voice_id="T7QGPtToiqH4S8VlIkMJ",
model_id="eleven_v4",
previous_request_ids=request_ids
) as response:
request_ids.append(response._response.headers.get("request-id"))
# response._response.headers also contains useful information like 'character-cost',
# which shows the cost of the generation in characters.
audio_data = b''.join(chunk for chunk in response.data)
audio_buffers.append(BytesIO(audio_data))
combined_stream = BytesIO(b''.join(buffer.getvalue() for buffer in audio_buffers))
play(combined_stream)
2

执行代码

python example.py

应能听到播放拼接后的完整音频。

常见问题

若要使用先前请求的请求 ID 进行条件控制,该请求必须已完全处理完毕。对于流式传输,这意味着必须从响应正文中完整读取音频。

效果取决于所使用的模型、音色和音色设置。

请求 ID 的时间不应超过 2 小时。

是,除非你是有更高隐私要求的企业版用户。

后续步骤