文本转语音快速入门

了解如何将语音音频转换为文本。

本指南将介绍如何使用文本转语音 API 将语音音频转换为文本。

使用 ElevenLabs speech-to-text skill,通过 AI 编程助手转录音频:

npx skills add elevenlabs/skills --skill speech-to-text

本教程将演示如何使用批量文本转语音 API。如需了解如何使用实时文本转语音 API,请参阅客户端 流式传输或服务端 流式传输指南。

使用文本转语音 API

1

创建 API 密钥

在控制台创建 API 密钥,用于安全地访问 API。

将密钥存储为托管密钥,并根据需要通过 .env 文件将其作为环境变量传递给 SDK,或直接在应用配置中设置。

.env
ELEVENLABS_API_KEY=<your_api_key_here>
2

安装 SDK

我们还会使用 dotenv 库从环境变量加载 API 密钥。

pip install elevenlabs
pip install python-dotenv
3

发起 API 请求

根据所选语言,新建名为 example.py 或 example.mts 的文件,然后添加以下代码:

# example.py
import os
from dotenv import load_dotenv
from io import BytesIO
import requests
from elevenlabs.client import ElevenLabs
load_dotenv()
elevenlabs = ElevenLabs(
api_key=os.getenv("ELEVENLABS_API_KEY"),
)
audio_url = (
"https://storage.googleapis.com/eleven-public-cdn/audio/marketing/nicole.mp3"
)
response = requests.get(audio_url)
audio_data = BytesIO(response.content)
transcription = elevenlabs.speech_to_text.convert(
file=audio_data,
model_id="scribe_v2", # Model to use
tag_audio_events=True, # Tag audio events like laughter, applause, etc.
language_code="eng", # Language of the audio file. If set to None, the model will detect the language automatically.
diarize=True, # Whether to annotate who is speaking
)
print(transcription)

然后运行:

python example.py

控制台中将显示音频文件的转录文本。

对于医疗和临床音频,请将 model_id 设置为 scribe_v2_medical。请求格式 与 scribe_v2 相同,收费标准也相同。请参阅 Scribe v2 Medical。

后续步骤