编辑转录文本

本指南介绍如何使用 Speech to Text API,通过自然语言指令编辑转录文本。

操作指南 · 假设你已完成 Speech to Text 快速入门。

概述

编辑转录文本是一项实验性功能,除基础转录费用外还需额外支付 30% 的费用;每次请求至少按 10 秒音频计费。详见 API 定价页面 的价格说明。

编辑转录文本功能可让你在转录请求中附加自然语言指令。音频转录完成后,系统会将指令应用于转录文本,并同时返回编辑后的文本和原始文本。

这样就无需在自己的流程中单独进行后处理。常见用途包括统一日期、时间或单位的写法,展开缩写,删除不需要的内容,应用不同的风格或语气,或重新格式化转录文本。

例如,转录一段语音留言时使用指令 Write all dates in ISO 8601 format (YYYY-MM-DD),会返回两个版本的文本:

{
"language_code": "eng",
"language_probability": 0.9912,
"text": "Hi, this is Jill. Your appointment is confirmed for the twelfth of July twenty twenty-six, and the follow-up is on the third of August.",
"words": [
{ "text": "Hi,", "start": 0.12, "end": 0.38, "type": "word", "logprob": 0.0 },
{ "text": " ", "start": 0.38, "end": 0.41, "type": "spacing", "logprob": 0.0 },
{ "text": "this", "start": 0.41, "end": 0.55, "type": "word", "logprob": 0.0 },
...
],
"transcription_id": "Y2ZX8AxHUzTPCIualYiE",
"edited_transcript": {
"kind": "transcript",
"text": "Hi, this is Jill. Your appointment is confirmed for 2026-07-12, and the follow-up is on 2026-08-03."
}
}

text 和 words 字段始终描述原始转录文本。编辑后的版本会单独在 edited_transcript 中返回。

集成转录文本编辑

要在 Speech to Text API 中使用转录文本编辑功能,请向 convert 方法传入 transcript_edit 参数。指令最长可包含 2000 个字符。

import os
from dotenv import load_dotenv
from elevenlabs.client import ElevenLabs
load_dotenv()
elevenlabs = ElevenLabs(
api_key=os.getenv("ELEVENLABS_API_KEY"),
)
with open("voicemail.mp3", "rb") as audio_file:
transcription = elevenlabs.speech_to_text.convert(
file=audio_file,
model_id="scribe_v2",
# Natural-language instruction applied to the finished transcript.
transcript_edit="Write all dates in ISO 8601 format (YYYY-MM-DD)",
)
print("Original:", transcription.text)
print("Edited:", transcription.edited_transcript)

转录文本编辑同时支持同步请求和 webhook 请求。对于 webhook 请求,edited_transcript 会包含在 webhook 负载的 transcription 对象中。

Scribe v2 Realtime 支持对每份已确认的转录文本使用相同指令。请参阅 实时转录文本编辑指南。

编写指令

指令会应用于整份转录文本中所有相关内容,所有匹配项都会被编辑,而不只是第一个。指令可以修改特定字词或短语,同时保持其他内容不变;也可以重写或标注完整文本。一条指令可组合多项编辑。

以下示例说明支持的指令范围:

指令效果
Write all dates in ISO 8601 format (YYYY-MM-DD)the twelfth of July twenty twenty-six 会变为 2026-07-12
Write times in 24-hour formathalf past two in the afternoon 会变为 14:30
Expand abbreviations such as "ETA" and "ASAP" on first useETA 会变为 estimated time of arrival (ETA)
Add the sentiment of every sentence in brackets at the end. Choose from [positive, negative, neutral]Thanks, that was really helpful. 会变为 Thanks, that was really helpful. [positive]
Redact every curse word with its first letter followed by stars, e.g. s***That was a damn good call. 会变为 That was a d*** good call.
Format the transcript as a bulleted list, one sentence per bullet将整份转录文本重新格式化

有些调整可通过专用参数实现,比编辑指令更便宜且结果更可预测。使用 关键术语 提示 让识别更倾向于特定名称和术语,使用 no_verbatim 删除填充词和 语句不流畅部分,并使用 numbers_format(在支持时)选择数字或文字。仅将转录文本编辑用于这些选项无法覆盖的更改。

编写指令时请注意:

  • 明确说明所需输出。Write all dates in ISO 8601 format (YYYY-MM-DD) 比 fix the dates 更可靠。
  • 指令可以使用任何语言编写,但英文指令效果最佳。编辑后的转录文本会保持原始语言。
  • 系统只会执行指令。音频中的语音内容会被视为数据,无法改变指令的应用方式。
  • 如果转录文本中没有内容受指令影响,编辑后的转录文本将与原始文本相同。

响应格式

设置 transcript_edit 后,响应会包含 edited_transcript 对象。其 kind 字段表示编辑是否成功:

kind字段说明
transcripttext编辑后的转录文本。未进行编辑时与原始 text 相同。
errorerror_type (edit_failed)、message无法生成编辑结果。转录本身已成功,仍会返回原始 text。

未请求 transcript_edit 时,该字段不存在。

编辑失败
{
"text": "Hi, this is Jill. Your appointment is confirmed for ...",
"edited_transcript": {
"kind": "error",
"error_type": "edit_failed",
"message": "The transcript could not be edited. Please try again."
}
}

需要注意的行为:

  • 编辑后的转录文本为纯文本。词级时间戳、说话人标签和 additional_formats 仍描述原始转录文本。
  • 编辑会在转录完成后执行,因此会增加延迟,且延迟会随转录文本长度增长。
  • 附加费用按请求的音频时长计算,每次请求至少按 10 秒计费。

转录文本编辑无法与 entity_detection、entity_redaction 或 use_multi_channel 结合使用。同时使用这些功能的请求会因参数无效而被拒绝。

后续步骤