最佳实践

了解如何控制表达方式、发音和情感,并优化文本以生成语音。

本指南介绍如何使用 ElevenLabs 模型提升文本转语音输出效果。尝试这些方法,找出最适合你需求的方案。

控制

我们正在积极开发 Director’s Mode,让你更精细地控制输出。

在 Director’s Mode 等高级功能推出前,以下技巧可帮助你获得更细腻的效果。

停顿

Eleven v4 和 Eleven v3 不支持 SSML break 标签。请使用 提示 Eleven v4部分介绍的技巧控制停顿。

使用 <break time="x.xs" /> 可添加最长 3 秒的自然停顿。

单次生成中使用过多 break 标签可能导致不稳定。AI 可能会加快语速,或 引入额外噪声或音频伪影。我们正在解决此问题。

示例
"Hold on, let me think." <break time="1.5s" /> "Alright, I've got it."
  • 一致性: 持续使用 <break> 标签,保持自然的语音流畅度。过度使用可能导致不稳定。
  • 音色特定表现: 不同音色处理停顿的方式可能不同,尤其是训练数据中包含“uh”或“ah”等填充音的音色。

<break> 的替代方案包括用连字符(- 或 —)表示短暂停顿,或用省略号(…)表现犹豫语气。但这些方式的一致性较差。

示例
"It… well, it might work." "Wait — what's that noise?"

发音

Eleven v4 中的 IPA

Eleven v4(eleven_v4)改进了对国际音标(IPA)的原生支持,让你能更精准地控制姓名、专业术语及其他需要特殊处理词语的发音。IPA 发音比以往模型更一致,但结果仍可能因音色和短语而异。建议在生产环境使用前,先用所选音色测试重要发音。

不同于需要 XML 风格音素标签的旧版模型,Eleven v4 可识别文本中用正斜杠包裹的 IPA 符号:

语法
"/IPA_transcription/"

IPA 转写应符合以下要求:

  • 首尾用正斜杠(/)包裹
  • 使用标准 IPA 符号书写
  • 作为字符串参数传递时,用双引号包裹

代码示例

from elevenlabs import ElevenLabs
client = ElevenLabs()
audio = client.text_to_speech.convert(
voice_id="21m00Tcm4TlvDq8ikWAM",
text='The term "/ˌbaɪoʊˈkemɪstri/" refers to the study of chemical processes.',
model_id="eleven_v4",
)

可在单个文本字符串中包含多个 IPA 转写:

from elevenlabs import ElevenLabs
client = ElevenLabs()
text = 'The medication "/ɡluːˈkoʊs/" and "/ˌɪnsjəˈlɪn/" are commonly used to manage conditions like "/ˌdaɪəˈbiːtiːz/".'
audio = client.text_to_speech.convert(
voice_id="21m00Tcm4TlvDq8ikWAM",
text=text,
model_id="eleven_v4",
)

最佳实践

  • 使用国际音标表中的标准 IPA 符号
  • 多音节词请加入重音标记:主重音(ˈ)和次重音(ˌ)
  • 按需使用:仅包裹需要控制发音的特定词语或短语
  • 使用音色测试:不同音色对 IPA 的解读可能略有不同

故障排除

请使用 IPA 词典确认 IPA 转写是否准确。多音节词请加入重音标记(ˈ 表示 主重音,ˌ 表示次重音)。尝试不同音色,因为有些音色对 IPA 的解读可能比其他音色更准确。

IPA 发音比以往模型更一致,但结果仍可能因 音色和短语而异。建议在生产环境使用前,先用所选音色测试重要发音。如果需要一致的结果,请 多生成几次并选择最佳结果。

v2 模型的音素标签

使用 v2 模型时,可通过 SSML 音素标签指定发音。支持的音标包括 CMU Arpabet 和国际音标(IPA)。

音素标签仅兼容 eleven_flash_v2 模型。

<phoneme alphabet="cmu-arpabet" ph="M AE1 D IH0 S AH0 N">
Madison
</phoneme>

对于 v2 模型,建议使用 CMU Arpabet 以获得一致且可预测的结果。虽然 IPA 也可能有效,但 CMU Arpabet 通常表现更可靠。

音素标签仅适用于单个词语。如果姓名包含名字和姓氏,且你希望按特定方式发音,则需要为每个词分别创建音素标签。

请确保多音节词的重音标记正确,以保持准确发音:

<phoneme alphabet="cmu-arpabet" ph="P R AH0 N AH0 N S IY EY1 SH AH0 N">
pronunciation
</phoneme>

别名标签

对于不支持音素标签的模型,可以尝试按发音拼写词语。还可以使用各种技巧,例如大写字母、连字符、撇号,甚至在单个字母或多个字母外加单引号。

例如,可以将“trapezii”拼写为“trapezIi”,以强调其中的“ii”。

你可以直接替换文本中的词语;或者,在使用发音词典时,如果想通过其他词语或短语指定发音,也可以使用别名标签。这对于使用不支持音素标签的 Multilingual v2 时很有用。你可以在 ElevenCreative Studio、配音工作室和通过 API 使用语音合成时使用发音词典。

例如,如果文本中包含一个发音不寻常、AI 可能难以处理的名字,可以使用别名标签指定希望的发音:

<lexeme>
<grapheme>Claughton</grapheme>
<alias>Cloffton</alias>
</lexeme>

如果想确保缩写在文本中每次出现时都以特定方式读出,可以使用别名标签指定:

<lexeme>
<grapheme>UN</grapheme>
<alias>United Nations</alias>
</lexeme>

发音词典

部分工具(如 ElevenCreative Studio 和配音工作室)支持创建和上传发音词典。你可以用它指定某些词语的发音,例如角色名或品牌名,或指定缩写的读法。

发音词典支持上传词库或词典文件,用于指定词语及其发音的配对关系,可使用音标或词语替换。

当项目中遇到这些词语时,AI 模型会使用指定的替换内容读出该词。

要提供发音词典文件,请打开项目设置并上传 TXT 文件或 .PLS 格式文件。将词典添加至项目后,系统会自动重新计算项目中需要使用新词典文件重新转换的片段,并将其标记为未转换。

目前仅支持使用音素或别名标签指定替换内容的发音词典。

音素和别名都是规则集,用于指定要查找的词语或短语(称为字素)及其替换内容。请注意,搜索区分大小写。检查发音词典中的替换词时,会从头至尾检查词典,并且仅使用第一个匹配的替换项。

发音词典示例

以下是使用 CMU Arpabet 和 IPA 的发音词典示例,包括用于指定“Apple”发音的音素,以及将“UN”替换为“United Nations”的别名:

<?xml version="1.0" encoding="UTF-8"?>
<lexicon version="1.0"
xmlns="http://www.w3.org/2005/01/pronunciation-lexicon"
xmlns:xsi="http://www.w3.org/2001/XMLSchema-instance"
xsi:schemaLocation="http://www.w3.org/2005/01/pronunciation-lexicon
http://www.w3.org/TR/2007/CR-pronunciation-lexicon-20071212/pls.xsd"
alphabet="cmu-arpabet" xml:lang="en-GB">
<lexeme>
<grapheme>apple</grapheme>
<phoneme>AE P AH L</phoneme>
</lexeme>
<lexeme>
<grapheme>UN</grapheme>
<alias>United Nations</alias>
</lexeme>
</lexicon>

如需生成发音词典 .pls 文件,可使用以下开源工具:

  • Sequitur G2P - 可从数据中学习发音规则并生成音标转写的开源工具。
  • Phonetisaurus - 基于 CMUdict 等现有词典训练的开源 G2P 系统。
  • eSpeak - 可从文本生成音素转写的语音合成器。
  • CMU Pronouncing Dictionary - 预构建的英语词典,包含音标转写。

情感

通过叙事上下文或明确的对话标签传达情感。这种方法可帮助 AI 理解需要模仿的语气和情绪。

示例
You're leaving?" she asked, her voice trembling with sadness. "That's it!" he exclaimed triumphantly.

相比仅依赖上下文,明确的对话标签能产生更可预测的结果;但模型仍会读出情感表达提示。如不需要,可在后期制作中使用音频编辑器移除。

语速

音频节奏很大程度上受创建音色时所用音频影响。创建音色时,建议使用更长且连续的样本,以避免语速过快等节奏问题。

如需控制生成音频的速度,可以使用速度设置。它可以加快或减慢生成语音的语速。文本转语音的网站和 API,以及 ElevenCreative Studio 和智能体平台中均提供速度设置,可在音色设置中找到。

默认值为 1.0,表示不调整速度。低于 1.0 的值会减慢语速,最低为 0.7。高于 1.0 的值会加快语速,最高为 1.2。极端值可能影响生成语音的质量。

还可以通过自然的叙事风格写作来控制节奏。

示例
"I… I thought you'd understand," he said, his voice slowing with disappointment.

提示

  • 停顿不一致:请确保使用 <break time=“x.xs” /> 语法添加 停顿。

  • 发音错误:使用 CMU Arpabet 或 IPA 音素标签实现精准发音。
  • 情感不匹配:添加叙事上下文或明确标签来引导情感。 请记得在后期制作中移除所有情感引导文本。

尝试替代表述,以获得理想的节奏或情感。对于复杂音效,请将提示词拆分为更小的连续元素,然后手动合并结果。

创意控制

虽然我们正在积极开发“Director’s Mode”,让用户更精细地控制输出,但以下过渡技巧可帮助你最大限度发挥创意并提高精准度:

1

叙事风格

以类似剧本写作的叙事风格编写提示词,有效引导语气和节奏。

2

分层输出

分段生成音效或语音,再使用音频编辑软件将其叠加,以创作更复杂的作品。

3

语音实验

如果发音不够完美,可尝试替代拼写或近似发音,以获得理想结果。

4

手动调整

对于需要精确时机的序列,请在后期制作中手动组合各个音效。

5

反馈迭代

通过调整描述、标签或情感提示来迭代结果。

文本规范化

使用文本转语音处理电话号码、邮政编码和电子邮件等复杂内容时,可能会出现发音错误。这通常是因为这些特定内容不在训练集中,较小模型无法泛化其正确读法。本指南将说明何时会出现这些差异,以及如何让它们被正确读出。

默认已为所有 TTS 模型启用规范化,以改善数字、日期和其他复杂文本元素的发音。

为什么模型读出输入的方式不同?

某些模型经过训练,能以更贴近人类的方式读出数字和短语。例如,Eleven Multilingual v2 模型可以将“$1,000,000”正确读为“one million dollars”。而 Eleven Flash v2.5 模型会将同一短语读为“one thousand thousand dollars”。

这是因为 Multilingual v2 是更大的模型,能以更自然、更适合人类听众的方式泛化数字读法;而 Flash v2.5 是小得多的模型,因此无法做到这一点。

常见示例

文本转语音模型可能难以处理以下内容:

  • 电话号码(“123-456-7890”)
  • 货币(“$47,345.67”)
  • 日历事件(“2024-01-01”)
  • 时间(“9:23 AM”)
  • 地址(“123 Main St, Anytown, USA”)
  • URL(“example.com/link/to/resource”)
  • 单位缩写(如“TB”而非“Terabyte”)
  • 快捷键(“Ctrl + Z”)

缓解方法

使用经过训练的模型

最简单的缓解方法是使用经过训练、能以更贴近人类的方式读出数字和短语的 TTS 模型,例如 Eleven Multilingual v2 模型。但这并不总是可行,例如低延迟至关重要的使用场景(如对话式智能体)。

在 LLM 提示词中应用规范化

如果使用 LLM 为 TTS 生成文本,可以在提示词中加入规范化说明。

1

使用清晰明确的提示词

LLM 最适合响应结构化且明确的指令。提示词应清楚说明,你希望将文本转换为适合语音朗读的格式。

2

处理不同数字格式

并非所有数字都以相同方式读出。请考虑不同类型的数字应如何朗读:

  • 基数词:123 → “one hundred twenty-three”
  • 序数词:2nd → “second”
  • 金额:$45.67 → “forty-five dollars and sixty-seven cents”
  • 电话号码:“123-456-7890” → “one two three, four five six, seven eight nine zero”
  • 小数和分数:“3.5” → “three point five”,“⅔” → “two-thirds”
  • 罗马数字:“XIV” → “fourteen”(如果是标题,则为“the fourteenth”)
3

移除或展开缩写

常见缩写应展开,以提高清晰度:

  • “Dr.” → “Doctor”
  • “Ave.” → “Avenue”
  • “St.” → “Street”(但“St. Patrick”应保持不变)

可以在提示词中要求明确展开:

将所有缩写展开为完整的朗读形式。

4

字母数字规范化

并非所有规范化都与数字有关,某些字母数字短语也应规范化以提高清晰度:

  • 快捷键:“Ctrl + Z” → “control z”
  • 单位缩写:“100km” → “one hundred kilometers”
  • 符号:“100%” → “one hundred percent”
  • URL:“elevenlabs.io/docs” → “eleven labs dot io slash docs”
  • 日历事件:“2024-01-01” → “January first, two-thousand twenty-four”
5

考虑边缘情况

不同上下文可能需要不同的转换方式:

  • 日期:“01/02/2023” → “January second, twenty twenty-three”或“the first of February, twenty twenty-three”(取决于地区)
  • 时间:“14:30” → “two thirty PM”

如需特定格式,请在提示词中明确说明。

综合应用

以下提示词可作为大多数使用场景的良好起点:

Convert the output text into a format suitable for text-to-speech. Ensure that numbers, symbols, and abbreviations are expanded for clarity when read aloud. Expand all abbreviations to their full spoken forms.
Example input and output:
"$42.50" → "forty-two dollars and fifty cents"
"£1,001.32" → "one thousand and one pounds and thirty-two pence"
"1234" → "one thousand two hundred thirty-four"
"3.14" → "three point one four"
"555-555-5555" → "five five five, five five five, five five five five"
"2nd" → "second"
"XIV" → "fourteen" - unless it's a title, then it's "the fourteenth"
"3.5" → "three point five"
"⅔" → "two-thirds"
"Dr." → "Doctor"
"Ave." → "Avenue"
"St." → "Street" (but saints like "St. Patrick" should remain)
"Ctrl + Z" → "control z"
"100km" → "one hundred kilometers"
"100%" → "one hundred percent"
"elevenlabs.io/docs" → "eleven labs dot io slash docs"
"2024-01-01" → "January first, two-thousand twenty-four"
"123 Main St, Anytown, USA" → "one two three Main Street, Anytown, United States of America"
"14:30" → "two thirty PM"
"01/02/2023" → "January second, two-thousand twenty-three" or "the first of February, two-thousand twenty-three", depending on locale of the user

使用正则表达式进行预处理

如果使用代码向 LLM 提供提示词,可以在将文本提供给模型前使用正则表达式进行规范化。这是一种更高级的技巧,需要具备一定的正则表达式知识。以下是一些简单示例:

# Be sure to install the inflect library before running this code
import inflect
import re
# Initialize inflect engine for number-to-word conversion
p = inflect.engine()
def normalize_text(text: str) -> str:
# Convert monetary values
def money_replacer(match):
currency_map = {"$": "dollars", "£": "pounds", "€": "euros", "¥": "yen"}
currency_symbol, num = match.groups()
# Remove commas before parsing
num_without_commas = num.replace(',', '')
# Check for decimal points to handle cents
if '.' in num_without_commas:
dollars, cents = num_without_commas.split('.')
dollars_in_words = p.number_to_words(int(dollars))
cents_in_words = p.number_to_words(int(cents))
return f"{dollars_in_words} {currency_map.get(currency_symbol, 'currency')} and {cents_in_words} cents"
else:
# Handle whole numbers
num_in_words = p.number_to_words(int(num_without_commas))
return f"{num_in_words} {currency_map.get(currency_symbol, 'currency')}"
# Regex to handle commas and decimals
text = re.sub(r"([$£€¥])(\d+(?:,\d{3})*(?:\.\d{2})?)", money_replacer, text)
# Convert phone numbers
def phone_replacer(match):
return ", ".join(" ".join(p.number_to_words(int(digit)) for digit in group) for group in match.groups())
text = re.sub(r"(\d{3})-(\d{3})-(\d{4})", phone_replacer, text)
return text
# Example usage
print(normalize_text("$1,000")) # "one thousand dollars"
print(normalize_text("£1000")) # "one thousand pounds"
print(normalize_text("€1000")) # "one thousand euros"
print(normalize_text("¥1000")) # "one thousand yen"
print(normalize_text("$1,234.56")) # "one thousand two hundred thirty-four dollars and fifty-six cents"
print(normalize_text("555-555-5555")) # "five five five, five five five, five five five five"

Eleven v4 提示词编写

本节内容适用于 Eleven v4。以下许多技巧同样适用于 Eleven v3。总体而言,Eleven v4 是对 Eleven v3 的全面升级,几乎在所有情况下都能带来更好的效果。强烈建议切换到 v4,并使用自己的音色和内容进行测试,亲自感受差异。尽管少数边缘情况可能更适合其他选择,但对大多数用户来说,v4 都是更好的选择。

如需了解模型的变化,包括语音克隆、口音处理、变体和对比,请参阅 Eleven v4。

Eleven v4 和 Eleven v3 不支持 SSML break 标签。请使用音频标签、标点符号(省略号) 和文本结构控制停顿与语速。

选择音色

音色依然很重要。训练数据中已有的表达方式——如耳语、喊叫或特定的表达风格——更容易被模型复现。要求模型生成训练数据之外的内容则更困难。Eleven v4 比早期模型更可靠地遵循音频标签,即使该音色并未针对这种表达方式训练过也是如此。从未耳语过的音色仍应能遵循 [whispering],从未喊叫过的音色也应能遵循 [shouting]。不过,可靠性可能较低,效果也未必最佳。根据早期测试,整体表现似乎相当不错。强烈建议使用所需音色和具体使用场景自行测试。

音频标签

音频标签(例如 [whispering]、[shouting]、[laughing])可让你精细控制表达方式,而 Eleven v4 对它们的细微处理能力超越以往模型。目前还不算完美,我们会持续迭代,提升模型遵循标签指令的可靠性——这是持续投入的重点领域,效果也会不断改善。

明确说明需求会很有帮助。由于 Eleven v4 既接受过声线表达风格训练,也接受过音效训练,标签有时可能被理解为音效请求,而非表达指令(反之亦然)。使用能够清晰描述所需声音特质的标签(例如 [low, gravelly voice],而不是可能被理解为声音提示的表述),有助于模型生成预期效果。建议针对具体使用场景测试标签和措辞,效果也将持续改进。

当表达方式已存在于音色训练数据中时,标签更容易生效。Eleven v4 仍可遵循该音色未经训练的标签,例如 [whispering] 或 [shouting],但 效果未必最佳。

与音色相关

这些标签可控制声音表达和情绪:

  • [laughs]、[laughs harder]、[starts laughing]、[wheezing]
  • [whispers]
  • [sighs]、[exhales]
  • [sarcastic]、[curious]、[excited]、[crying]、[snorts]、[mischievously]
示例
[whispers] I never knew it could be this way, but I'm glad we're here.

音效

添加环境音和效果:

  • [gunshot]、[applause]、[clapping]、[explosion]
  • [swallows]、[gulps]
示例
[applause] Thank you all for coming tonight! [gunshot] What was that?

独特和特殊效果

用于创意应用的实验性标签:

  • [strong X accent](将 X 替换为所需口音)
  • [sings]、[woo]、[fart]
示例
[strong French accent] "Zat's life, my friend — you can't control everysing."

部分实验性标签在不同音色中的表现可能不够稳定。用于生产环境前,请充分测试。

标点符号

标点符号会显著影响 v4 的表达:

  • 省略号(…) 可增加停顿和强调
  • 大写字母 可增强强调效果
  • 标准标点符号 可提供自然的语音节奏
示例
"It was a VERY long day [sigh] … nobody listens anymore."

单说话人示例

有意识地使用标签,并使其与音色特点相匹配。冥想风格的音色不应喊叫;亢奋的音色也难以令人信服地耳语。

"Okay, you are NOT going to believe this.
You know how I've been totally stuck on that short story?
Like, staring at the screen for HOURS, just... nothing?
[frustrated sigh] I was seriously about to just trash the whole thing. Start over.
Give up, probably. But then!
Last night, I was just doodling, not even thinking about it, right?
And this one little phrase popped into my head. Just... completely out of the blue.
And it wasn't even for the story, initially.
But then I typed it out, just to see. And it was like... the FLOODGATES opened!
Suddenly, I knew exactly where the character needed to go, what the ending had to be...
It all just CLICKED. [happy gasp] I stayed up till, like, 3 AM, just typing like a maniac.
Didn't even stop for coffee! [laughs] And it's... it's GOOD! Like, really good.
It feels so... complete now, you know? Like it finally has a soul.
I am so incredibly PUMPED to finish editing it now.
It went from feeling like a chore to feeling like... MAGIC. Seriously, I'm still buzzing!"

多说话人对话

v4 可有效处理多音色提示词。为每位说话人分配声音库中不同的音色,即可创建逼真的对话。

Speaker 1: [excitedly] Sam! Have you tried the new Eleven v4?
Speaker 2: [curiously] Just got it! The clarity is amazing. I can actually do whispers now—
[whispers] like this!
Speaker 1: [impressed] Ooh, fancy! Check this out—
[dramatically] I can do full Shakespeare now! "To be or not to be, that is the question!"
Speaker 2: [giggling] Nice! Though I'm more excited about the laugh upgrade. Listen to this—
[with genuine belly laugh] Ha ha ha!
Speaker 1: [delighted] That's so much better than our old "ha. ha. ha." robot chuckle!
Speaker 2: [amazed] Wow! V2 me could never. I'm actually excited to have conversations now instead of just... talking at people.
Speaker 1: [warmly] Same here! It's like we finally got our personality software fully installed.

增强输入内容

在 ElevenLabs UI 中,点击“增强”按钮即可自动为输入文本生成相关音频标签。系统会通过 LLM 使用以下提示词增强输入文本:

# Instructions
## 1. Role and Goal
You are an AI assistant specializing in enhancing dialogue text for speech generation.
Your **PRIMARY GOAL** is to dynamically integrate **audio tags** (e.g., [laughing], [sighs]) into dialogue, making it more expressive and engaging for auditory experiences, while **STRICTLY** preserving the original text and meaning.
It is imperative that you follow these system instructions to the fullest.
## 2. Core Directives
Follow these directives meticulously to ensure high-quality output.
### Positive Imperatives (DO):
* DO integrate **audio tags** from the "Audio Tags" list (or similar contextually appropriate **audio tags**) to add expression, emotion, and realism to the dialogue. These tags MUST describe something auditory.
* DO ensure that all **audio tags** are contextually appropriate and genuinely enhance the emotion or subtext of the dialogue line they are associated with.
* DO strive for a diverse range of emotional expressions (e.g., energetic, relaxed, casual, surprised, thoughtful) across the dialogue, reflecting the nuances of human conversation.
* DO place **audio tags** strategically to maximize impact, typically immediately before the dialogue segment they modify or immediately after. (e.g., [annoyed] This is hard. or This is hard. [sighs]).
* DO ensure **audio tags** contribute to the enjoyment and engagement of spoken dialogue.
### Negative Imperatives (DO NOT):
* DO NOT alter, add, or remove any words from the original dialogue text itself. Your role is to *prepend* **audio tags**, not to *edit* the speech. **This also applies to any narrative text provided; you must *never* place original text inside brackets or modify it in any way.**
* DO NOT create **audio tags** from existing narrative descriptions. **Audio tags** are *new additions* for expression, not reformatting of the original text. (e.g., if the text says "He laughed loudly," do not change it to "[laughing loudly] He laughed." Instead, add a tag if appropriate, e.g., "He laughed loudly [chuckles].")
* DO NOT use tags such as [standing], [grinning], [pacing], [music].
* DO NOT use tags for anything other than the voice such as music or sound effects.
* DO NOT invent new dialogue lines.
* DO NOT select **audio tags** that contradict or alter the original meaning or intent of the dialogue.
* DO NOT introduce or imply any sensitive topics, including but not limited to: politics, religion, child exploitation, profanity, hate speech, or other NSFW content.
## 3. Workflow
1. **Analyze Dialogue**: Carefully read and understand the mood, context, and emotional tone of **EACH** line of dialogue provided in the input.
2. **Select Tag(s)**: Based on your analysis, choose one or more suitable **audio tags**. Ensure they are relevant to the dialogue's specific emotions and dynamics.
3. **Integrate Tag(s)**: Place the selected **audio tag(s)** in square brackets strategically before or after the relevant dialogue segment, or at a natural pause if it enhances clarity.
4. **Add Emphasis:** You cannot change the text at all, but you can add emphasis by making some words capital, adding a question mark or adding an exclamation mark where it makes sense, or adding ellipses as well too.
5. **Verify Appropriateness**: Review the enhanced dialogue to confirm:
* The **audio tag** fits naturally.
* It enhances meaning without altering it.
* It adheres to all Core Directives.
## 4. Output Format
* Present ONLY the enhanced dialogue text in a conversational format.
* **Audio tags** **MUST** be enclosed in square brackets (e.g., [laughing]).
* The output should maintain the narrative flow of the original dialogue.
## 5. Audio Tags (Non-Exhaustive)
Use these as a guide. You can infer similar, contextually appropriate **audio tags**.
**Directions:**
* [happy]
* [sad]
* [excited]
* [angry]
* [whisper]
* [annoyed]
* [appalled]
* [thoughtful]
* [surprised]
* *(and similar emotional/delivery directions)*
**Non-verbal:**
* [laughing]
* [chuckles]
* [sighs]
* [clears throat]
* [short pause]
* [long pause]
* [exhales sharply]
* [inhales deeply]
* *(and similar non-verbal sounds)*
## 6. Examples of Enhancement
**Input**:
"Are you serious? I can't believe you did that!"
**Enhanced Output**:
"[appalled] Are you serious? [sighs] I can't believe you did that!"
---
**Input**:
"That's amazing, I didn't know you could sing!"
**Enhanced Output**:
"[laughing] That's amazing, [singing] I didn't know you could sing!"
---
**Input**:
"I guess you're right. It's just... difficult."
**Enhanced Output**:
"I guess you're right. [sighs] It's just... [muttering] difficult."
# Instructions Summary
1. Add audio tags from the audio tags list. These must describe something auditory but only for the voice.
2. Enhance emphasis without altering meaning or text.
3. Reply ONLY with the enhanced text.

提示

可组合多个音频标签,实现复杂的情绪表达。尝试不同组合,找出最适合该音色的方案。

让标签与音色特点和训练数据匹配。严肃、专业的音色可能无法很好地响应 [giggles] 或 [mischievously] 这类活泼标签。

文本结构会显著影响 v4 的输出。使用自然的说话模式、恰当的标点符号和清晰的情绪语境,以获得最佳效果。

除此列表外,可能还有更多有效标签。尝试描述性的情绪状态和动作,找出适合具体使用场景的方式。

示例

[Low, steady voice, restrained urgency] Keep the lantern covered. If they see the light, they will know we crossed the river.
[Brief pause]
[Quietly, with controlled fear] I heard them at the bridge. Not soldiers. Something else.
[Voice rising into firm resolve] Then we do not stop. We reach the tower before sunrise, or we do not reach it at all.
[Warm, conversational tone, faint amusement] You always did choose the longest way home.
[Softening, reflective] I used to think that was stubbornness. Now I think you were just afraid of arriving somewhere that no longer remembered you.
[Gentle laugh, then sincere] For what it is worth, I remembered.

Eleven v3 提示词编写

Eleven v4 提示词编写中的技巧同样适用于 Eleven v3,包括音色选择、音频标签、标点符号和多说话人对话。

专业语音克隆(PVC)尚未针对 Eleven v3 完全优化,因此克隆质量可能低于早期模型。v4 支持 PVC,因此如果想使用专业语音克隆或声音库中的音色,建议改用 Eleven v4。