コンポジションプラン

構造化JSONによる音楽生成の精密な制御

コンポジションプランを使うと、音楽生成をきめ細かく制御できます。music_v2またはmusic_v2_5のプランは、順序付けられたチャンクのリストです。各チャンクは、独自のスタイル、歌詞、長さを持つ楽曲セクションを定義します。素早いプロトタイピングにはテキストプロンプトを、特定のチャンク構造、正確な歌詞のタイミング、複雑なアレンジが必要な場合にはコンポジションプランを使用してください。

コンポジションプランとテキストプロンプトは同時に使用できません。どちらか一方を使用してください。

{
"chunks": [
{
"text": "[Verse 1]\nWoke up today with a feeling inside\nSomething is changing I cannot hide\nThe sun on my face and the wind at my back\nI'm finally ready to get on track",
"duration_ms": 16000,
"positive_styles": [
"upbeat pop",
"female vocalist with clear tone",
"acoustic guitar and light synths",
"gentle and conversational vocals",
"light drums in background",
"polished production",
"120 BPM",
"C major"
],
"negative_styles": ["dark", "aggressive", "slow tempo", "a cappella"],
"context_adherence": "high"
},
{
"text": "[Verse 2]\nUsed to be scared of the world outside\nBuilding up walls where I used to hide\nBut now I see clearly what I need to do\nTake that first step into something new",
"duration_ms": 16000,
"positive_styles": [
"confident vocals",
"fuller guitar strumming",
"steady drum beat",
"bass joins in"
],
"negative_styles": ["a cappella", "sparse", "quiet"],
"context_adherence": "high"
},
{
"text": "[Pre-Chorus]\nNo more waiting for tomorrow\nThis is my time now",
"duration_ms": 8000,
"positive_styles": [
"building intensity",
"rising synth melody",
"driving drums",
"full band playing"
],
"negative_styles": ["a cappella", "dropping out"],
"context_adherence": "high"
},
{
"text": "[Chorus]\nI'm breaking through\nNothing's gonna stop me now\nI'm breaking through\nFinally found out how",
"duration_ms": 16000,
"positive_styles": [
"powerful and anthemic vocals",
"full band at maximum energy",
"punchy drums and bass",
"layered synths and guitar"
],
"negative_styles": ["a cappella", "minimal", "stripped back"],
"context_adherence": "high"
},
{
"text": "[Outro]",
"duration_ms": 8000,
"positive_styles": [
"instrumental fade out",
"guitar melody repeating",
"drums softening",
"gentle ending"
],
"negative_styles": ["vocals", "abrupt ending", "building"],
"context_adherence": "high"
}
]
}

チャンクベースのコンポジションプランにはmusic_v2またはmusic_v2_5が必要です。作曲時にmodel_id="music_v2_5"を渡してください。

クイックスタート

music quickstartに従ってAPIキーを設定し、SDKをインストールしたら、コンポジションプランでより細かく制御できます。

1

コンポジションプランで音楽を生成する

import os
from dotenv import load_dotenv
from elevenlabs.client import ElevenLabs
load_dotenv()
elevenlabs = ElevenLabs(api_key=os.environ.get("ELEVENLABS_API_KEY"))
composition_plan = {
"chunks": [
{
"text": "[Verse]\nWalking down an empty street\nWondering who I'll meet",
"duration_ms": 15000,
"positive_styles": ["pop", "upbeat", "female vocals", "soft vocals", "acoustic guitar"],
"negative_styles": ["dark", "slow"],
"context_adherence": "high"
},
{
"text": "[Chorus]\nThis is my moment\nI won't let it go",
"duration_ms": 15000,
"positive_styles": ["powerful vocals", "full band"],
"negative_styles": [],
"context_adherence": "high"
}
]
}
audio = elevenlabs.music.compose(
composition_plan=composition_plan,
model_id="music_v2_5",
# with_timestamps=True, # Optional: return word-level timestamps
)
with open("output.mp3", "wb") as f:
for chunk in audio:
f.write(chunk)
2

プロンプトからプランを生成する

テキストの説明からコンポジションプランを生成し、生成前に変更できます。

plan = elevenlabs.music.composition_plan.create(
prompt="An upbeat pop song about summer adventures",
music_length_ms=60000,
model_id="music_v2_5"
)
# Modify the generated plan
plan["chunks"][0]["text"] = "[Verse 1]\nCustom lyrics here"
audio = elevenlabs.music.compose(composition_plan=plan, model_id="music_v2_5")

構造リファレンス

チャンク

コンポジションプランは、最大30個のチャンクを順序付けたリストです。各チャンクは、textとスタイルに基づいて楽曲の1セクションを生成します。最初のチャンクが最も重要で、そのスタイルが楽曲全体の雰囲気とジャンルを決めます。

フィールド型説明
textstring角括弧内のセクション名([Verse 1])、歌詞の行、中括弧内のインライン指示({scratching})。
duration_msnumber長さ(ミリ秒、3,000~120,000)。
positive_stylesarray含めるスタイルと指示(最大50件)。
negative_stylesarray避けるスタイルと指示(最大50件)。デフォルトは空です。
context_adherencestringlow、medium、またはhigh(デフォルト)。チャンクが前後のチャンクにどれだけ忠実に従うかを指定します。

楽曲には最大30個のチャンクを含められます。合計時間は3秒から10分の範囲で、各チャンクは3秒から120秒の範囲にする必要があります。

最初のチャンクのスタイルは、全体の雰囲気とジャンルを決めるため、最も重要です。方向性が定まるまで、初期のチャンクには少なくとも6~7個のスタイルを指定してください。“great production quality”のような汎用的なスタイルは、リストに追加するデフォルトとして適しています。

チャンクは保存済み楽曲のオーディオを参照して、既存セクションを変更せずに維持したり、それをもとに新しいオーディオを条件付けたりすることもできます。既存の楽曲の編集や結合については、music inpaintingを参照してください。

歌詞を書く

textフィールドには、セクション名、歌詞、インライン指示を組み合わせます。

  • 角括弧内のセクション名:[Verse 1]、[Chorus]、[Bridge]
  • 改行(\n)で各行を区切ったプレーンテキストの歌詞
  • 括弧内の発音表現:(hmmm hmmm)、(ooh)、(yeah)
  • 中括弧内のインライン指示:{guitar solo}、{scratching}、{instrumental break}

短いインラインのキューには中括弧を使用してください。チャンク全体に適用されるジャンル、楽器編成、全体的なボーカルスタイルなどの幅広い特性には、代わりにpositive_stylesを使用してください。

{
"text": "[Verse]\n(soft female vocals) I've been waiting\n(instrumental break)\nfor you"
}

修正後の例では、全体的なボーカルスタイルをpositive_stylesに移し、短いインラインキューは括弧ではなく中括弧を使ってtext内に残しています。

スタイルのヒント

スタイル記述は具体的にしてください。

{
"positive_styles": [
"warm acoustic guitar with light fingerpicking",
"soft female vocals with intimate delivery",
"gentle percussion with brushed snare",
"80 BPM"
]
}

不要な音を防ぐため、ネガティブスタイルは積極的に使用してください。スタイルは英語で指定する必要があります(歌詞はどの言語でも使用できます)。

スタイルに著作権で保護されたコンテンツを含めると、APIは代替案を含むbad_composition_planエラーを返します。著作権で保護された素材の扱いを参照してください。

例

シネマティックなインストゥルメンタル

{
"chunks": [
{
"text": "[Tension Build]",
"duration_ms": 15000,
"positive_styles": [
"cinematic",
"orchestral",
"epic",
"low strings tremolo",
"building intensity",
"80 BPM",
"D minor"
],
"negative_styles": ["vocals", "lyrics", "pop", "electronic", "bright"],
"context_adherence": "high"
},
{
"text": "[Climax]",
"duration_ms": 15000,
"positive_styles": ["full orchestra", "brass fanfare", "triumphant"],
"negative_styles": ["quiet", "vocals"],
"context_adherence": "high"
},
{
"text": "[Resolution]",
"duration_ms": 10000,
"positive_styles": ["gentle strings", "piano melody", "fading out"],
"negative_styles": ["intense", "vocals"],
"context_adherence": "high"
}
]
}

ボイスオーバー付き広告

{
"chunks": [
{
"text": "[Intro]",
"duration_ms": 5000,
"positive_styles": [
"upbeat",
"modern pop",
"energetic",
"120 BPM",
"instrumental",
"catchy hook"
],
"negative_styles": ["sad", "slow", "dark", "vocals"],
"context_adherence": "high"
},
{
"text": "[Voiceover]\nIntroducing the future of productivity\nWork smarter, not harder",
"duration_ms": 10000,
"positive_styles": ["spoken voiceover", "confident male voice", "background music"],
"negative_styles": ["singing"],
"context_adherence": "high"
},
{
"text": "[Outro]",
"duration_ms": 5000,
"positive_styles": ["musical sting", "memorable"],
"negative_styles": ["vocals"],
"context_adherence": "high"
}
]
}

次のステップ