音楽インペインティング

既存の楽曲セクションを編集・結合する

music_v2またはmusic_v2_5モデルによる音楽インペインティングでは、曲の特定部分を変更しながら、残りの部分はそのまま維持できます。生成した曲を保存してから、コンポジションプラン内でその一部を参照することで、変更せずに維持したり、再生成したり、元のオーディオをもとに新しいオーディオを条件付けしたりできます。

仕組み

music_v2またはmusic_v2_5のコンポジションプランは、チャンクを順番に並べたリストです。チャンクには次の2種類があります。

  • 生成チャンク:textとスタイルに基づいて新しいオーディオを生成します。セクションの再生成や新しい素材の追加に使用します。
  • オーディオ参照チャンク:保存された曲の一部を変更せずに挿入します。既存の曲のセクションをそのまま維持するために使用します。

クイックスタート

1

インペインティング用に曲を保存する

インペインティング用に曲を保存する方法は2つあります。store_for_inpaintingを使用して新しい曲を生成するか、既存のオーディオファイルをアップロードします。

import os
from dotenv import load_dotenv
from elevenlabs.client import ElevenLabs
load_dotenv()
elevenlabs = ElevenLabs(api_key=os.environ.get("ELEVENLABS_API_KEY"))
# Generate a song and store it for later inpainting
response = elevenlabs.music.compose_detailed(
prompt="An upbeat pop song with verse and chorus",
music_length_ms=60000,
model_id="music_v2_5",
store_for_inpainting=True
)
song_id = response.song_id
# Save the audio
with open("original.mp3", "wb") as f:
f.write(response.audio)
2

チャンクを維持・再生成する

オーディオ参照チャンク(維持)と生成チャンク(再生成)を組み合わせたプランを作成し、model_id="music_v2_5"を指定してcomposeに渡します。

# Keep the first 30 seconds, regenerate the rest with a new style
composition_plan = {
"chunks": [
# Keep the original first 30 seconds unchanged
{
"song_id": song_id,
"range": {"start_ms": 0, "end_ms": 30000}
},
# Regenerate the chorus with a new style
{
"text": "[Chorus]\nWe're rising up tonight\nNothing can stop us now",
"duration_ms": 30000,
"positive_styles": ["bigger drums", "layered vocals", "anthemic"],
"negative_styles": ["sparse", "minimal"],
"context_adherence": "high"
}
]
}
audio = elevenlabs.music.compose(
composition_plan=composition_plan,
model_id="music_v2_5",
)
with open("edited.mp3", "wb") as f:
for chunk in audio:
f.write(chunk)

条件付け

生成チャンクは、conditioning_refを使用して保存済みオーディオの一部を条件として指定できます。モデルは参照元の音楽的特性に近づけながら、チャンクを再生成します。condition_strength(low、medium、high、またはxhigh)で、チャンクが参照にどの程度厳密に従うかを制御します。

{
"text": "[Chorus]\nThis is my moment\nI won't let it go",
"duration_ms": 15000,
"positive_styles": ["powerful vocals", "full band", "anthemic"],
"negative_styles": [],
"context_adherence": "high",
"conditioning_ref": {
"song_id": "vVtPM1Sas70E2LIhQFch",
"range": { "start_ms": 30000, "end_ms": 45000 }
},
"condition_strength": "high"
}

最初のチャンクは、それ以降のすべてのチャンクの生成に影響します。曲全体を参照に基づいて 条件付けするには、最初のチャンクからconditioning_refを適用してください。

条件付け参照の長さは最大30秒(30,000ms)です。

コンテキスト追従性

各生成チャンクにはcontext_adherenceレベルがあり、隣接するチャンクにどの程度厳密に従うかを制御します。

  • high(デフォルト):周囲のチャンクとの一貫性を維持します。維持したオーディオと再生成したオーディオ間を滑らかに遷移させるために使用します。
  • medium:一貫性と創造性の自由度のバランスを取ります。
  • low:チャンクをコンテキストからより自由に逸脱させ、創造性を高めます。

例

1つのセクションを編集する

映画予告編を生成し、アウトロだけを別の歌詞で再生成します。

1

元の曲を生成する

composition_plan = {
"chunks": [
{
"text": "[Intro]\nIn a world beyond code\nWhere sound becomes life",
"duration_ms": 15000,
"positive_styles": ["cinematic", "epic", "orchestral", "low strings", "suspenseful"],
"negative_styles": ["acoustic", "pop", "minimalistic"],
"context_adherence": "high"
},
{
"text": "[Build]\nTechnology awakens the future\nShaping every word into power",
"duration_ms": 20000,
"positive_styles": ["rising brass", "full orchestra", "epic"],
"negative_styles": ["acoustic", "pop"],
"context_adherence": "high"
},
{
"text": "[Bridge]\n(ah ah ah ah)",
"duration_ms": 15000,
"positive_styles": ["ethereal choir", "crescendo"],
"negative_styles": [],
"context_adherence": "high"
},
{
"text": "[Outro]\nThe voice of tomorrow, unleashed\nElevenLabs",
"duration_ms": 10000,
"positive_styles": ["deep narration", "epic finale"],
"negative_styles": [],
"context_adherence": "high"
}
]
}
response = elevenlabs.music.compose_detailed(
composition_plan=composition_plan,
model_id="music_v2_5",
store_for_inpainting=True
)
song_id = response.song_id
2

アウトロだけを編集する

最初の3セクション(0~50秒)を1つのオーディオ参照チャンクで保持し、アウトロを再生成します。

edited_plan = {
"chunks": [
# Keep the intro, build, and bridge unchanged
{
"song_id": song_id,
"range": {"start_ms": 0, "end_ms": 50000}
},
# Regenerate the outro with new lyrics
{
"text": "[Outro]\nThe future has arrived\nElevenLabs",
"duration_ms": 10000,
"positive_styles": ["deep narration", "epic finale"],
"negative_styles": [],
"context_adherence": "high"
}
]
}
audio = elevenlabs.music.compose(composition_plan=edited_plan, model_id="music_v2_5")

曲を延長する

既存の曲に新しいイントロとアウトロを追加します。

1

元の曲を生成する

response = elevenlabs.music.compose_detailed(
prompt="Berlin night club techno",
music_length_ms=60000,
model_id="music_v2_5",
store_for_inpainting=True
)
song_id = response.song_id
2

新しいイントロとアウトロで延長する

元の曲から保持するスライスを、2つの新しい生成チャンクで挟みます。

extend_plan = {
"chunks": [
# New intro
{
"text": "[Intro]",
"duration_ms": 30000,
"positive_styles": ["techno", "building tension", "filtered synths"],
"negative_styles": [],
"context_adherence": "high"
},
# Keep the core of the original (seconds 10-50)
{
"song_id": song_id,
"range": {"start_ms": 10000, "end_ms": 50000}
},
# New outro
{
"text": "[Outro]",
"duration_ms": 30000,
"positive_styles": ["techno", "fading out", "sparse"],
"negative_styles": [],
"context_adherence": "high"
}
]
}
audio = elevenlabs.music.compose(composition_plan=extend_plan, model_id="music_v2_5")

シームレスなループを作成する

音楽フレーズを生成し、同じスライスを2回繰り返す間をつなぐ「グルー」チャンクを使ってループを作成します。

1

短いクリップを生成する

composition_plan = {
"chunks": [
{
"text": "[Solo Acoustic]",
"duration_ms": 10000,
"positive_styles": ["acoustic guitar", "fingerpicking", "warm tone", "soft dynamics"],
"negative_styles": ["electric", "drums", "electronic"],
"context_adherence": "high"
}
]
}
response = elevenlabs.music.compose_detailed(
composition_plan=composition_plan,
model_id="music_v2_5",
store_for_inpainting=True
)
song_id = response.song_id
2

グルーチャンクでループを作成する

loop_plan = {
"chunks": [
# Loop start - keep a slice of the original
{
"song_id": song_id,
"range": {"start_ms": 3000, "end_ms": 8000}
},
# Glue - generate a smooth transition between the two slices
{
"text": "[Glue]",
"duration_ms": 3000,
"positive_styles": ["acoustic guitar", "fingerpicking", "smooth transition"],
"negative_styles": [],
"context_adherence": "high"
},
# Loop end - keep the same slice again
{
"song_id": song_id,
"range": {"start_ms": 3000, "end_ms": 8000}
}
]
}
audio = elevenlabs.music.compose(composition_plan=loop_plan, model_id="music_v2_5")

類似した曲を生成する

既存の曲の短いスライスをもとに、音楽的な特徴を引き継いだまったく新しい曲を生成します。最初のチャンクは後続するすべてのチャンクに影響するため、最初のチャンクにconditioning_refを適用すると、保存済みオーディオを参照するのがそのチャンクだけであっても、生成全体の方向性が決まります。

1

元の曲を生成する

response = elevenlabs.music.compose_detailed(
prompt="An upbeat pop song with bright synths and driving drums",
music_length_ms=60000,
model_id="music_v2_5",
store_for_inpainting=True
)
song_id = response.song_id
2

元の曲を条件として新しい曲を生成する

similar_plan = {
"chunks": [
# The first chunk conditions the whole song on the reference
{
"text": "[Verse]\nSalt on my skin from a borrowed sea\nCounting the heartbeats it takes to break free",
"duration_ms": 30000,
"positive_styles": ["pop", "energetic", "bright synths", "driving drums"],
"negative_styles": ["sparse", "minimal"],
"context_adherence": "high",
"conditioning_ref": {
"song_id": song_id,
"range": {"start_ms": 0, "end_ms": 10000}
},
"condition_strength": "high"
},
{
"text": "[Chorus]\nWe're rising up tonight\nNothing can stop us now",
"duration_ms": 30000,
"positive_styles": ["bigger drums", "layered vocals", "anthemic"],
"negative_styles": ["sparse", "minimal"],
"context_adherence": "high"
}
]
}
audio = elevenlabs.music.compose(composition_plan=similar_plan, model_id="music_v2_5")

チャンクリファレンス

music_v2プランには最大30個のチャンクを含められます。各チャンクは、生成チャンクまたはオーディオ参照チャンクのいずれかです。

生成チャンク

フィールド型説明
textstring角括弧内のセクション名([Verse 1])、歌詞の行、中括弧内のインライン指示({scratching})。
duration_msnumber長さ(ミリ秒、3,000~120,000)。
positive_stylesarray含めるスタイルと指示(最大50個)。
negative_stylesarray避けるスタイルと指示(最大50個)。デフォルトは空です。
context_adherencestringlow、medium、またはhigh(デフォルト)。チャンクが前後のチャンクにどの程度忠実に従うかを指定します。
conditioning_refobject | null条件付けに使用する、保存済みオーディオの任意の{ song_id, range }スライス。デフォルトはnullです。
condition_strengthstring | nulllow、medium(デフォルト)、high、またはxhigh。チャンクが条件付け参照にどの程度強く従うかを指定します。

最初のチャンクのスタイルは、全体のトーンとジャンルを決めるため、最も重要です。方向性が定まるまで、 初期のチャンクには少なくとも6~7個のスタイルを設定してください。

オーディオ参照チャンク

フィールド型説明
song_idstringオーディオの取得元となる保存済み曲のID。
rangeobject変更せずに挿入する、保存済み曲の{ start_ms, end_ms }スライス。

制約

制約値
プランあたりの最大チャンク数30
最小チャンク長3秒(3,000ms)
最大チャンク長2分(120,000ms)
最大条件付け参照30秒(30,000ms)
最小時間範囲50ms

次のステップ