음악 인페인팅

기존 노래의 섹션 편집 및 결합

music_v2 또는 music_v2_5 모델을 사용한 음악 인페인팅을 통해 노래의 특정 부분을 수정하면서 나머지는 그대로 유지할 수 있습니다. 생성된 노래를 저장한 다음 작곡 계획에서 해당 부분을 참조하여 변경 없이 유지하거나, 다시 생성하거나, 원본을 조건으로 새 오디오를 생성할 수 있습니다.

작동 방식

music_v2 또는 music_v2_5 작곡 계획은 청크의 순서가 있는 목록입니다. 각 청크는 다음 두 유형 중 하나입니다.

  • 생성 청크 — text와 스타일을 기반으로 새 오디오를 생성합니다. 섹션을 다시 생성하거나 새 콘텐츠를 추가할 때 사용합니다.
  • 오디오 참조 청크 — 저장된 노래의 일부를 변경 없이 삽입합니다. 기존 노래의 섹션을 정확히 그대로 유지할 때 사용합니다.

빠른 시작

1

인페인팅용 노래 저장

인페인팅용 노래는 두 가지 방법으로 저장할 수 있습니다. store_for_inpainting으로 새 노래를 생성하거나 기존 오디오 파일을 업로드하세요.

import os
from dotenv import load_dotenv
from elevenlabs.client import ElevenLabs
load_dotenv()
elevenlabs = ElevenLabs(api_key=os.environ.get("ELEVENLABS_API_KEY"))
# Generate a song and store it for later inpainting
response = elevenlabs.music.compose_detailed(
prompt="An upbeat pop song with verse and chorus",
music_length_ms=60000,
model_id="music_v2_5",
store_for_inpainting=True
)
song_id = response.song_id
# Save the audio
with open("original.mp3", "wb") as f:
f.write(response.audio)
2

청크 유지 및 재생성

오디오 참조 청크(유지)와 생성 청크(재생성)를 혼합한 계획을 만든 다음, model_id="music_v2_5"와 함께 compose에 전달하세요.

# Keep the first 30 seconds, regenerate the rest with a new style
composition_plan = {
"chunks": [
# Keep the original first 30 seconds unchanged
{
"song_id": song_id,
"range": {"start_ms": 0, "end_ms": 30000}
},
# Regenerate the chorus with a new style
{
"text": "[Chorus]\nWe're rising up tonight\nNothing can stop us now",
"duration_ms": 30000,
"positive_styles": ["bigger drums", "layered vocals", "anthemic"],
"negative_styles": ["sparse", "minimal"],
"context_adherence": "high"
}
]
}
audio = elevenlabs.music.compose(
composition_plan=composition_plan,
model_id="music_v2_5",
)
with open("edited.mp3", "wb") as f:
for chunk in audio:
f.write(chunk)

조건 지정

생성 청크는 conditioning_ref를 사용해 저장된 오디오의 일부를 조건으로 지정할 수 있습니다. 모델은 참조의 음악적 특성을 가깝게 유지하면서 청크를 다시 생성합니다. condition_strength(low, medium, high 또는 xhigh)로 청크가 참조를 얼마나 밀접하게 따를지 제어하세요.

{
"text": "[Chorus]\nThis is my moment\nI won't let it go",
"duration_ms": 15000,
"positive_styles": ["powerful vocals", "full band", "anthemic"],
"negative_styles": [],
"context_adherence": "high",
"conditioning_ref": {
"song_id": "vVtPM1Sas70E2LIhQFch",
"range": { "start_ms": 30000, "end_ms": 45000 }
},
"condition_strength": "high"
}

첫 번째 청크는 이후 모든 청크의 생성에 영향을 줍니다. 참조를 기반으로 전체 노래에 조건을 지정하려면 첫 번째 청크부터 conditioning_ref를 적용하세요.

조건 참조는 최대 30초(30,000ms)까지 가능합니다.

컨텍스트 준수

각 생성 청크에는 인접한 청크를 얼마나 밀접하게 따를지 제어하는 context_adherence 수준이 있습니다.

  • high(기본값) — 주변 청크와 일관성을 유지합니다. 유지된 오디오와 재생성된 오디오 사이를 부드럽게 전환할 때 사용하세요.
  • medium — 일관성과 창의적 자유의 균형을 맞춥니다.
  • low — 청크가 컨텍스트에서 벗어나 더 창의적으로 표현할 수 있게 합니다.

예제

단일 섹션 편집

영화 예고편을 생성한 다음, 다른 가사로 아웃트로만 다시 생성합니다.

1

원본 생성

composition_plan = {
"chunks": [
{
"text": "[Intro]\nIn a world beyond code\nWhere sound becomes life",
"duration_ms": 15000,
"positive_styles": ["cinematic", "epic", "orchestral", "low strings", "suspenseful"],
"negative_styles": ["acoustic", "pop", "minimalistic"],
"context_adherence": "high"
},
{
"text": "[Build]\nTechnology awakens the future\nShaping every word into power",
"duration_ms": 20000,
"positive_styles": ["rising brass", "full orchestra", "epic"],
"negative_styles": ["acoustic", "pop"],
"context_adherence": "high"
},
{
"text": "[Bridge]\n(ah ah ah ah)",
"duration_ms": 15000,
"positive_styles": ["ethereal choir", "crescendo"],
"negative_styles": [],
"context_adherence": "high"
},
{
"text": "[Outro]\nThe voice of tomorrow, unleashed\nElevenLabs",
"duration_ms": 10000,
"positive_styles": ["deep narration", "epic finale"],
"negative_styles": [],
"context_adherence": "high"
}
]
}
response = elevenlabs.music.compose_detailed(
composition_plan=composition_plan,
model_id="music_v2_5",
store_for_inpainting=True
)
song_id = response.song_id
2

아웃트로만 편집

단일 오디오 참조 청크로 처음 세 섹션(0~50초)을 유지하고 아웃트로를 다시 생성합니다.

edited_plan = {
"chunks": [
# Keep the intro, build, and bridge unchanged
{
"song_id": song_id,
"range": {"start_ms": 0, "end_ms": 50000}
},
# Regenerate the outro with new lyrics
{
"text": "[Outro]\nThe future has arrived\nElevenLabs",
"duration_ms": 10000,
"positive_styles": ["deep narration", "epic finale"],
"negative_styles": [],
"context_adherence": "high"
}
]
}
audio = elevenlabs.music.compose(composition_plan=edited_plan, model_id="music_v2_5")

노래 확장

기존 노래에 새 인트로와 아웃트로를 추가합니다.

1

원본 생성

response = elevenlabs.music.compose_detailed(
prompt="Berlin night club techno",
music_length_ms=60000,
model_id="music_v2_5",
store_for_inpainting=True
)
song_id = response.song_id
2

새 인트로와 아웃트로로 확장

원본의 유지된 일부를 두 개의 새 생성 청크 사이에 배치합니다.

extend_plan = {
"chunks": [
# New intro
{
"text": "[Intro]",
"duration_ms": 30000,
"positive_styles": ["techno", "building tension", "filtered synths"],
"negative_styles": [],
"context_adherence": "high"
},
# Keep the core of the original (seconds 10-50)
{
"song_id": song_id,
"range": {"start_ms": 10000, "end_ms": 50000}
},
# New outro
{
"text": "[Outro]",
"duration_ms": 30000,
"positive_styles": ["techno", "fading out", "sparse"],
"negative_styles": [],
"context_adherence": "high"
}
]
}
audio = elevenlabs.music.compose(composition_plan=extend_plan, model_id="music_v2_5")

매끄러운 루프 만들기

음악 구절을 생성하고, 동일한 일부를 두 번 반복하는 부분을 연결하는 “글루” 청크를 사용해 루프를 만듭니다.

1

짧은 클립 생성

composition_plan = {
"chunks": [
{
"text": "[Solo Acoustic]",
"duration_ms": 10000,
"positive_styles": ["acoustic guitar", "fingerpicking", "warm tone", "soft dynamics"],
"negative_styles": ["electric", "drums", "electronic"],
"context_adherence": "high"
}
]
}
response = elevenlabs.music.compose_detailed(
composition_plan=composition_plan,
model_id="music_v2_5",
store_for_inpainting=True
)
song_id = response.song_id
2

글루 청크로 루프 만들기

loop_plan = {
"chunks": [
# Loop start - keep a slice of the original
{
"song_id": song_id,
"range": {"start_ms": 3000, "end_ms": 8000}
},
# Glue - generate a smooth transition between the two slices
{
"text": "[Glue]",
"duration_ms": 3000,
"positive_styles": ["acoustic guitar", "fingerpicking", "smooth transition"],
"negative_styles": [],
"context_adherence": "high"
},
# Loop end - keep the same slice again
{
"song_id": song_id,
"range": {"start_ms": 3000, "end_ms": 8000}
}
]
}
audio = elevenlabs.music.compose(composition_plan=loop_plan, model_id="music_v2_5")

유사한 노래 생성

기존 노래의 짧은 일부를 조건으로 완전히 새로운 노래를 생성하여 음악적 특성을 이어갑니다. 첫 번째 청크는 이후의 모든 청크에 영향을 주므로, 첫 번째 청크에 conditioning_ref를 적용하면 해당 청크만 저장된 오디오를 참조하더라도 전체 생성 결과가 형성됩니다.

1

원본 생성

response = elevenlabs.music.compose_detailed(
prompt="An upbeat pop song with bright synths and driving drums",
music_length_ms=60000,
model_id="music_v2_5",
store_for_inpainting=True
)
song_id = response.song_id
2

원본을 조건으로 새 노래 생성

similar_plan = {
"chunks": [
# The first chunk conditions the whole song on the reference
{
"text": "[Verse]\nSalt on my skin from a borrowed sea\nCounting the heartbeats it takes to break free",
"duration_ms": 30000,
"positive_styles": ["pop", "energetic", "bright synths", "driving drums"],
"negative_styles": ["sparse", "minimal"],
"context_adherence": "high",
"conditioning_ref": {
"song_id": song_id,
"range": {"start_ms": 0, "end_ms": 10000}
},
"condition_strength": "high"
},
{
"text": "[Chorus]\nWe're rising up tonight\nNothing can stop us now",
"duration_ms": 30000,
"positive_styles": ["bigger drums", "layered vocals", "anthemic"],
"negative_styles": ["sparse", "minimal"],
"context_adherence": "high"
}
]
}
audio = elevenlabs.music.compose(composition_plan=similar_plan, model_id="music_v2_5")

청크 레퍼런스

music_v2 플랜에는 최대 30개의 청크가 포함됩니다. 각 청크는 생성 청크 또는 오디오 레퍼런스 청크입니다.

생성 청크

필드유형설명
textstring대괄호 안의 섹션 이름([Verse 1]), 가사 줄, 중괄호 안의 인라인 지시문({scratching})입니다.
duration_msnumber길이(밀리초)입니다(3,000~120,000).
positive_stylesarray포함할 스타일 및 지시문입니다(최대 50개).
negative_stylesarray피할 스타일 및 지시문입니다(최대 50개). 기본값은 비어 있음입니다.
context_adherencestringlow, medium 또는 high(기본값)입니다. 청크가 주변 청크를 따르는 정도입니다.
conditioning_refobject | null조건으로 사용할 저장된 오디오의 선택적 { song_id, range } 슬라이스입니다. 기본값은 null입니다.
condition_strengthstring | nulllow, medium(기본값), high 또는 xhigh입니다. 청크가 조건 레퍼런스를 따르는 강도입니다.

첫 번째 청크의 스타일은 전체 분위기와 장르를 설정하므로 가장 중요합니다. 방향성이 정해질 때까지 초기 청크에는 최소 6~7개의 스타일을 사용하세요.

오디오 레퍼런스 청크

필드유형설명
song_idstring오디오를 가져올 저장된 곡의 ID입니다.
rangeobject변경 없이 삽입할 저장된 곡의 { start_ms, end_ms } 슬라이스입니다.

제약 조건

제약 조건값
플랜당 최대 청크 수30
최소 청크 길이3초(3,000ms)
최대 청크 길이2분(120,000ms)
최대 조건 레퍼런스 길이30초(30,000ms)
최소 시간 범위50ms

다음 단계