म्यूज़िक इनपेंटिंग

मौजूदा गानों के सेक्शन एडिट और मिलाएं

music_v2 या music_v2_5 मॉडल के साथ म्यूज़िक इनपेंटिंग आपको गाने के खास हिस्सों में बदलाव करने देती है, जबकि बाकी गाना वैसा ही रहता है। जनरेट किया गया गाना स्टोर करें, फिर उसके हिस्सों को कंपोज़िशन प्लान में रेफरेंस करें ताकि वे बिना बदलाव के रहें, दोबारा जनरेट हों या मूल ऑडियो के आधार पर नया ऑडियो कंडीशन किया जा सके।

यह कैसे काम करता है

music_v2 या music_v2_5 कंपोज़िशन प्लान चंक्स की एक क्रमबद्ध सूची है। हर चंक दो प्रकारों में से एक होता है:

  • जनरेशन चंक — text और स्टाइल से नया ऑडियो जनरेट करता है। किसी सेक्शन को फिर से जनरेट करने या नई सामग्री जोड़ने के लिए इसका इस्तेमाल करें।
  • ऑडियो रेफरेंस चंक — स्टोर किए गए गाने का एक हिस्सा बिना बदलाव के जोड़ता है। मौजूदा गाने के किसी सेक्शन को बिल्कुल वैसा ही रखने के लिए इसका इस्तेमाल करें।

क्विकस्टार्ट

1

इनपेंटिंग के लिए गाना स्टोर करें

इनपेंटिंग के लिए गाना दो तरीकों से स्टोर कर सकते हैं: store_for_inpainting के साथ नया गाना जनरेट करें या मौजूदा ऑडियो फ़ाइल अपलोड करें।

import os
from dotenv import load_dotenv
from elevenlabs.client import ElevenLabs
load_dotenv()
elevenlabs = ElevenLabs(api_key=os.environ.get("ELEVENLABS_API_KEY"))
# Generate a song and store it for later inpainting
response = elevenlabs.music.compose_detailed(
prompt="An upbeat pop song with verse and chorus",
music_length_ms=60000,
model_id="music_v2_5",
store_for_inpainting=True
)
song_id = response.song_id
# Save the audio
with open("original.mp3", "wb") as f:
f.write(response.audio)
2

चंक्स रखें और फिर से जनरेट करें

ऐसा प्लान बनाएं जिसमें ऑडियो रेफरेंस चंक्स (रखे गए) और जनरेशन चंक्स (दोबारा जनरेट किए गए) दोनों शामिल हों, फिर इसे model_id="music_v2_5" के साथ compose में पास करें:

# Keep the first 30 seconds, regenerate the rest with a new style
composition_plan = {
"chunks": [
# Keep the original first 30 seconds unchanged
{
"song_id": song_id,
"range": {"start_ms": 0, "end_ms": 30000}
},
# Regenerate the chorus with a new style
{
"text": "[Chorus]\nWe're rising up tonight\nNothing can stop us now",
"duration_ms": 30000,
"positive_styles": ["bigger drums", "layered vocals", "anthemic"],
"negative_styles": ["sparse", "minimal"],
"context_adherence": "high"
}
]
}
audio = elevenlabs.music.compose(
composition_plan=composition_plan,
model_id="music_v2_5",
)
with open("edited.mp3", "wb") as f:
for chunk in audio:
f.write(chunk)

कंडीशनिंग

किसी जनरेशन चंक को conditioning_ref के साथ स्टोर किए गए ऑडियो के एक हिस्से पर कंडीशन किया जा सकता है। मॉडल रेफरेंस की संगीत विशेषताओं के करीब रहते हुए चंक को फिर से जनरेट करता है। condition_strength (low, medium, high, या xhigh) से नियंत्रित करें कि चंक रेफरेंस का कितनी बारीकी से पालन करे।

{
"text": "[Chorus]\nThis is my moment\nI won't let it go",
"duration_ms": 15000,
"positive_styles": ["powerful vocals", "full band", "anthemic"],
"negative_styles": [],
"context_adherence": "high",
"conditioning_ref": {
"song_id": "vVtPM1Sas70E2LIhQFch",
"range": { "start_ms": 30000, "end_ms": 45000 }
},
"condition_strength": "high"
}

पहला चंक, उसके बाद के सभी चंक्स के जनरेशन को प्रभावित करता है। पूरे गाने को किसी रेफरेंस पर कंडीशन करने के लिए, पहले चंक से conditioning_ref लागू करें।

कंडीशनिंग रेफरेंस की अधिकतम लंबाई 30 सेकंड (30,000ms) हो सकती है।

कॉन्टेक्स्ट अड्हीरेंस

हर जनरेशन चंक में context_adherence स्तर होता है, जो नियंत्रित करता है कि वह अपने आस-पास के चंक्स का कितनी बारीकी से पालन करे:

  • high (डिफ़ॉल्ट) — आसपास के चंक्स के साथ एकरूप रहता है। रखे गए और फिर से जनरेट किए गए ऑडियो के बीच सहज ट्रांज़िशन के लिए इसका इस्तेमाल करें।
  • medium — एकरूपता और रचनात्मक स्वतंत्रता के बीच संतुलन रखता है।
  • low — चंक को अपने कॉन्टेक्स्ट से अलग होने और ज़्यादा रचनात्मक होने देता है।

उदाहरण

एक सेक्शन एडिट करें

मूवी ट्रेलर जनरेट करें, फिर अलग बोलों के साथ सिर्फ़ आउट्रो को दोबारा जनरेट करें।

1

ओरिजिनल जनरेट करें

composition_plan = {
"chunks": [
{
"text": "[Intro]\nIn a world beyond code\nWhere sound becomes life",
"duration_ms": 15000,
"positive_styles": ["cinematic", "epic", "orchestral", "low strings", "suspenseful"],
"negative_styles": ["acoustic", "pop", "minimalistic"],
"context_adherence": "high"
},
{
"text": "[Build]\nTechnology awakens the future\nShaping every word into power",
"duration_ms": 20000,
"positive_styles": ["rising brass", "full orchestra", "epic"],
"negative_styles": ["acoustic", "pop"],
"context_adherence": "high"
},
{
"text": "[Bridge]\n(ah ah ah ah)",
"duration_ms": 15000,
"positive_styles": ["ethereal choir", "crescendo"],
"negative_styles": [],
"context_adherence": "high"
},
{
"text": "[Outro]\nThe voice of tomorrow, unleashed\nElevenLabs",
"duration_ms": 10000,
"positive_styles": ["deep narration", "epic finale"],
"negative_styles": [],
"context_adherence": "high"
}
]
}
response = elevenlabs.music.compose_detailed(
composition_plan=composition_plan,
model_id="music_v2_5",
store_for_inpainting=True
)
song_id = response.song_id
2

सिर्फ़ आउट्रो एडिट करें

एक ऑडियो रेफ़रेंस चंक के साथ पहले तीन सेक्शन (0–50 सेकंड) रखें और आउट्रो को दोबारा जनरेट करें:

edited_plan = {
"chunks": [
# Keep the intro, build, and bridge unchanged
{
"song_id": song_id,
"range": {"start_ms": 0, "end_ms": 50000}
},
# Regenerate the outro with new lyrics
{
"text": "[Outro]\nThe future has arrived\nElevenLabs",
"duration_ms": 10000,
"positive_styles": ["deep narration", "epic finale"],
"negative_styles": [],
"context_adherence": "high"
}
]
}
audio = elevenlabs.music.compose(composition_plan=edited_plan, model_id="music_v2_5")

गाना बढ़ाएं

मौजूदा गाने में नया इंट्रो और आउट्रो जोड़ें।

1

ओरिजिनल जनरेट करें

response = elevenlabs.music.compose_detailed(
prompt="Berlin night club techno",
music_length_ms=60000,
model_id="music_v2_5",
store_for_inpainting=True
)
song_id = response.song_id
2

नए इंट्रो और आउट्रो के साथ बढ़ाएं

ओरिजिनल के रखे गए हिस्से को दो नए जनरेशन चंक्स के बीच रखें:

extend_plan = {
"chunks": [
# New intro
{
"text": "[Intro]",
"duration_ms": 30000,
"positive_styles": ["techno", "building tension", "filtered synths"],
"negative_styles": [],
"context_adherence": "high"
},
# Keep the core of the original (seconds 10-50)
{
"song_id": song_id,
"range": {"start_ms": 10000, "end_ms": 50000}
},
# New outro
{
"text": "[Outro]",
"duration_ms": 30000,
"positive_styles": ["techno", "fading out", "sparse"],
"negative_styles": [],
"context_adherence": "high"
}
]
}
audio = elevenlabs.music.compose(composition_plan=extend_plan, model_id="music_v2_5")

एक सहज लूप बनाएं

संगीत का एक अंश जनरेट करें और उसी स्लाइस को दो बार दोहराने वाले “ग्लू” चंक का इस्तेमाल करके लूप बनाएं।

1

छोटा क्लिप जनरेट करें

composition_plan = {
"chunks": [
{
"text": "[Solo Acoustic]",
"duration_ms": 10000,
"positive_styles": ["acoustic guitar", "fingerpicking", "warm tone", "soft dynamics"],
"negative_styles": ["electric", "drums", "electronic"],
"context_adherence": "high"
}
]
}
response = elevenlabs.music.compose_detailed(
composition_plan=composition_plan,
model_id="music_v2_5",
store_for_inpainting=True
)
song_id = response.song_id
2

ग्लू चंक के साथ लूप बनाएं

loop_plan = {
"chunks": [
# Loop start - keep a slice of the original
{
"song_id": song_id,
"range": {"start_ms": 3000, "end_ms": 8000}
},
# Glue - generate a smooth transition between the two slices
{
"text": "[Glue]",
"duration_ms": 3000,
"positive_styles": ["acoustic guitar", "fingerpicking", "smooth transition"],
"negative_styles": [],
"context_adherence": "high"
},
# Loop end - keep the same slice again
{
"song_id": song_id,
"range": {"start_ms": 3000, "end_ms": 8000}
}
]
}
audio = elevenlabs.music.compose(composition_plan=loop_plan, model_id="music_v2_5")

मिलता-जुलता गाना जनरेट करें

किसी मौजूदा गाने के छोटे हिस्से के आधार पर बिल्कुल नया गाना बनाएं, ताकि उसकी संगीत संबंधी विशेषताएं बनी रहें। पहला चंक उसके बाद के हर चंक को प्रभावित करता है, इसलिए पहले चंक पर conditioning_ref लागू करने से पूरी जनरेशन आकार लेती है, भले ही सिर्फ़ वही चंक स्टोर किए गए ऑडियो को रेफ़रेंस करता हो।

1

ओरिजिनल जनरेट करें

response = elevenlabs.music.compose_detailed(
prompt="An upbeat pop song with bright synths and driving drums",
music_length_ms=60000,
model_id="music_v2_5",
store_for_inpainting=True
)
song_id = response.song_id
2

ओरिजिनल के आधार पर नया गाना जनरेट करें

similar_plan = {
"chunks": [
# The first chunk conditions the whole song on the reference
{
"text": "[Verse]\nSalt on my skin from a borrowed sea\nCounting the heartbeats it takes to break free",
"duration_ms": 30000,
"positive_styles": ["pop", "energetic", "bright synths", "driving drums"],
"negative_styles": ["sparse", "minimal"],
"context_adherence": "high",
"conditioning_ref": {
"song_id": song_id,
"range": {"start_ms": 0, "end_ms": 10000}
},
"condition_strength": "high"
},
{
"text": "[Chorus]\nWe're rising up tonight\nNothing can stop us now",
"duration_ms": 30000,
"positive_styles": ["bigger drums", "layered vocals", "anthemic"],
"negative_styles": ["sparse", "minimal"],
"context_adherence": "high"
}
]
}
audio = elevenlabs.music.compose(composition_plan=similar_plan, model_id="music_v2_5")

चंक रेफ़रेंस

music_v2 प्लान में अधिकतम 30 चंक्स होते हैं। हर चंक या तो जनरेशन चंक होता है या ऑडियो रेफ़रेंस चंक।

जनरेशन चंक

फ़ील्डटाइपविवरण
textstringवर्ग कोष्ठक में सेक्शन का नाम ([Verse 1]), बोल की पंक्तियां और कर्ली ब्रेसेज़ में इनलाइन निर्देश ({scratching})।
duration_msnumberमिलीसेकंड में लंबाई (3,000 - 120,000)।
positive_stylesarrayशामिल करने के लिए स्टाइल और निर्देश (अधिकतम 50)।
negative_stylesarrayजिन स्टाइल और निर्देशों से बचना है (अधिकतम 50)। डिफ़ॉल्ट रूप से खाली।
context_adherencestringlow, medium या high (डिफ़ॉल्ट)। चंक अपने आसपास के चंक्स का कितनी बारीकी से पालन करता है।
conditioning_refobject | nullकंडीशनिंग के लिए स्टोर किए गए ऑडियो का वैकल्पिक { song_id, range } स्लाइस। डिफ़ॉल्ट null है।
condition_strengthstring | nulllow, medium (डिफ़ॉल्ट), high या xhigh। चंक कंडीशनिंग रेफ़रेंस का कितनी मज़बूती से पालन करता है।

पहले चंक के स्टाइल सबसे अहम होते हैं, क्योंकि वे ओवरऑल टोन और जॉनर तय करते हैं। दिशा तय होने तक शुरुआती चंक्स में कम से कम 6-7 स्टाइल रखने की कोशिश करें।

ऑडियो रेफ़रेंस चंक

फ़ील्डटाइपविवरण
song_idstringऑडियो स्रोत के लिए स्टोर किए गए गाने की ID।
rangeobjectबिना बदलाव डालने के लिए स्टोर किए गए गाने का { start_ms, end_ms } स्लाइस।

सीमाएं

सीमामान
प्रति प्लान अधिकतम चंक्स30
चंक की न्यूनतम अवधि3 सेकंड (3,000ms)
चंक की अधिकतम अवधि2 मिनट (120,000ms)
अधिकतम कंडीशनिंग रेफ़रेंस30 सेकंड (30,000ms)
न्यूनतम समय सीमा50ms

अगले चरण