प्रोफेशनल वॉइस क्लोनिंग क्विकस्टार्ट

इस गाइड में बताया गया है कि PVC API का इस्तेमाल करके प्रोफेशनल वॉइस क्लोन कैसे बनाएं।

यह गाइड आपको PVC API का इस्तेमाल करके प्रोफेशनल वॉइस क्लोन (PVC) बनाने का तरीका दिखाएगी। डैशबोर्ड के ज़रिए PVC बनाने के लिए प्रोफेशनल वॉइस क्लोन प्रोडक्ट गाइड देखें।

PVC बनाने के लिए आपके पास Creator प्लान या उससे ऊपर होना चाहिए।

IVC और PVC अंदरूनी तौर पर कैसे काम करते हैं और किसे कब चुनना चाहिए, इसकी विस्तृत जानकारी के लिए वॉइस क्लोनिंग: यह कैसे काम करती है देखें।

अगर आपको कानूनी तौर पर अनुमत चीज़ों को लेकर संदेह है, तो ज़्यादा जानकारी के लिए कृपया सेवा की शर्तें और हमारी AI सुरक्षा जानकारी देखें।

API के ज़रिए PVC बनाने में इंस्टेंट वॉइस क्लोन बनाने की तुलना में काफ़ी ज़्यादा चरण होते हैं। ऐसा इसलिए है क्योंकि PVC ज़्यादा जटिल होते हैं और उच्च-गुणवत्ता वाला क्लोन बनाने के लिए ज़्यादा डेटा और फाइन-ट्यूनिंग की ज़रूरत होती है।

प्रोफेशनल वॉइस क्लोन API का इस्तेमाल

इस गाइड में माना गया है कि आपने अपनी API कुंजी और SDK सेट अप कर ली है। अगर नहीं की है, तो पहले क्विकस्टार्ट पूरा करें।

1

PVC वॉइस बनाएं

अपनी पसंद की भाषा के अनुसार example.py या example.mts नाम की नई फ़ाइल बनाएं और PVC वॉइस बनाने के लिए इसमें यह कोड जोड़ें:

# example.py
import os
import time
import base64
from contextlib import ExitStack
from io import BytesIO
from dotenv import load_dotenv
from elevenlabs.client import ElevenLabs
load_dotenv()
elevenlabs = ElevenLabs(
api_key=os.getenv("ELEVENLABS_API_KEY"),
)
voice = elevenlabs.voices.pvc.create(
name="My Professional Voice Clone",
language="en",
description="A professional voice clone of my voice"
)
print(voice)
2

ऑडियो फ़ाइलें अपलोड करें

अब हम वे ऑडियो सैंपल फ़ाइलें अपलोड करेंगे जिनका इस्तेमाल PVC को ट्रेन करने के लिए होगा। अपनी ऑडियो फ़ाइलों से सबसे अच्छे नतीजे पाने के बारे में ज़्यादा जानकारी के लिए PVC प्रोडक्ट गाइड का सुझाव और टिप्स सेक्शन देखें।

# Define the list of file paths explicitly
# Replace with the paths to your audio and/or video files.
# The more files you add, the better the clone will be.
sample_file_paths = [
"/path/to/your/first_sample.mp3",
"/path/to/your/second_sample.wav",
"relative/path/to/another_sample.mp4"
]
samples = None
files_to_upload = []
# Use ExitStack to manage multiple open files
with ExitStack() as stack:
for filepath in sample_file_paths:
# Open each file and add it to the stack
audio_file = stack.enter_context(open(filepath, "rb"))
filename = os.path.basename(filepath)
# Create a File object for the SDK
files_to_upload.append(
BytesIO(audio_file.read())
)
samples = elevenlabs.voices.pvc.samples.create(
voice_id=voice.voice_id,
files=files_to_upload # Pass the list of File objects
)
3

स्पीकर सेपरेशन शुरू करें

यह चरण ऑडियो फ़ाइलों को अलग-अलग स्पीकर में विभाजित करने की कोशिश करेगा। अगर आप कई स्पीकर वाली ऑडियो अपलोड कर रहे हैं, तो यह ज़रूरी है।

sample_ids_to_check = []
for sample in samples:
if sample.sample_id:
print(f"Starting separation for sample: {sample.sample_id}")
elevenlabs.voices.pvc.samples.speakers.separate(
voice_id=voice.voice_id,
sample_id=sample.sample_id
)
sample_ids_to_check.append(sample.sample_id)
while sample_ids_to_check:
# Create a copy of the list to iterate over, so we can remove items from the original
ids_in_batch = list(sample_ids_to_check)
for sample_id in ids_in_batch:
status_response = elevenlabs.voices.pvc.samples.speakers.get(
voice_id=voice.voice_id,
sample_id=sample_id
)
status = status_response.status
print(f"Sample {sample_id} status: {status}")
if status == "completed" or status == "failed":
sample_ids_to_check.remove(sample_id)
if sample_ids_to_check:
# Wait before the next poll cycle
time.sleep(5) # Wait for 5 seconds
print("All samples have been processed or removed from polling.")
4

स्पीकर ऑडियो प्राप्त करें

पिछले चरण को पूरा होने में कुछ समय लगेगा, इसलिए अगला चरण पिछले चरण के पूरा होने के बाद अलग प्रोसेस में चलाया जाना चाहिए।

स्पीकर सेपरेशन पूरा होने के बाद, आपको हर सैंपल के लिए स्पीकर की सूची मिलेगी। कई स्पीकर वाले सैंपल में, आपको PVC के लिए इस्तेमाल करने वाला स्पीकर चुनना होगा। स्पीकर की पहचान करने के लिए, आप हर स्पीकर का ऑडियो प्राप्त करके सुन सकते हैं।

# Get the list of samples from the voice created in Step 3
voice = elevenlabs.voices.get(voice_id=voice_id)
samples = voice.samples
# Loop over each sample and save the audio for each speaker to a file
speaker_audio_output_dir = "path/to/speakers/"
if not os.path.exists(speaker_audio_output_dir):
os.makedirs(speaker_audio_output_dir)
for sample in samples:
speaker_info = elevenlabs.voices.pvc.samples.speakers.get(
voice_id=voice.voice_id,
sample_id=sample.sample_id
)
# Proceed only if separation is actually complete
if getattr(speaker_info, 'status', 'unknown') != "completed":
continue
if hasattr(speaker_info, 'speakers') and speaker_info.speakers:
speaker_list = speaker_info.speakers
if isinstance(speaker_info.speakers, dict):
speaker_list = speaker_info.speakers.values()
for speaker in speaker_list:
audio_response = elevenlabs.voices.pvc.samples.speakers.audio.get(
voice_id=voice.voice_id,
sample_id=sample.sample_id,
speaker_id=speaker.speaker_id
)
audio_base64 = audio_response.audio_base_64
audio_data = base64.b64decode(audio_base64)
output_filename = os.path.join(speaker_audio_output_dir, f"sample_{sample.sample_id}_speaker_{speaker.speaker_id}.mp3")
with open(output_filename, "wb") as f:
f.write(audio_data)
5

स्पीकर ID के साथ सैंपल अपडेट करें

स्पीकर सेपरेशन पूरा होने के बाद, आप PVC के लिए इस्तेमाल करने वाला स्पीकर चुनने हेतु सैंपल अपडेट कर सकते हैं।

elevenlabs.voices.pvc.samples.update(
voice_id=voice.voice_id,
sample_id=sample.sample_id,
selected_speaker_ids=[speaker.speaker_id]
)
6

PVC सत्यापित करें

ट्रेनिंग शुरू करने से पहले, यह सुनिश्चित करने के लिए सत्यापन ज़रूरी है कि आपके पास वॉइस इस्तेमाल करने की अनुमति है। पहले सत्यापन CAPTCHA का अनुरोध करें।

captcha_response = elevenlabs.voices.pvc.verification.captcha.get(voice.voice_id)
# Save captcha image to file
captcha_buffer = base64.b64decode(captcha_response)
with open('captcha.png', 'wb') as f:
f.write(captcha_buffer)

इमेज में टेक्स्ट की कई पंक्तियां होती हैं जिन्हें वॉइस के मालिक को ज़ोर से पढ़कर रिकॉर्ड करना होगा। रिकॉर्डिंग पूरी होने पर, वॉइस के मालिक की पहचान सत्यापित करने के लिए उसे सबमिट करें।

elevenlabs.voices.pvc.verification.captcha.verify(
voice_id=voice.voice_id,
recording=open('path/to/recording.mp3', 'rb')
)
7

(वैकल्पिक) मैन्युअल सत्यापन का अनुरोध करें

अगर आप CAPTCHA सत्यापित नहीं कर पा रहे हैं, तो मैन्युअल सत्यापन का अनुरोध कर सकते हैं। ध्यान दें कि इसे प्रोसेस होने में ज़्यादा समय लगेगा।

इसका इस्तेमाल सिर्फ़ तब करें जब पिछले सत्यापन चरण असफल हो गए हों या संभव न हों, जैसे कि वॉइस के मालिक को देखने में परेशानी हो।

मैन्युअल सत्यापन के लिए आवश्यक फ़ाइलों की सूची के लिए सहायता टीम से संपर्क करें, क्योंकि हर मामला अलग हो सकता है।

elevenlabs.voices.pvc.verification.request(
voice_id=voice.voice_id,
files=[open('path/to/verification/files.txt', 'rb')],
)
8

PVC को ट्रेन करें

अब ट्रेनिंग प्रक्रिया शुरू करें। दिए गए सैंपल की लंबाई और संख्या के आधार पर इसे पूरा होने में कुछ समय लगेगा।

elevenlabs.voices.pvc.train(
voice_id=voice.voice_id,
# Specify the model the PVC should be trained on
model_id="eleven_multilingual_v2"
)
# Poll the fine tuning status until it is complete or fails
# This example specifically checks for the eleven_multilingual_v2 model
while True:
voice_details = elevenlabs.voices.get(voice_id=voice.voice_id)
fine_tuning_state = None
if voice_details.fine_tuning and voice_details.fine_tuning.state:
fine_tuning_state = voice_details.fine_tuning.state.get("eleven_multilingual_v2")
if fine_tuning_state:
progress = None
if voice_details.fine_tuning.progress and voice_details.fine_tuning.progress.get("eleven_multilingual_v2"):
progress = voice_details.fine_tuning.progress.get("eleven_multilingual_v2")
print(f"Fine tuning progress: {progress}")
if fine_tuning_state == "fine_tuned" or fine_tuning_state == "failed":
print("Fine tuning completed or failed")
break
# Wait for 5 seconds before polling again
time.sleep(5)
9

नई बनाई गई वॉइस इस्तेमाल करें

PVC सत्यापित हो जाने के बाद, आप इसे किसी भी अन्य वॉइस की तरह इस्तेमाल कर सकते हैं। वॉइस इस्तेमाल करने के बारे में ज़्यादा जानकारी के लिए टेक्स्ट टू स्पीच क्विकस्टार्ट देखें।

अगले चरण