프로페셔널 음성 복제 빠른 시작

이 가이드에서는 PVC API를 사용해 프로페셔널 음성 복제를 만드는 방법을 알아봅니다.

이 가이드에서는 PVC API를 사용해 프로페셔널 음성 복제(PVC)를 만드는 방법을 알아봅니다. 대시보드에서 PVC를 만들려면 프로페셔널 음성 복제 제품 가이드를 참조하세요.

PVC를 만들려면 크리에이터 플랜 이상을 사용해야 합니다.

IVC와 PVC의 내부 작동 방식 및 각각을 선택해야 하는 경우에 대한 자세한 설명은 음성 복제: 작동 방식을 참조하세요.

법적으로 허용되는 범위가 확실하지 않다면 자세한 내용은 서비스 약관 및 AI 안전 정보를 확인하세요.

API를 통해 PVC를 만드는 과정은 즉석 음성 복제를 만드는 것보다 훨씬 더 많은 단계로 이루어집니다. PVC는 더 복잡하며 고품질 복제를 만들기 위해 더 많은 데이터와 미세 조정이 필요하기 때문입니다.

프로페셔널 음성 복제 API 사용하기

이 가이드에서는 API 키 및 SDK를 설정했다고 가정합니다. 아직 완료하지 않았다면 먼저 빠른 시작을 완료하세요.

1

PVC 음성 만들기

사용하는 언어에 따라 example.py 또는 example.mts라는 새 파일을 만들고 PVC 음성을 생성하는 다음 코드를 추가하세요.

# example.py
import os
import time
import base64
from contextlib import ExitStack
from io import BytesIO
from dotenv import load_dotenv
from elevenlabs.client import ElevenLabs
load_dotenv()
elevenlabs = ElevenLabs(
api_key=os.getenv("ELEVENLABS_API_KEY"),
)
voice = elevenlabs.voices.pvc.create(
name="My Professional Voice Clone",
language="en",
description="A professional voice clone of my voice"
)
print(voice)
2

오디오 파일 업로드

다음으로 PVC 학습에 사용할 오디오 샘플 파일을 업로드합니다. 오디오 파일에서 최상의 결과를 얻는 방법에 관한 자세한 내용은 PVC 제품 가이드의 팁 및 제안 섹션을 확인하세요.

# Define the list of file paths explicitly
# Replace with the paths to your audio and/or video files.
# The more files you add, the better the clone will be.
sample_file_paths = [
"/path/to/your/first_sample.mp3",
"/path/to/your/second_sample.wav",
"relative/path/to/another_sample.mp4"
]
samples = None
files_to_upload = []
# Use ExitStack to manage multiple open files
with ExitStack() as stack:
for filepath in sample_file_paths:
# Open each file and add it to the stack
audio_file = stack.enter_context(open(filepath, "rb"))
filename = os.path.basename(filepath)
# Create a File object for the SDK
files_to_upload.append(
BytesIO(audio_file.read())
)
samples = elevenlabs.voices.pvc.samples.create(
voice_id=voice.voice_id,
files=files_to_upload # Pass the list of File objects
)
3

화자 분리 시작

이 단계에서는 오디오 파일을 개별 화자로 분리합니다. 여러 화자가 포함된 오디오를 업로드하는 경우 필요합니다.

sample_ids_to_check = []
for sample in samples:
if sample.sample_id:
print(f"Starting separation for sample: {sample.sample_id}")
elevenlabs.voices.pvc.samples.speakers.separate(
voice_id=voice.voice_id,
sample_id=sample.sample_id
)
sample_ids_to_check.append(sample.sample_id)
while sample_ids_to_check:
# Create a copy of the list to iterate over, so we can remove items from the original
ids_in_batch = list(sample_ids_to_check)
for sample_id in ids_in_batch:
status_response = elevenlabs.voices.pvc.samples.speakers.get(
voice_id=voice.voice_id,
sample_id=sample_id
)
status = status_response.status
print(f"Sample {sample_id} status: {status}")
if status == "completed" or status == "failed":
sample_ids_to_check.remove(sample_id)
if sample_ids_to_check:
# Wait before the next poll cycle
time.sleep(5) # Wait for 5 seconds
print("All samples have been processed or removed from polling.")
4

화자 오디오 가져오기

이전 단계를 완료하는 데 시간이 걸리므로, 다음 단계는 이전 단계가 완료된 후 별도의 프로세스에서 실행해야 합니다.

화자 분리가 완료되면 각 샘플의 화자 목록을 받게 됩니다. 여러 화자가 있는 샘플의 경우 PVC에 사용할 화자를 선택해야 합니다. 화자를 식별하려면 각 화자의 오디오를 가져와 들어 볼 수 있습니다.

# Get the list of samples from the voice created in Step 3
voice = elevenlabs.voices.get(voice_id=voice_id)
samples = voice.samples
# Loop over each sample and save the audio for each speaker to a file
speaker_audio_output_dir = "path/to/speakers/"
if not os.path.exists(speaker_audio_output_dir):
os.makedirs(speaker_audio_output_dir)
for sample in samples:
speaker_info = elevenlabs.voices.pvc.samples.speakers.get(
voice_id=voice.voice_id,
sample_id=sample.sample_id
)
# Proceed only if separation is actually complete
if getattr(speaker_info, 'status', 'unknown') != "completed":
continue
if hasattr(speaker_info, 'speakers') and speaker_info.speakers:
speaker_list = speaker_info.speakers
if isinstance(speaker_info.speakers, dict):
speaker_list = speaker_info.speakers.values()
for speaker in speaker_list:
audio_response = elevenlabs.voices.pvc.samples.speakers.audio.get(
voice_id=voice.voice_id,
sample_id=sample.sample_id,
speaker_id=speaker.speaker_id
)
audio_base64 = audio_response.audio_base_64
audio_data = base64.b64decode(audio_base64)
output_filename = os.path.join(speaker_audio_output_dir, f"sample_{sample.sample_id}_speaker_{speaker.speaker_id}.mp3")
with open(output_filename, "wb") as f:
f.write(audio_data)
5

화자 ID로 샘플 업데이트

화자 분리가 완료되면 샘플을 업데이트하여 PVC에 사용할 화자를 선택할 수 있습니다.

elevenlabs.voices.pvc.samples.update(
voice_id=voice.voice_id,
sample_id=sample.sample_id,
selected_speaker_ids=[speaker.speaker_id]
)
6

PVC 검증

학습을 시작하기 전에 음성을 사용할 권한이 있는지 확인하는 검증 단계가 필요합니다. 먼저 검증 CAPTCHA를 요청하세요.

captcha_response = elevenlabs.voices.pvc.verification.captcha.get(voice.voice_id)
# Save captcha image to file
captcha_buffer = base64.b64decode(captcha_response)
with open('captcha.png', 'wb') as f:
f.write(captcha_buffer)

이미지에는 음성 소유자가 소리 내어 읽고 녹음해야 하는 여러 줄의 텍스트가 있습니다. 완료되면 녹음을 제출하여 음성 소유자의 신원을 검증하세요.

elevenlabs.voices.pvc.verification.captcha.verify(
voice_id=voice.voice_id,
recording=open('path/to/recording.mp3', 'rb')
)
7

(선택 사항) 수동 검증 요청

CAPTCHA를 검증할 수 없는 경우 수동 검증을 요청할 수 있습니다. 처리 시간이 더 오래 걸립니다.

이 방법은 이전 검증 단계가 실패했거나 불가능한 경우에만 사용해야 합니다. 예를 들어 음성 소유자가 시각 장애가 있는 경우가 이에 해당합니다.

수동 검증에 필요한 파일 목록은 사례마다 다를 수 있으므로 고객 지원팀에 문의하세요.

elevenlabs.voices.pvc.verification.request(
voice_id=voice.voice_id,
files=[open('path/to/verification/files.txt', 'rb')],
)
8

PVC 학습

다음으로 학습 프로세스를 시작합니다. 제공된 샘플의 길이와 수에 따라 완료까지 시간이 걸립니다.

elevenlabs.voices.pvc.train(
voice_id=voice.voice_id,
# Specify the model the PVC should be trained on
model_id="eleven_multilingual_v2"
)
# Poll the fine tuning status until it is complete or fails
# This example specifically checks for the eleven_multilingual_v2 model
while True:
voice_details = elevenlabs.voices.get(voice_id=voice.voice_id)
fine_tuning_state = None
if voice_details.fine_tuning and voice_details.fine_tuning.state:
fine_tuning_state = voice_details.fine_tuning.state.get("eleven_multilingual_v2")
if fine_tuning_state:
progress = None
if voice_details.fine_tuning.progress and voice_details.fine_tuning.progress.get("eleven_multilingual_v2"):
progress = voice_details.fine_tuning.progress.get("eleven_multilingual_v2")
print(f"Fine tuning progress: {progress}")
if fine_tuning_state == "fine_tuned" or fine_tuning_state == "failed":
print("Fine tuning completed or failed")
break
# Wait for 5 seconds before polling again
time.sleep(5)
9

새로 만든 음성 사용

PVC가 검증되면 다른 음성과 동일한 방식으로 사용할 수 있습니다. 음성 사용 방법에 관한 자세한 내용은 텍스트 음성 변환 빠른 시작을 참조하세요.

다음 단계