プロフェッショナルボイスクローン クイックスタート

このガイドでは、PVC APIを使用してプロフェッショナルボイスクローンを作成する方法を説明します。

このガイドでは、PVC APIを使用してプロフェッショナルボイスクローン(PVC)を作成する方法を説明します。ダッシュボードからPVCを作成する場合は、プロフェッショナルボイスクローンのプロダクトガイドを参照してください。

PVCを作成するには、クリエイタープラン以上に加入している必要があります。

IVCとPVCの仕組み、それぞれを選ぶタイミングについて詳しくは、ボイスクローン:仕組みを参照してください。

法的に許可される範囲が不明な場合は、詳細について利用規約とAIセーフティに関する情報を確認してください。

API経由でPVCを作成する場合、インスタントボイスクローンの作成よりもかなり多くの手順が必要です。PVCはより複雑で、高品質なクローンを作成するにはより多くのデータと微調整が必要なためです。

プロフェッショナルボイスクローンAPIを使用する

このガイドでは、APIキーとSDKを設定済みであることを前提としています。まだの場合は、先にクイックスタートを完了してください。

1

PVC音声を作成する

使用する言語に応じてexample.pyまたはexample.mtsという新しいファイルを作成し、以下のコードを追加してPVC音声を作成します。

# example.py
import os
import time
import base64
from contextlib import ExitStack
from io import BytesIO
from dotenv import load_dotenv
from elevenlabs.client import ElevenLabs
load_dotenv()
elevenlabs = ElevenLabs(
api_key=os.getenv("ELEVENLABS_API_KEY"),
)
voice = elevenlabs.voices.pvc.create(
name="My Professional Voice Clone",
language="en",
description="A professional voice clone of my voice"
)
print(voice)
2

オーディオファイルをアップロードする

次に、PVCのトレーニングに使用するオーディオサンプルファイルをアップロードします。オーディオファイルから最良の結果を得る方法については、PVCプロダクトガイドのヒントと推奨事項セクションを確認してください。

# Define the list of file paths explicitly
# Replace with the paths to your audio and/or video files.
# The more files you add, the better the clone will be.
sample_file_paths = [
"/path/to/your/first_sample.mp3",
"/path/to/your/second_sample.wav",
"relative/path/to/another_sample.mp4"
]
samples = None
files_to_upload = []
# Use ExitStack to manage multiple open files
with ExitStack() as stack:
for filepath in sample_file_paths:
# Open each file and add it to the stack
audio_file = stack.enter_context(open(filepath, "rb"))
filename = os.path.basename(filepath)
# Create a File object for the SDK
files_to_upload.append(
BytesIO(audio_file.read())
)
samples = elevenlabs.voices.pvc.samples.create(
voice_id=voice.voice_id,
files=files_to_upload # Pass the list of File objects
)
3

話者分離を開始する

この手順では、オーディオファイルを個々の話者に分離します。複数の話者を含むオーディオをアップロードする場合に必要です。

sample_ids_to_check = []
for sample in samples:
if sample.sample_id:
print(f"Starting separation for sample: {sample.sample_id}")
elevenlabs.voices.pvc.samples.speakers.separate(
voice_id=voice.voice_id,
sample_id=sample.sample_id
)
sample_ids_to_check.append(sample.sample_id)
while sample_ids_to_check:
# Create a copy of the list to iterate over, so we can remove items from the original
ids_in_batch = list(sample_ids_to_check)
for sample_id in ids_in_batch:
status_response = elevenlabs.voices.pvc.samples.speakers.get(
voice_id=voice.voice_id,
sample_id=sample_id
)
status = status_response.status
print(f"Sample {sample_id} status: {status}")
if status == "completed" or status == "failed":
sample_ids_to_check.remove(sample_id)
if sample_ids_to_check:
# Wait before the next poll cycle
time.sleep(5) # Wait for 5 seconds
print("All samples have been processed or removed from polling.")
4

話者のオーディオを取得する

前の手順の完了には時間がかかるため、次の手順は前の手順が完了した後に別のプロセスで実行してください。

話者分離が完了すると、サンプルごとに話者のリストが取得できます。複数の話者を含むサンプルの場合、PVCに使用する話者を選択する必要があります。話者を識別するには、各話者のオーディオを取得して聞いてください。

# Get the list of samples from the voice created in Step 3
voice = elevenlabs.voices.get(voice_id=voice_id)
samples = voice.samples
# Loop over each sample and save the audio for each speaker to a file
speaker_audio_output_dir = "path/to/speakers/"
if not os.path.exists(speaker_audio_output_dir):
os.makedirs(speaker_audio_output_dir)
for sample in samples:
speaker_info = elevenlabs.voices.pvc.samples.speakers.get(
voice_id=voice.voice_id,
sample_id=sample.sample_id
)
# Proceed only if separation is actually complete
if getattr(speaker_info, 'status', 'unknown') != "completed":
continue
if hasattr(speaker_info, 'speakers') and speaker_info.speakers:
speaker_list = speaker_info.speakers
if isinstance(speaker_info.speakers, dict):
speaker_list = speaker_info.speakers.values()
for speaker in speaker_list:
audio_response = elevenlabs.voices.pvc.samples.speakers.audio.get(
voice_id=voice.voice_id,
sample_id=sample.sample_id,
speaker_id=speaker.speaker_id
)
audio_base64 = audio_response.audio_base_64
audio_data = base64.b64decode(audio_base64)
output_filename = os.path.join(speaker_audio_output_dir, f"sample_{sample.sample_id}_speaker_{speaker.speaker_id}.mp3")
with open(output_filename, "wb") as f:
f.write(audio_data)
5

話者IDを使用してサンプルを更新する

話者分離が完了したら、サンプルを更新してPVCに使用する話者を選択できます。

elevenlabs.voices.pvc.samples.update(
voice_id=voice.voice_id,
sample_id=sample.sample_id,
selected_speaker_ids=[speaker.speaker_id]
)
6

PVCを検証する

トレーニングを開始する前に、音声を使用する権限があることを確認するための検証手順が必要です。まず、検証用CAPTCHAをリクエストします。

captcha_response = elevenlabs.voices.pvc.verification.captcha.get(voice.voice_id)
# Save captcha image to file
captcha_buffer = base64.b64decode(captcha_response)
with open('captcha.png', 'wb') as f:
f.write(captcha_buffer)

画像には、音声のオーナーが声に出して読み上げて録音する必要がある複数行のテキストが含まれています。完了したら、録音を送信して音声オーナーの本人確認を行います。

elevenlabs.voices.pvc.verification.captcha.verify(
voice_id=voice.voice_id,
recording=open('path/to/recording.mp3', 'rb')
)
7

(任意)手動検証をリクエストする

CAPTCHAを検証できない場合は、手動検証をリクエストできます。処理にはより時間がかかることに注意してください。

これは、前の検証手順が失敗した場合、または音声のオーナーに視覚障害がある場合など、検証が不可能な場合にのみ使用してください。

必要なファイルはケースごとに異なる場合があるため、手動検証に必要なファイルの一覧についてはサポートまでお問い合わせください。

elevenlabs.voices.pvc.verification.request(
voice_id=voice.voice_id,
files=[open('path/to/verification/files.txt', 'rb')],
)
8

PVCをトレーニングする

次に、トレーニングプロセスを開始します。完了までの時間は、提供するサンプルの長さと数によって異なります。

elevenlabs.voices.pvc.train(
voice_id=voice.voice_id,
# Specify the model the PVC should be trained on
model_id="eleven_multilingual_v2"
)
# Poll the fine tuning status until it is complete or fails
# This example specifically checks for the eleven_multilingual_v2 model
while True:
voice_details = elevenlabs.voices.get(voice_id=voice.voice_id)
fine_tuning_state = None
if voice_details.fine_tuning and voice_details.fine_tuning.state:
fine_tuning_state = voice_details.fine_tuning.state.get("eleven_multilingual_v2")
if fine_tuning_state:
progress = None
if voice_details.fine_tuning.progress and voice_details.fine_tuning.progress.get("eleven_multilingual_v2"):
progress = voice_details.fine_tuning.progress.get("eleven_multilingual_v2")
print(f"Fine tuning progress: {progress}")
if fine_tuning_state == "fine_tuned" or fine_tuning_state == "failed":
print("Fine tuning completed or failed")
break
# Wait for 5 seconds before polling again
time.sleep(5)
9

新しく作成した音声を使用する

PVCが検証されると、他の音声と同じように使用できます。音声の使用方法については、テキスト読み上げクイックスタートを参照してください。

次のステップ