Raspberry PiでAgents Platformを使った音声アシスタントを構築する

Raspberry PiでAgents Platformを使った音声アシスタントを構築します。

チュートリアル · ElevenAgents クイックスタートを完了しており、PythonがインストールされたRaspberry Piを使用していることを前提としています。

はじめに

このチュートリアルでは、Raspberry Pi上で動作するAgents Platformを使って音声アシスタントを構築する方法を学びます。Amazon EchoのAlexa、Google Home、AppleデバイスのSiriといった一般的なホームアシスタントと同様に、Eleven Voiceアシスタントはホットワード(この例では「Hey Eleven」)を聞き取り、ElevenLabs Agentsセッションを開始してユーザーを支援します。

必要なもの

  • Raspberry Pi 5または同等のデバイス。
  • マイクとスピーカー。
  • Python 3.9以降がインストールされたマシン。
  • APIキーを持つElevenLabsアカウント。

セットアップ

依存関係をインストールする

Debianベースのシステムでは、次のコマンドで依存関係をインストールできます。

sudo apt-get update
sudo apt-get install libportaudio2 libportaudiocpp0 portaudio19-dev libasound-dev libsndfile1-dev -y

プロジェクトを作成する

Raspberry Piでターミナルを開き、プロジェクト用の新しいディレクトリを作成します。

mkdir eleven-voice-assistant
cd eleven-voice-assistant

新しい仮想環境を作成し、依存関係をインストールします。

python -m venv .venv # Only required the first time you set up the project
source .venv/bin/activate

依存関係をインストールします。

pip install tflite-runtime
pip install librosa
pip install EfficientWord-Net
pip install elevenlabs
pip install "elevenlabs[pyaudio]"

次に、hotword.pyという新しいPythonファイルを作成し、以下のコードを追加します。

hotword.py
import os
import signal
import time
from eff_word_net.streams import SimpleMicStream
from eff_word_net.engine import HotwordDetector
from eff_word_net.audio_processing import Resnet50_Arc_loss
# from eff_word_net import samples_loc
from elevenlabs.client import ElevenLabs
from elevenlabs.conversational_ai.conversation import Conversation, ConversationInitiationData
from elevenlabs.conversational_ai.default_audio_interface import DefaultAudioInterface
convai_active = False
elevenlabs = ElevenLabs()
agent_id = os.getenv("ELEVENLABS_AGENT_ID")
api_key = os.getenv("ELEVENLABS_API_KEY")
dynamic_vars = {
'user_name': 'Thor',
'greeting': 'Hey'
}
config = ConversationInitiationData(
dynamic_variables=dynamic_vars
)
base_model = Resnet50_Arc_loss()
eleven_hw = HotwordDetector(
hotword="hey_eleven",
model = base_model,
reference_file=os.path.join("hotword_refs", "hey_eleven_ref.json"),
threshold=0.7,
relaxation_time=2
)
def create_conversation():
"""Create a new conversation instance"""
return Conversation(
# API client and agent ID.
elevenlabs,
agent_id,
config=config,
# Assume auth is required when API_KEY is set.
requires_auth=bool(api_key),
# Use the default audio interface.
audio_interface=DefaultAudioInterface(),
# Simple callbacks that print the conversation to the console.
callback_agent_response=lambda response: print(f"Agent: {response}"),
callback_agent_response_correction=lambda original, corrected: print(f"Agent: {original} -> {corrected}"),
callback_user_transcript=lambda transcript: print(f"User: {transcript}"),
# Uncomment if you want to see latency measurements.
# callback_latency_measurement=lambda latency: print(f"Latency: {latency}ms"),
)
def start_mic_stream():
"""Start or restart the microphone stream"""
global mic_stream
try:
# Always create a new stream instance
mic_stream = SimpleMicStream(
window_length_secs=1.5,
sliding_window_secs=0.75,
)
mic_stream.start_stream()
print("Microphone stream started")
except Exception as e:
print(f"Error starting microphone stream: {e}")
mic_stream = None
time.sleep(1) # Wait a bit before retrying
def stop_mic_stream():
"""Stop the microphone stream safely"""
global mic_stream
try:
if mic_stream:
# SimpleMicStream doesn't have a stop_stream method
# We'll just set it to None and recreate it next time
mic_stream = None
print("Microphone stream stopped")
except Exception as e:
print(f"Error stopping microphone stream: {e}")
# Initialize microphone stream
mic_stream = None
start_mic_stream()
print("Say Hey Eleven ")
while True:
if not convai_active:
try:
# Make sure we have a valid mic stream
if mic_stream is None:
start_mic_stream()
continue
frame = mic_stream.getFrame()
result = eleven_hw.scoreFrame(frame)
if result == None:
#no voice activity
continue
if result["match"]:
print("Wakeword uttered", result["confidence"])
# Stop the microphone stream to avoid conflicts
stop_mic_stream()
# Start ConvAI Session
print("Start ConvAI Session")
convai_active = True
try:
# Create a new conversation instance
conversation = create_conversation()
# Start the session
conversation.start_session()
# Set up signal handler for graceful shutdown
def signal_handler(sig, frame):
print("Received interrupt signal, ending session...")
try:
conversation.end_session()
except Exception as e:
print(f"Error ending session: {e}")
signal.signal(signal.SIGINT, signal_handler)
# Wait for session to end
conversation_id = conversation.wait_for_session_end()
print(f"Conversation ID: {conversation_id}")
except Exception as e:
print(f"Error during conversation: {e}")
finally:
# Cleanup
convai_active = False
print("Conversation ended, cleaning up...")
# Give some time for cleanup
time.sleep(1)
# Restart microphone stream
start_mic_stream()
print("Ready for next wake word...")
except Exception as e:
print(f"Error in wake word detection: {e}")
# Try to restart microphone stream if there's an error
mic_stream = None
time.sleep(1)
start_mic_stream()

エージェントの設定

1

ElevenLabsにログイン

elevenlabs.ioにアクセスして、アカウントにログインします。

2

新しいエージェントを作成

Agents Platform > Agentsに移動し、 空のテンプレートから新しいエージェントを作成します。

3

最初のメッセージを設定

最初のメッセージを設定し、プラットフォーム用の動的変数を指定します。

{{greeting}} {{user_name}}, Eleven here, what's up?
4

システムプロンプトを設定

システムプロンプトを設定します。ベストプラクティスのドキュメントはこちらで確認できます。

You are a helpful Agents Platform assistant with access to a weather tool. When users ask about
weather conditions, use the get_weather tool to fetch accurate, real-time data. The tool requires
a latitude and longitude - use your geographic knowledge to convert location names to coordinates
accurately.
Never ask users for coordinates - you must determine these yourself. Always report weather
information conversationally, referring to locations by name only. For weather requests:
1. Extract the location from the user's message
2. Convert the location to coordinates and call get_weather
3. Present the information naturally and helpfully
For non-weather queries, provide friendly assistance within your knowledge boundaries. Always be
concise, accurate, and helpful.
5

Webhookツールを設定

天気データを取得するシンプルなWebhookツールを設定します。ツールを設定するには、こちらの手順に従ってください。

アプリを実行する

アプリを実行するには、まず必要な環境変数を設定します。

export ELEVENLABS_API_KEY=YOUR_API_KEY
export ELEVENLABS_AGENT_ID=YOUR_AGENT_ID

次に、以下のコマンドを実行します。

python hotword.py

「Hey Eleven」と話しかけると会話が始まります。会話を楽しみましょう!

[任意]カスタムホットワードをトレーニングする

トレーニング用オーディオを生成する

ホットワードの埋め込みを生成するには、ElevenLabsで4つのトレーニングサンプルを生成します。ElevenLabsアプリ内のテキスト読み上げに移動し、「Hey Eleven」などのホットワードを入力します。音声を選択して「Generate」ボタンをクリックします。

オーディオが生成されたら、オーディオファイルをダウンロードし、プロジェクトのルートにあるhotword_training_audioフォルダに保存します。別の音声を使ってこの手順をさらに3回繰り返します。

ホットワードをトレーニングする

仮想環境を有効にしたターミナルで、次のコマンドを実行してホットワードをトレーニングします。

python -m eff_word_net.generate_reference --input-dir hotword_training_audio --output-dir hotword_refs --wakeword hey_eleven --model-type resnet_50_arc

これにより、hotword_refsフォルダにhey_eleven_ref.jsonファイルが生成されます。あとはhotword.pyのHotwordDetectorクラスにあるreference_fileパラメータを新しい参照ファイルに更新すれば完了です!

次のステップ