在 Raspberry Pi 上使用 Agents Platform 构建语音助手

在 Raspberry Pi 上使用 Agents Platform 构建语音助手。

教程 · 假设你已完成 ElevenAgents 快速入门,并拥有安装了 Python 的 Raspberry Pi。

简介

本教程将介绍如何在 Raspberry Pi 上使用 Agents Platform 构建语音助手。与 Amazon Echo 上的 Alexa、Google Home 或 Apple 设备上的 Siri 等传统家庭助手一样,Eleven 语音助手会监听唤醒词(本例为“Hey Eleven”),然后启动 ElevenLabs Agents 会话来协助用户。

要求

  • Raspberry Pi 5 或类似设备。
  • 麦克风和扬声器。
  • 设备上已安装 Python 3.9 或更高版本。
  • 拥有 API 密钥的 ElevenLabs 账户。

设置

安装依赖项

在基于 Debian 的系统上,可使用以下命令安装依赖项:

sudo apt-get update
sudo apt-get install libportaudio2 libportaudiocpp0 portaudio19-dev libasound-dev libsndfile1-dev -y

创建项目

在 Raspberry Pi 上打开终端,为项目创建新目录。

mkdir eleven-voice-assistant
cd eleven-voice-assistant

创建新的虚拟环境并安装依赖项:

python -m venv .venv # Only required the first time you set up the project
source .venv/bin/activate

安装依赖项:

pip install tflite-runtime
pip install librosa
pip install EfficientWord-Net
pip install elevenlabs
pip install "elevenlabs[pyaudio]"

现在创建名为 hotword.py 的新 Python 文件,并添加以下代码:

hotword.py
import os
import signal
import time
from eff_word_net.streams import SimpleMicStream
from eff_word_net.engine import HotwordDetector
from eff_word_net.audio_processing import Resnet50_Arc_loss
# from eff_word_net import samples_loc
from elevenlabs.client import ElevenLabs
from elevenlabs.conversational_ai.conversation import Conversation, ConversationInitiationData
from elevenlabs.conversational_ai.default_audio_interface import DefaultAudioInterface
convai_active = False
elevenlabs = ElevenLabs()
agent_id = os.getenv("ELEVENLABS_AGENT_ID")
api_key = os.getenv("ELEVENLABS_API_KEY")
dynamic_vars = {
'user_name': 'Thor',
'greeting': 'Hey'
}
config = ConversationInitiationData(
dynamic_variables=dynamic_vars
)
base_model = Resnet50_Arc_loss()
eleven_hw = HotwordDetector(
hotword="hey_eleven",
model = base_model,
reference_file=os.path.join("hotword_refs", "hey_eleven_ref.json"),
threshold=0.7,
relaxation_time=2
)
def create_conversation():
"""Create a new conversation instance"""
return Conversation(
# API client and agent ID.
elevenlabs,
agent_id,
config=config,
# Assume auth is required when API_KEY is set.
requires_auth=bool(api_key),
# Use the default audio interface.
audio_interface=DefaultAudioInterface(),
# Simple callbacks that print the conversation to the console.
callback_agent_response=lambda response: print(f"Agent: {response}"),
callback_agent_response_correction=lambda original, corrected: print(f"Agent: {original} -> {corrected}"),
callback_user_transcript=lambda transcript: print(f"User: {transcript}"),
# Uncomment if you want to see latency measurements.
# callback_latency_measurement=lambda latency: print(f"Latency: {latency}ms"),
)
def start_mic_stream():
"""Start or restart the microphone stream"""
global mic_stream
try:
# Always create a new stream instance
mic_stream = SimpleMicStream(
window_length_secs=1.5,
sliding_window_secs=0.75,
)
mic_stream.start_stream()
print("Microphone stream started")
except Exception as e:
print(f"Error starting microphone stream: {e}")
mic_stream = None
time.sleep(1) # Wait a bit before retrying
def stop_mic_stream():
"""Stop the microphone stream safely"""
global mic_stream
try:
if mic_stream:
# SimpleMicStream doesn't have a stop_stream method
# We'll just set it to None and recreate it next time
mic_stream = None
print("Microphone stream stopped")
except Exception as e:
print(f"Error stopping microphone stream: {e}")
# Initialize microphone stream
mic_stream = None
start_mic_stream()
print("Say Hey Eleven ")
while True:
if not convai_active:
try:
# Make sure we have a valid mic stream
if mic_stream is None:
start_mic_stream()
continue
frame = mic_stream.getFrame()
result = eleven_hw.scoreFrame(frame)
if result == None:
#no voice activity
continue
if result["match"]:
print("Wakeword uttered", result["confidence"])
# Stop the microphone stream to avoid conflicts
stop_mic_stream()
# Start ConvAI Session
print("Start ConvAI Session")
convai_active = True
try:
# Create a new conversation instance
conversation = create_conversation()
# Start the session
conversation.start_session()
# Set up signal handler for graceful shutdown
def signal_handler(sig, frame):
print("Received interrupt signal, ending session...")
try:
conversation.end_session()
except Exception as e:
print(f"Error ending session: {e}")
signal.signal(signal.SIGINT, signal_handler)
# Wait for session to end
conversation_id = conversation.wait_for_session_end()
print(f"Conversation ID: {conversation_id}")
except Exception as e:
print(f"Error during conversation: {e}")
finally:
# Cleanup
convai_active = False
print("Conversation ended, cleaning up...")
# Give some time for cleanup
time.sleep(1)
# Restart microphone stream
start_mic_stream()
print("Ready for next wake word...")
except Exception as e:
print(f"Error in wake word detection: {e}")
# Try to restart microphone stream if there's an error
mic_stream = None
time.sleep(1)
start_mic_stream()

智能体配置

1

登录 ElevenLabs

前往 elevenlabs.io 并登录账户。

2

创建新智能体

前往 Agents Platform > Agents, 通过空白模板创建新智能体。

3

设置首条消息

设置首条消息,并指定平台的动态变量。

{{greeting}} {{user_name}}, Eleven here, what's up?
4

设置系统提示词

设置系统提示词。可在此处查看我们的最佳实践文档。

You are a helpful Agents Platform assistant with access to a weather tool. When users ask about
weather conditions, use the get_weather tool to fetch accurate, real-time data. The tool requires
a latitude and longitude - use your geographic knowledge to convert location names to coordinates
accurately.
Never ask users for coordinates - you must determine these yourself. Always report weather
information conversationally, referring to locations by name only. For weather requests:
1. Extract the location from the user's message
2. Convert the location to coordinates and call get_weather
3. Present the information naturally and helpfully
For non-weather queries, provide friendly assistance within your knowledge boundaries. Always be
concise, accurate, and helpful.
5

设置 webhook 工具

我们将设置一个简单的 webhook 工具来获取天气数据。请按照此处的设置步骤配置该工具。

运行应用

要运行应用,先设置所需的环境变量:

export ELEVENLABS_API_KEY=YOUR_API_KEY
export ELEVENLABS_AGENT_ID=YOUR_AGENT_ID

然后运行以下命令:

python hotword.py

现在说“Hey Eleven”即可开始对话。祝你聊天愉快!

【可选】训练自定义唤醒词

生成训练音频

要生成唤醒词嵌入向量,可使用 ElevenLabs 生成 4 个训练样本。只需在 ElevenLabs 应用中前往文本转语音,然后输入唤醒词,例如“Hey Eleven”。选择一个音色并点击“生成”按钮。

生成音频后,下载音频文件并保存到项目根目录下名为 hotword_training_audio 的文件夹。使用不同音色再重复此过程 3 次。

训练唤醒词

在已激活虚拟环境的终端中,运行以下命令训练唤醒词:

python -m eff_word_net.generate_reference --input-dir hotword_training_audio --output-dir hotword_refs --wakeword hey_eleven --model-type resnet_50_arc

这会在 hotword_refs 文件夹中生成 hey_eleven_ref.json 文件。现在只需更新 hotword.py 中 HotwordDetector 类的 reference_file 参数,使其指向新的参考文件,即可开始使用!

后续步骤