Voice Assistant mit Agents Platform auf einem Raspberry Pi erstellen

Erstellen Sie einen Voice Assistant mit Agents Platform auf einem Raspberry Pi.

Tutorial · Setzt voraus, dass Sie den ElevenAgents- Schnellstart abgeschlossen haben und einen Raspberry Pi mit installiertem Python besitzen.

Einführung

In diesem Tutorial erfahren Sie, wie Sie einen Voice Assistant mit Agents Platform auf einem Raspberry Pi erstellen. Wie herkömmliche Home-Assistenten wie Alexa auf Amazon Echo, Google Home oder Siri auf Apple-Geräten hört Ihr Eleven Voice Assistant auf ein Aktivierungswort – in unserem Fall „Hey Eleven“ – und startet dann eine ElevenLabs-Agents-Sitzung, um den Nutzer zu unterstützen.

Voraussetzungen

  • Ein Raspberry Pi 5 oder ein ähnliches Gerät.
  • Ein Mikrofon und Lautsprecher.
  • Python 3.9 oder höher muss auf Ihrem Gerät installiert sein.
  • Ein ElevenLabs-Konto mit einem API-Schlüssel.

Einrichtung

Abhängigkeiten installieren

Auf Debian-basierten Systemen können Sie die Abhängigkeiten wie folgt installieren:

sudo apt-get update
sudo apt-get install libportaudio2 libportaudiocpp0 portaudio19-dev libasound-dev libsndfile1-dev -y

Projekt erstellen

Öffnen Sie auf Ihrem Raspberry Pi das Terminal und erstellen Sie ein neues Verzeichnis für Ihr Projekt.

mkdir eleven-voice-assistant
cd eleven-voice-assistant

Erstellen Sie eine neue virtuelle Umgebung und installieren Sie die Abhängigkeiten:

python -m venv .venv # Only required the first time you set up the project
source .venv/bin/activate

Installieren Sie die Abhängigkeiten:

pip install tflite-runtime
pip install librosa
pip install EfficientWord-Net
pip install elevenlabs
pip install "elevenlabs[pyaudio]"

Erstellen Sie nun eine neue Python-Datei mit dem Namen hotword.py und fügen Sie den folgenden Code hinzu:

hotword.py
import os
import signal
import time
from eff_word_net.streams import SimpleMicStream
from eff_word_net.engine import HotwordDetector
from eff_word_net.audio_processing import Resnet50_Arc_loss
# from eff_word_net import samples_loc
from elevenlabs.client import ElevenLabs
from elevenlabs.conversational_ai.conversation import Conversation, ConversationInitiationData
from elevenlabs.conversational_ai.default_audio_interface import DefaultAudioInterface
convai_active = False
elevenlabs = ElevenLabs()
agent_id = os.getenv("ELEVENLABS_AGENT_ID")
api_key = os.getenv("ELEVENLABS_API_KEY")
dynamic_vars = {
'user_name': 'Thor',
'greeting': 'Hey'
}
config = ConversationInitiationData(
dynamic_variables=dynamic_vars
)
base_model = Resnet50_Arc_loss()
eleven_hw = HotwordDetector(
hotword="hey_eleven",
model = base_model,
reference_file=os.path.join("hotword_refs", "hey_eleven_ref.json"),
threshold=0.7,
relaxation_time=2
)
def create_conversation():
"""Create a new conversation instance"""
return Conversation(
# API client and agent ID.
elevenlabs,
agent_id,
config=config,
# Assume auth is required when API_KEY is set.
requires_auth=bool(api_key),
# Use the default audio interface.
audio_interface=DefaultAudioInterface(),
# Simple callbacks that print the conversation to the console.
callback_agent_response=lambda response: print(f"Agent: {response}"),
callback_agent_response_correction=lambda original, corrected: print(f"Agent: {original} -> {corrected}"),
callback_user_transcript=lambda transcript: print(f"User: {transcript}"),
# Uncomment if you want to see latency measurements.
# callback_latency_measurement=lambda latency: print(f"Latency: {latency}ms"),
)
def start_mic_stream():
"""Start or restart the microphone stream"""
global mic_stream
try:
# Always create a new stream instance
mic_stream = SimpleMicStream(
window_length_secs=1.5,
sliding_window_secs=0.75,
)
mic_stream.start_stream()
print("Microphone stream started")
except Exception as e:
print(f"Error starting microphone stream: {e}")
mic_stream = None
time.sleep(1) # Wait a bit before retrying
def stop_mic_stream():
"""Stop the microphone stream safely"""
global mic_stream
try:
if mic_stream:
# SimpleMicStream doesn't have a stop_stream method
# We'll just set it to None and recreate it next time
mic_stream = None
print("Microphone stream stopped")
except Exception as e:
print(f"Error stopping microphone stream: {e}")
# Initialize microphone stream
mic_stream = None
start_mic_stream()
print("Say Hey Eleven ")
while True:
if not convai_active:
try:
# Make sure we have a valid mic stream
if mic_stream is None:
start_mic_stream()
continue
frame = mic_stream.getFrame()
result = eleven_hw.scoreFrame(frame)
if result == None:
#no voice activity
continue
if result["match"]:
print("Wakeword uttered", result["confidence"])
# Stop the microphone stream to avoid conflicts
stop_mic_stream()
# Start ConvAI Session
print("Start ConvAI Session")
convai_active = True
try:
# Create a new conversation instance
conversation = create_conversation()
# Start the session
conversation.start_session()
# Set up signal handler for graceful shutdown
def signal_handler(sig, frame):
print("Received interrupt signal, ending session...")
try:
conversation.end_session()
except Exception as e:
print(f"Error ending session: {e}")
signal.signal(signal.SIGINT, signal_handler)
# Wait for session to end
conversation_id = conversation.wait_for_session_end()
print(f"Conversation ID: {conversation_id}")
except Exception as e:
print(f"Error during conversation: {e}")
finally:
# Cleanup
convai_active = False
print("Conversation ended, cleaning up...")
# Give some time for cleanup
time.sleep(1)
# Restart microphone stream
start_mic_stream()
print("Ready for next wake word...")
except Exception as e:
print(f"Error in wake word detection: {e}")
# Try to restart microphone stream if there's an error
mic_stream = None
time.sleep(1)
start_mic_stream()

Agent-Konfiguration

1

Bei ElevenLabs anmelden

Rufen Sie elevenlabs.io auf und melden Sie sich bei Ihrem Konto an.

2

Neuen Agenten erstellen

Navigieren Sie zu Agents Platform > Agents und erstellen Sie einen neuen Agenten aus der leeren Vorlage.

3

Erste Nachricht festlegen

Legen Sie die erste Nachricht fest und geben Sie die dynamische Variable für die Plattform an.

{{greeting}} {{user_name}}, Eleven here, what's up?
4

System-Prompt festlegen

Legen Sie den System-Prompt fest. Unsere Dokumentation zu Best Practices finden Sie hier.

You are a helpful Agents Platform assistant with access to a weather tool. When users ask about
weather conditions, use the get_weather tool to fetch accurate, real-time data. The tool requires
a latitude and longitude - use your geographic knowledge to convert location names to coordinates
accurately.
Never ask users for coordinates - you must determine these yourself. Always report weather
information conversationally, referring to locations by name only. For weather requests:
1. Extract the location from the user's message
2. Convert the location to coordinates and call get_weather
3. Present the information naturally and helpfully
For non-weather queries, provide friendly assistance within your knowledge boundaries. Always be
concise, accurate, and helpful.
5

Webhook-Tool einrichten

Wir richten ein einfaches Webhook-Tool ein, das die Wetterdaten für uns abruft. Folgen Sie den Einrichtungsschritten hier, um das Tool einzurichten.

App ausführen

Legen Sie zum Ausführen der App zunächst die erforderlichen Umgebungsvariablen fest:

export ELEVENLABS_API_KEY=YOUR_API_KEY
export ELEVENLABS_AGENT_ID=YOUR_AGENT_ID

Führen Sie dann einfach den folgenden Befehl aus:

python hotword.py

Sagen Sie jetzt „Hey Eleven“, um die Unterhaltung zu starten. Viel Spaß beim Chatten.

[Optional] Eigenes Aktivierungswort trainieren

Trainingsaudio generieren

Um die Aktivierungswort-Einbettungen zu generieren, können Sie mit ElevenLabs vier Trainingsbeispiele erstellen. Navigieren Sie in Ihrer ElevenLabs-App zu Text to Speech und geben Sie Ihr Aktivierungswort ein, z. B. „Hey Eleven“. Wählen Sie eine Stimme aus und klicken Sie auf „Generate“.

Nachdem das Audio generiert wurde, laden Sie die Audiodatei herunter und speichern Sie sie in einem Ordner namens hotword_training_audio im Stammverzeichnis Ihres Projekts. Wiederholen Sie diesen Vorgang dreimal mit verschiedenen Stimmen.

Aktivierungswort trainieren

Führen Sie in Ihrem Terminal bei aktivierter virtueller Umgebung den folgenden Befehl aus, um das Aktivierungswort zu trainieren:

python -m eff_word_net.generate_reference --input-dir hotword_training_audio --output-dir hotword_refs --wakeword hey_eleven --model-type resnet_50_arc

Dadurch wird die Datei hey_eleven_ref.json im Ordner hotword_refs erstellt. Jetzt müssen Sie nur noch den Parameter reference_file in der Klasse HotwordDetector in hotword.py aktualisieren, sodass er auf die neue Referenzdatei verweist. Anschließend ist alles bereit.

Nächste Schritte