Vai alla navigazione

Crea un assistente vocale con Agents Platform su Raspberry Pi

Crea un assistente vocale con Agents Platform su Raspberry Pi.

Tutorial · Presuppone che tu abbia completato la guida rapida di ElevenAgents e disponga di un Raspberry Pi con Python installato.

Introduzione

In questo tutorial imparerai a creare un assistente vocale con Agents Platform in esecuzione su un Raspberry Pi. Proprio come gli assistenti domestici tradizionali come Alexa su Amazon Echo, Google Home o Siri sui dispositivi Apple, il tuo assistente vocale Eleven ascolterà una hotword, nel nostro caso “Hey Eleven”, e poi avvierà una sessione ElevenLabs Agents per assistere l’utente.

Requisiti

  • Un Raspberry Pi 5 o un dispositivo simile.
  • Un microfono e un altoparlante.
  • Python 3.9 o versione successiva installato sul tuo dispositivo.
  • Un account ElevenLabs con una chiave API.

Configurazione

Installa le dipendenze

Sui sistemi basati su Debian puoi installare le dipendenze con:

sudo apt-get update
sudo apt-get install libportaudio2 libportaudiocpp0 portaudio19-dev libasound-dev libsndfile1-dev -y

Crea il progetto

Sul tuo Raspberry Pi, apri il terminale e crea una nuova directory per il progetto.

mkdir eleven-voice-assistant
cd eleven-voice-assistant

Crea un nuovo ambiente virtuale e installa le dipendenze:

python -m venv .venv # Only required the first time you set up the project
source .venv/bin/activate

Installa le dipendenze:

pip install tflite-runtime
pip install librosa
pip install EfficientWord-Net
pip install elevenlabs
pip install "elevenlabs[pyaudio]"

Ora crea un nuovo file Python chiamato hotword.py e aggiungi il codice seguente:

hotword.py
import os
import signal
import time
from eff_word_net.streams import SimpleMicStream
from eff_word_net.engine import HotwordDetector
from eff_word_net.audio_processing import Resnet50_Arc_loss
# from eff_word_net import samples_loc
from elevenlabs.client import ElevenLabs
from elevenlabs.conversational_ai.conversation import Conversation, ConversationInitiationData
from elevenlabs.conversational_ai.default_audio_interface import DefaultAudioInterface
convai_active = False
elevenlabs = ElevenLabs()
agent_id = os.getenv("ELEVENLABS_AGENT_ID")
api_key = os.getenv("ELEVENLABS_API_KEY")
dynamic_vars = {
'user_name': 'Thor',
'greeting': 'Hey'
}
config = ConversationInitiationData(
dynamic_variables=dynamic_vars
)
base_model = Resnet50_Arc_loss()
eleven_hw = HotwordDetector(
hotword="hey_eleven",
model = base_model,
reference_file=os.path.join("hotword_refs", "hey_eleven_ref.json"),
threshold=0.7,
relaxation_time=2
)
def create_conversation():
"""Create a new conversation instance"""
return Conversation(
# API client and agent ID.
elevenlabs,
agent_id,
config=config,
# Assume auth is required when API_KEY is set.
requires_auth=bool(api_key),
# Use the default audio interface.
audio_interface=DefaultAudioInterface(),
# Simple callbacks that print the conversation to the console.
callback_agent_response=lambda response: print(f"Agent: {response}"),
callback_agent_response_correction=lambda original, corrected: print(f"Agent: {original} -> {corrected}"),
callback_user_transcript=lambda transcript: print(f"User: {transcript}"),
# Uncomment if you want to see latency measurements.
# callback_latency_measurement=lambda latency: print(f"Latency: {latency}ms"),
)
def start_mic_stream():
"""Start or restart the microphone stream"""
global mic_stream
try:
# Always create a new stream instance
mic_stream = SimpleMicStream(
window_length_secs=1.5,
sliding_window_secs=0.75,
)
mic_stream.start_stream()
print("Microphone stream started")
except Exception as e:
print(f"Error starting microphone stream: {e}")
mic_stream = None
time.sleep(1) # Wait a bit before retrying
def stop_mic_stream():
"""Stop the microphone stream safely"""
global mic_stream
try:
if mic_stream:
# SimpleMicStream doesn't have a stop_stream method
# We'll just set it to None and recreate it next time
mic_stream = None
print("Microphone stream stopped")
except Exception as e:
print(f"Error stopping microphone stream: {e}")
# Initialize microphone stream
mic_stream = None
start_mic_stream()
print("Say Hey Eleven ")
while True:
if not convai_active:
try:
# Make sure we have a valid mic stream
if mic_stream is None:
start_mic_stream()
continue
frame = mic_stream.getFrame()
result = eleven_hw.scoreFrame(frame)
if result == None:
#no voice activity
continue
if result["match"]:
print("Wakeword uttered", result["confidence"])
# Stop the microphone stream to avoid conflicts
stop_mic_stream()
# Start ConvAI Session
print("Start ConvAI Session")
convai_active = True
try:
# Create a new conversation instance
conversation = create_conversation()
# Start the session
conversation.start_session()
# Set up signal handler for graceful shutdown
def signal_handler(sig, frame):
print("Received interrupt signal, ending session...")
try:
conversation.end_session()
except Exception as e:
print(f"Error ending session: {e}")
signal.signal(signal.SIGINT, signal_handler)
# Wait for session to end
conversation_id = conversation.wait_for_session_end()
print(f"Conversation ID: {conversation_id}")
except Exception as e:
print(f"Error during conversation: {e}")
finally:
# Cleanup
convai_active = False
print("Conversation ended, cleaning up...")
# Give some time for cleanup
time.sleep(1)
# Restart microphone stream
start_mic_stream()
print("Ready for next wake word...")
except Exception as e:
print(f"Error in wake word detection: {e}")
# Try to restart microphone stream if there's an error
mic_stream = None
time.sleep(1)
start_mic_stream()

Configurazione dell’agente

1

Accedi a ElevenLabs

Vai su elevenlabs.io e accedi al tuo account.

2

Crea un nuovo agente

Vai a Agents Platform > Agents e crea un nuovo agente dal template vuoto.

3

Imposta il primo messaggio

Imposta il primo messaggio e specifica la variabile dinamica per la piattaforma.

{{greeting}} {{user_name}}, Eleven here, what's up?
4

Imposta il prompt di sistema

Imposta il prompt di sistema. Trovi qui la nostra documentazione sulle best practice qui.

You are a helpful Agents Platform assistant with access to a weather tool. When users ask about
weather conditions, use the get_weather tool to fetch accurate, real-time data. The tool requires
a latitude and longitude - use your geographic knowledge to convert location names to coordinates
accurately.
Never ask users for coordinates - you must determine these yourself. Always report weather
information conversationally, referring to locations by name only. For weather requests:
1. Extract the location from the user's message
2. Convert the location to coordinates and call get_weather
3. Present the information naturally and helpfully
For non-weather queries, provide friendly assistance within your knowledge boundaries. Always be
concise, accurate, and helpful.
5

Configura uno strumento webhook

Configureremo un semplice strumento webhook che recupererà per noi i dati meteo. Segui i passaggi di configurazione qui per configurarlo.

Esegui l’app

Per eseguire l’app, imposta prima le variabili d’ambiente richieste:

export ELEVENLABS_API_KEY=YOUR_API_KEY
export ELEVENLABS_AGENT_ID=YOUR_AGENT_ID

Poi esegui questo comando:

python hotword.py

Ora di’ “Hey Eleven” per avviare la conversazione. Buona conversazione!

[Facoltativo] Addestra la tua hotword personalizzata

Genera l’audio di addestramento

Per generare gli embedding della hotword, puoi usare ElevenLabs per generare quattro campioni di addestramento. Vai a Text to Speech nella tua app ElevenLabs e inserisci la tua hotword, ad esempio “Hey Eleven”. Seleziona una voce e fai clic sul pulsante “Genera”.

Dopo aver generato l’audio, scarica il file audio e salvalo in una cartella chiamata hotword_training_audio nella root del progetto. Ripeti il processo altre tre volte con voci diverse.

Addestra la hotword

Nel terminale, con l’ambiente virtuale attivato, esegui il comando seguente per addestrare la hotword:

python -m eff_word_net.generate_reference --input-dir hotword_training_audio --output-dir hotword_refs --wakeword hey_eleven --model-type resnet_50_arc

Verrà generato il file hey_eleven_ref.json nella cartella hotword_refs. Ora devi solo aggiornare il parametro reference_file nella classe HotwordDetector in hotword.py affinché punti al nuovo file di riferimento e sei pronto!

Passaggi successivi