Crea un asistente de voz con Agents Platform en una Raspberry Pi

Crea un asistente de voz con Agents Platform en una Raspberry Pi.

Tutorial · Da por hecho que has completado la guía de inicio rápido de ElevenAgents y que tienes una Raspberry Pi con Python instalado.

Introducción

En este tutorial aprenderás a crear un asistente de voz con Agents Platform ejecutándose en una Raspberry Pi. Al igual que los asistentes domésticos convencionales, como Alexa en Amazon Echo, Google Home o Siri en dispositivos Apple, tu asistente de voz Eleven escuchará una palabra de activación, en nuestro caso «Hey Eleven», y հետո iniciará una sesión de ElevenLabs Agents para ayudar al usuario.

Requisitos

  • Una Raspberry Pi 5 o un dispositivo similar.
  • Un micrófono y un altavoz.
  • Python 3.9 o una versión posterior instalado en tu equipo.
  • Una cuenta de ElevenLabs con una clave de API.

Configuración

Instala las dependencias

En sistemas basados en Debian puedes instalar las dependencias con:

sudo apt-get update
sudo apt-get install libportaudio2 libportaudiocpp0 portaudio19-dev libasound-dev libsndfile1-dev -y

Crea el proyecto

En tu Raspberry Pi, abre el terminal y crea un directorio nuevo para tu proyecto.

mkdir eleven-voice-assistant
cd eleven-voice-assistant

Crea un entorno virtual e instala las dependencias:

python -m venv .venv # Only required the first time you set up the project
source .venv/bin/activate

Instala las dependencias:

pip install tflite-runtime
pip install librosa
pip install EfficientWord-Net
pip install elevenlabs
pip install "elevenlabs[pyaudio]"

Ahora crea un archivo de Python llamado hotword.py y añade el siguiente código:

hotword.py
import os
import signal
import time
from eff_word_net.streams import SimpleMicStream
from eff_word_net.engine import HotwordDetector
from eff_word_net.audio_processing import Resnet50_Arc_loss
# from eff_word_net import samples_loc
from elevenlabs.client import ElevenLabs
from elevenlabs.conversational_ai.conversation import Conversation, ConversationInitiationData
from elevenlabs.conversational_ai.default_audio_interface import DefaultAudioInterface
convai_active = False
elevenlabs = ElevenLabs()
agent_id = os.getenv("ELEVENLABS_AGENT_ID")
api_key = os.getenv("ELEVENLABS_API_KEY")
dynamic_vars = {
'user_name': 'Thor',
'greeting': 'Hey'
}
config = ConversationInitiationData(
dynamic_variables=dynamic_vars
)
base_model = Resnet50_Arc_loss()
eleven_hw = HotwordDetector(
hotword="hey_eleven",
model = base_model,
reference_file=os.path.join("hotword_refs", "hey_eleven_ref.json"),
threshold=0.7,
relaxation_time=2
)
def create_conversation():
"""Create a new conversation instance"""
return Conversation(
# API client and agent ID.
elevenlabs,
agent_id,
config=config,
# Assume auth is required when API_KEY is set.
requires_auth=bool(api_key),
# Use the default audio interface.
audio_interface=DefaultAudioInterface(),
# Simple callbacks that print the conversation to the console.
callback_agent_response=lambda response: print(f"Agent: {response}"),
callback_agent_response_correction=lambda original, corrected: print(f"Agent: {original} -> {corrected}"),
callback_user_transcript=lambda transcript: print(f"User: {transcript}"),
# Uncomment if you want to see latency measurements.
# callback_latency_measurement=lambda latency: print(f"Latency: {latency}ms"),
)
def start_mic_stream():
"""Start or restart the microphone stream"""
global mic_stream
try:
# Always create a new stream instance
mic_stream = SimpleMicStream(
window_length_secs=1.5,
sliding_window_secs=0.75,
)
mic_stream.start_stream()
print("Microphone stream started")
except Exception as e:
print(f"Error starting microphone stream: {e}")
mic_stream = None
time.sleep(1) # Wait a bit before retrying
def stop_mic_stream():
"""Stop the microphone stream safely"""
global mic_stream
try:
if mic_stream:
# SimpleMicStream doesn't have a stop_stream method
# We'll just set it to None and recreate it next time
mic_stream = None
print("Microphone stream stopped")
except Exception as e:
print(f"Error stopping microphone stream: {e}")
# Initialize microphone stream
mic_stream = None
start_mic_stream()
print("Say Hey Eleven ")
while True:
if not convai_active:
try:
# Make sure we have a valid mic stream
if mic_stream is None:
start_mic_stream()
continue
frame = mic_stream.getFrame()
result = eleven_hw.scoreFrame(frame)
if result == None:
#no voice activity
continue
if result["match"]:
print("Wakeword uttered", result["confidence"])
# Stop the microphone stream to avoid conflicts
stop_mic_stream()
# Start ConvAI Session
print("Start ConvAI Session")
convai_active = True
try:
# Create a new conversation instance
conversation = create_conversation()
# Start the session
conversation.start_session()
# Set up signal handler for graceful shutdown
def signal_handler(sig, frame):
print("Received interrupt signal, ending session...")
try:
conversation.end_session()
except Exception as e:
print(f"Error ending session: {e}")
signal.signal(signal.SIGINT, signal_handler)
# Wait for session to end
conversation_id = conversation.wait_for_session_end()
print(f"Conversation ID: {conversation_id}")
except Exception as e:
print(f"Error during conversation: {e}")
finally:
# Cleanup
convai_active = False
print("Conversation ended, cleaning up...")
# Give some time for cleanup
time.sleep(1)
# Restart microphone stream
start_mic_stream()
print("Ready for next wake word...")
except Exception as e:
print(f"Error in wake word detection: {e}")
# Try to restart microphone stream if there's an error
mic_stream = None
time.sleep(1)
start_mic_stream()

Configuración del agente

1

Inicia sesión en ElevenLabs

Ve a elevenlabs.io e inicia sesión en tu cuenta.

2

Crea un agente nuevo

Ve a Agents Platform > Agents y crea un agente nuevo a partir de la plantilla en blanco.

3

Configura el primer mensaje

Configura el primer mensaje y especifica la variable dinámica para la plataforma.

{{greeting}} {{user_name}}, Eleven here, what's up?
4

Configura el prompt del sistema

Configura el prompt del sistema. Puedes consultar nuestra documentación de buenas prácticas aquí.

You are a helpful Agents Platform assistant with access to a weather tool. When users ask about
weather conditions, use the get_weather tool to fetch accurate, real-time data. The tool requires
a latitude and longitude - use your geographic knowledge to convert location names to coordinates
accurately.
Never ask users for coordinates - you must determine these yourself. Always report weather
information conversationally, referring to locations by name only. For weather requests:
1. Extract the location from the user's message
2. Convert the location to coordinates and call get_weather
3. Present the information naturally and helpfully
For non-weather queries, provide friendly assistance within your knowledge boundaries. Always be
concise, accurate, and helpful.
5

Configura una herramienta webhook

Configuraremos una sencilla herramienta webhook que obtendrá los datos meteorológicos. Sigue los pasos de configuración aquí para configurar la herramienta.

Ejecuta la app

Para ejecutar la app, primero configura las variables de entorno necesarias:

export ELEVENLABS_API_KEY=YOUR_API_KEY
export ELEVENLABS_AGENT_ID=YOUR_AGENT_ID

Después, ejecuta simplemente el siguiente comando:

python hotword.py

Ahora di «Hey Eleven» para iniciar la conversación. ¡Que disfrutes charlando!

[Opcional] Entrena tu palabra de activación personalizada

Genera audio de entrenamiento

Para generar las incrustaciones de la palabra de activación, puedes usar ElevenLabs para generar cuatro muestras de entrenamiento. Ve a Texto a Voz en tu app de ElevenLabs e introduce tu palabra de activación, por ejemplo, «Hey Eleven». Selecciona una voz y haz clic en el botón «Generate».

Cuando se haya generado el audio, descarga el archivo de audio y guárdalo en una carpeta llamada hotword_training_audio en la raíz de tu proyecto. Repite este proceso tres veces más con voces diferentes.

Entrena la palabra de activación

En el terminal, con el entorno virtual activado, ejecuta el siguiente comando para entrenar la palabra de activación:

python -m eff_word_net.generate_reference --input-dir hotword_training_audio --output-dir hotword_refs --wakeword hey_eleven --model-type resnet_50_arc

Esto generará el archivo hey_eleven_ref.json en la carpeta hotword_refs. Ahora solo tienes que actualizar el parámetro reference_file de la clase HotwordDetector en hotword.py para que apunte al nuevo archivo de referencia. ¡Y listo!

Siguientes pasos