> This is a page from the ElevenLabs documentation. For a complete page index, fetch https://elevenlabs.io/docs/llms.txt. For the full documentation in a single file, fetch https://elevenlabs.io/docs/llms-full.txt.

# Integrazione LiveKit

Questa guida spiega come usare ElevenLabs Speech Engine come livello vocale per una stanza LiveKit. Un worker LiveKit Agents entra nella stanza come partecipante, si iscrive alla traccia audio dell'utente, apre un WebSocket verso Speech Engine e pubblica l'audio sintetizzato da Speech Engine nella stanza come propria traccia.

## Architettura

Speech Engine accetta due tipi di connessioni WebSocket:

* Il **WebSocket brain** a cui si connette l'API ElevenLabs. Il tuo server lo esegue con l'SDK Speech Engine (`engine.serve()` / `engine.attach()`) e riceve trascrizioni a cui rispondere.
* Il **WebSocket di conversazione** a cui si connettono i client. I browser si connettono tramite un token WebRTC; i client non browser (come un worker LiveKit Agents) si connettono tramite un URL firmato e trasmettono audio PCM raw in entrambe le direzioni.

Il worker LiveKit usa la seconda connessione. Agisce come "client" di Speech Engine per conto dei partecipanti nella stanza LiveKit.

```mermaid
sequenceDiagram
    participant Browser
    participant LK as LiveKit Room
    participant Worker as Agents Worker
    participant EL as ElevenLabs (conversation WS)
    participant Brain as Brain Server

    Browser->>LK: Join room (LiveKit token)
    Worker->>LK: Join room (dispatched)
    Worker->>EL: Open conversation WebSocket (signed URL)

    loop Conversation
        Browser->>LK: Microphone audio (Opus)
        LK->>Worker: Decoded PCM frames
        Worker->>EL: user_audio_chunk (base64 PCM)
        EL->>Brain: user_transcript
        Brain-->>EL: agent_response (streamed)
        EL->>Worker: audio (base64 PCM)
        Worker->>LK: Publish PCM frames
        LK->>Browser: Audio (Opus)
    end
```

Il server brain rimane invariato rispetto alla [guida rapida di Speech Engine](/docs/it/eleven-api/guides/cookbooks/speech-engine): il worker LiveKit sostituisce il browser come sorgente audio, ma la logica LLM resta la stessa.

## Quando usare questo schema

Usa il bridge LiveKit quando la stanza stessa fa parte dell'esperienza:

* Sessioni con più partecipanti in cui gli utenti parlano con l'agente insieme
* Implementazioni LiveKit esistenti in cui cambiare trasporto interromperebbe i client
* Agenti vocali che condividono una stanza con condivisione schermo, video o chat testuale
* Chiamate inviate da SIP a LiveKit che richiedono un agente IA in linea

Se ti serve solo un loop vocale dal browser a Speech Engine senza altri partecipanti, il client WebRTC nella [guida rapida di Speech Engine](/docs/it/eleven-api/guides/cookbooks/speech-engine#client-setup) è più semplice: Speech Engine comunica direttamente via WebRTC con il browser, senza bisogno di una stanza LiveKit.

## Prerequisiti

* Un progetto LiveKit ([LiveKit Cloud](https://cloud.livekit.io/) o un server ospitato autonomamente). Il worker richiede `LIVEKIT_URL`, `LIVEKIT_API_KEY` e `LIVEKIT_API_SECRET`.
* Un ElevenLabs Speech Engine. Segui la [guida rapida di Speech Engine](/docs/it/eleven-api/guides/cookbooks/speech-engine) per crearne uno ed eseguire il server brain.
* Python 3.9+ o Node.js 18+.

> **Note**
>
> Il worker bridge Node usa
> [`@livekit/rtc-node`](https://www.npmjs.com/package/@livekit/rtc-node), attualmente in
> Developer Preview. Per le implementazioni in produzione, preferisci il worker Python.

## Configura i formati audio di Speech Engine

`AudioStream` di LiveKit ricampiona le tracce Opus in entrata a qualsiasi frequenza di campionamento PCM richiesta, quindi puoi adattarla direttamente all'input di Speech Engine. Aggiorna Speech Engine affinché accetti PCM a 16 kHz per l'input ASR ed emetta PCM a 24 kHz per l'output TTS.

**`configure_engine.py`**

```python title="configure_engine.py"
import asyncio
import os
from elevenlabs import AsyncElevenLabs

elevenlabs = AsyncElevenLabs(api_key=os.environ["ELEVENLABS_API_KEY"])


async def update_engine():
    await elevenlabs.speech_engine.update(
        speech_engine_id="seng_8k3m9xr4hjnfg983brhmhkd98n6",
        asr={"user_input_audio_format": "pcm_16000"},
        tts={"agent_output_audio_format": "pcm_24000"},
    )


asyncio.run(update_engine())
```

**`configure-engine.mts`**

```typescript title="configure-engine.mts"
import { ElevenLabsClient } from "@elevenlabs/elevenlabs-js";
import "dotenv/config";

const elevenlabs = new ElevenLabsClient({
  apiKey: process.env.ELEVENLABS_API_KEY,
});

await elevenlabs.speechEngine.update("seng_8k3m9xr4hjnfg983brhmhkd98n6", {
  asr: { userInputAudioFormat: "pcm_16000" },
  tts: { agentOutputAudioFormat: "pcm_24000" },
});
```

Il PCM di Speech Engine è sempre little-endian con segno a 16 bit. Consulta il [riferimento dei formati audio](#audio-format-reference) per le altre frequenze supportate.

## Crea il worker bridge

Il worker è un processo a lunga esecuzione che si connette al tuo server LiveKit, attende i job, entra nelle stanze assegnate e fa da ponte per l'audio tra la stanza e Speech Engine.

#### Installa le dipendenze

**`Python`**

```bash title="Python"
pip install "livekit-agents" "livekit-api" "elevenlabs" "aiohttp" "python-dotenv"
```

**`Node`**

```bash title="Node"
npm install @livekit/agents @livekit/rtc-node @elevenlabs/elevenlabs-js ws dotenv
```

#### Genera un URL firmato Speech Engine

Il worker richiede un URL firmato a breve durata per il WebSocket di conversazione Speech Engine. L'URL firmato incorpora l'ID del motore e una firma monouso, così il worker può aprire il WebSocket senza esporre la tua chiave API.

**`bridge.py`**

```python title="bridge.py"
from elevenlabs import AsyncElevenLabs

elevenlabs = AsyncElevenLabs(api_key=os.environ["ELEVENLABS_API_KEY"])

async def signed_url() -> str:
    response = await elevenlabs.conversational_ai.conversations.get_signed_url(
        agent_id=os.environ["SPEECH_ENGINE_ID"],
    )
    return response.signed_url
```

**`bridge.mts`**

```typescript title="bridge.mts"
import { ElevenLabsClient } from "@elevenlabs/elevenlabs-js";

const elevenlabs = new ElevenLabsClient({
  apiKey: process.env.ELEVENLABS_API_KEY,
});

async function signedUrl(): Promise<string> {
  const response = await elevenlabs.conversationalAi.conversations.getSignedUrl({
    agentId: process.env.SPEECH_ENGINE_ID!,
  });
  return response.signedUrl;
}
```

#### Definisci l'entrypoint del worker

Ogni volta che il worker viene inviato a una stanza, viene eseguito il suo entrypoint. L'entrypoint si connette alla stanza, apre un WebSocket di conversazione Speech Engine e avvia due bridge audio: uno per l'audio del chiamante diretto a Speech Engine e uno per l'audio sintetizzato di ritorno.

**`bridge.py`**

```python title="bridge.py" maxLines=0
import asyncio
import base64
import json
import os

import aiohttp
from dotenv import load_dotenv
from elevenlabs import AsyncElevenLabs
from livekit import agents, rtc
from livekit.agents import JobContext, WorkerOptions, cli

load_dotenv()

elevenlabs = AsyncElevenLabs(api_key=os.environ["ELEVENLABS_API_KEY"])
SPEECH_ENGINE_ID = os.environ["SPEECH_ENGINE_ID"]

USER_INPUT_RATE = 16000
AGENT_OUTPUT_RATE = 24000


async def signed_url() -> str:
    response = await elevenlabs.conversational_ai.conversations.get_signed_url(
        agent_id=SPEECH_ENGINE_ID,
    )
    return response.signed_url


async def entrypoint(ctx: JobContext):
    el_ws_ready: asyncio.Future[aiohttp.ClientWebSocketResponse] = (
        asyncio.get_running_loop().create_future()
    )

    async def pump_user_audio(track: rtc.Track):
        el_ws = await el_ws_ready
        stream = rtc.AudioStream(
            track, sample_rate=USER_INPUT_RATE, num_channels=1,
        )
        async for event in stream:
            payload = base64.b64encode(bytes(event.frame.data)).decode()
            await el_ws.send_str(json.dumps({"user_audio_chunk": payload}))

    # Register the subscriber BEFORE ctx.connect() so we don't miss tracks
    # that get auto-subscribed during the connection handshake.
    @ctx.room.on("track_subscribed")
    def on_track_subscribed(track, publication, participant):
        if track.kind != rtc.TrackKind.KIND_AUDIO:
            return
        if participant.identity == ctx.room.local_participant.identity:
            return
        asyncio.create_task(pump_user_audio(track))

    await ctx.connect()

    # Publish a track for the agent's synthesized audio.
    source = rtc.AudioSource(sample_rate=AGENT_OUTPUT_RATE, num_channels=1)
    track = rtc.LocalAudioTrack.create_audio_track("elevenlabs-agent", source)
    await ctx.room.local_participant.publish_track(
        track,
        rtc.TrackPublishOptions(source=rtc.TrackSource.SOURCE_MICROPHONE),
    )

    # Open the Speech Engine conversation WebSocket.
    http = aiohttp.ClientSession()
    el_ws = await http.ws_connect(await signed_url())
    await el_ws.send_str(json.dumps({"type": "conversation_initiation_client_data"}))
    el_ws_ready.set_result(el_ws)

    async def el_to_room():
        async for msg in el_ws:
            if msg.type != aiohttp.WSMsgType.TEXT:
                continue
            event = json.loads(msg.data)
            etype = event.get("type")
            if etype == "audio":
                pcm = base64.b64decode(event["audio_event"]["audio_base_64"])
                samples_per_channel = len(pcm) // 2
                frame = rtc.AudioFrame(
                    pcm, AGENT_OUTPUT_RATE, 1, samples_per_channel,
                )
                await source.capture_frame(frame)
            elif etype == "interruption":
                source.clear_queue()
            elif etype == "ping":
                event_id = event.get("ping_event", {}).get("event_id")
                await el_ws.send_str(json.dumps({
                    "type": "pong", "event_id": event_id,
                }))

    pump_task = asyncio.create_task(el_to_room())

    async def cleanup():
        pump_task.cancel()
        await el_ws.close()
        await http.close()

    ctx.add_shutdown_callback(cleanup)


if __name__ == "__main__":
    cli.run_app(WorkerOptions(
        entrypoint_fnc=entrypoint,
        agent_name="elevenlabs-bridge",
    ))
```

**`bridge.mts`**

```typescript title="bridge.mts" maxLines=0
import {
  type JobContext,
  WorkerOptions,
  cli,
  defineAgent,
} from "@livekit/agents";
import {
  AudioFrame,
  AudioSource,
  AudioStream,
  LocalAudioTrack,
  RoomEvent,
  TrackKind,
  TrackPublishOptions,
  TrackSource,
} from "@livekit/rtc-node";
import { ElevenLabsClient } from "@elevenlabs/elevenlabs-js";
import WebSocket from "ws";
import { fileURLToPath } from "node:url";
import "dotenv/config";

const elevenlabs = new ElevenLabsClient({
  apiKey: process.env.ELEVENLABS_API_KEY,
});
const SPEECH_ENGINE_ID = process.env.SPEECH_ENGINE_ID!;

const USER_INPUT_RATE = 16000;
const AGENT_OUTPUT_RATE = 24000;

async function signedUrl(): Promise<string> {
  const response = await elevenlabs.conversationalAi.conversations.getSignedUrl({
    agentId: SPEECH_ENGINE_ID,
  });
  return response.signedUrl;
}

export default defineAgent({
  entry: async (ctx: JobContext) => {
    let resolveElReady: (ws: WebSocket) => void;
    const elReady = new Promise<WebSocket>((resolve) => {
      resolveElReady = resolve;
    });

    // Register the subscriber BEFORE ctx.connect() so we don't miss
    // tracks that get auto-subscribed during the connection handshake.
    ctx.room.on(RoomEvent.TrackSubscribed, (track, _pub, participant) => {
      if (track.kind !== TrackKind.KIND_AUDIO) return;
      if (participant.identity === ctx.room.localParticipant?.identity) return;

      (async () => {
        const ws = await elReady;
        const stream = new AudioStream(track, {
          sampleRate: USER_INPUT_RATE,
          numChannels: 1,
        });
        for await (const frame of stream) {
          const payload = Buffer.from(
            frame.data.buffer,
            frame.data.byteOffset,
            frame.data.byteLength,
          ).toString("base64");
          ws.send(JSON.stringify({ user_audio_chunk: payload }));
        }
      })();
    });

    await ctx.connect();

    const source = new AudioSource(AGENT_OUTPUT_RATE, 1);
    const track = LocalAudioTrack.createAudioTrack("elevenlabs-agent", source);
    const publishOptions = new TrackPublishOptions();
    publishOptions.source = TrackSource.SOURCE_MICROPHONE;
    await ctx.room.localParticipant!.publishTrack(track, publishOptions);

    const ws = new WebSocket(await signedUrl());
    await new Promise<void>((resolve, reject) => {
      ws.once("open", () => resolve());
      ws.once("error", reject);
    });
    ws.send(JSON.stringify({ type: "conversation_initiation_client_data" }));
    resolveElReady!(ws);

    // Serialize captureFrame calls — concurrent captures throw
    // InvalidState in the rtc-node native layer.
    let captureChain: Promise<unknown> = Promise.resolve();

    ws.on("message", (raw) => {
      const event = JSON.parse(raw.toString());
      if (event.type === "audio") {
        const pcm = Buffer.from(event.audio_event.audio_base_64, "base64");
        const samples = new Int16Array(
          pcm.buffer, pcm.byteOffset, pcm.byteLength / 2,
        );
        const frame = new AudioFrame(
          samples, AGENT_OUTPUT_RATE, 1, samples.length,
        );
        captureChain = captureChain
          .then(() => source.captureFrame(frame))
          .catch((err) => console.warn("captureFrame:", err.message));
      } else if (event.type === "interruption") {
        source.clearQueue();
      } else if (event.type === "ping") {
        ws.send(JSON.stringify({
          type: "pong", event_id: event.ping_event?.event_id,
        }));
      }
    });

    ctx.addShutdownCallback(async () => {
      ws.close();
    });
  },
});

cli.runApp(new WorkerOptions({
  agent: fileURLToPath(import.meta.url),
  agentName: "elevenlabs-bridge",
}));
```

Il worker esclude il proprio audio pubblicato nell'handler `track_subscribed` confrontandolo con l'identità del partecipante locale. Senza questo controllo, il worker tenterebbe di inviare il proprio audio sintetizzato di nuovo a Speech Engine.

Due dettagli sull'ordinamento sono importanti per la correttezza:

* **Tempistica del listener**: `TrackSubscribed` viene registrato prima di `ctx.connect()`. LiveKit si iscrive automaticamente alle tracce esistenti durante l'handshake della connessione e un listener registrato successivamente potrebbe non ricevere l'evento. La pompa audio attende un `Future` / `Promise` per il WebSocket Speech Engine, così può iscriversi immediatamente e inoltrare l'audio non appena la connessione viene aperta.
* **Solo TypeScript — serializzazione dell'acquisizione**: `AudioSource.captureFrame` di `@livekit/rtc-node` genera `InvalidState` se viene chiamato contemporaneamente. L'handler TypeScript serializza le acquisizioni con una catena di promise. Il singolo loop `async for el_to_room` di Python è naturalmente sequenziale e non ne ha bisogno.

#### Avvia il worker

**`Python`**

```bash title="Python"
python bridge.py dev
```

**`Node`**

```bash title="Node"
npx tsx bridge.mts dev
```

`dev` abilita il ricaricamento a caldo e i log colorati. In produzione usa `start` per i log JSON e una chiusura ordinata.

Il worker si connette al tuo server LiveKit e attende l'assegnazione di job. Non entra in alcuna stanza finché non viene inviato.

## Invia il worker a una stanza

Poiché il worker ha un `agent_name`, usa l'invio esplicito: entra nelle stanze solo quando il tuo backend glielo indica. Lo schema più semplice consiste nell'includere un `RoomAgentDispatch` nel token di accesso LiveKit che il browser usa per connettersi.

**`token_server.py`**

```python title="token_server.py"
import os

from dotenv import load_dotenv
from flask import Flask, jsonify, request
from livekit.api import AccessToken, RoomAgentDispatch, VideoGrants

load_dotenv()

app = Flask(**name**)

@app.route("/api/livekit-token")
def get_token():
room_name = request.args.get("room", "demo-room")
identity = request.args.get("identity", "web-user")

    token = (
        AccessToken(
            os.environ["LIVEKIT_API_KEY"],
            os.environ["LIVEKIT_API_SECRET"],
        )
        .with_identity(identity)
        .with_grants(VideoGrants(room_join=True, room=room_name))
        .with_room_config(
            room_configuration={
                "agents": [RoomAgentDispatch(agent_name="elevenlabs-bridge")],
            },
        )
    )

    return jsonify(token=token.to_jwt(), url=os.environ["LIVEKIT_URL"])

if **name** == "**main**":
app.run(port=3002)

```

**`token-server.mts`**

```typescript title="token-server.mts"
import express from "express";
import { AccessToken } from "livekit-server-sdk";
import "dotenv/config";

const app = express();

app.get("/api/livekit-token", async (req, res) => {
  const room = (req.query.room as string) ?? "demo-room";
  const identity = (req.query.identity as string) ?? "web-user";

  const token = new AccessToken(
    process.env.LIVEKIT_API_KEY!,
    process.env.LIVEKIT_API_SECRET!,
    { identity },
  );
  token.addGrant({ roomJoin: true, room });
  token.roomConfig = {
    agents: [{ agentName: "elevenlabs-bridge" }],
  };

  res.json({
    token: await token.toJwt(),
    url: process.env.LIVEKIT_URL,
  });
});

app.listen(3002, () => {
  console.log("Token server listening on port 3002");
});
```

Quando un browser usa questo token per creare o entrare in una stanza, LiveKit invia automaticamente il worker bridge nella stessa stanza.

## Connettiti dal browser

Il browser necessita solo del client LiveKit standard: non interagisce direttamente con Speech Engine.

**`App.tsx`**

```typescript title="App.tsx"
import { Room, RoomEvent, Track } from "livekit-client";
import { useCallback, useState } from "react";

export default function App() {
  const [room] = useState(() => new Room());

  const join = useCallback(async () => {
    const response = await fetch("/api/livekit-token");
    const { token, url } = await response.json();

    room.on(RoomEvent.TrackSubscribed, (track) => {
      if (track.kind === Track.Kind.Audio) {
        document.body.appendChild(track.attach());
      }
    });

    await room.connect(url, token);
    await room.localParticipant.setMicrophoneEnabled(true);
  }, [room]);

  return <button onClick={join}>Start conversation</button>;
}
```

Quando si fa clic sul pulsante, il browser recupera un token LiveKit, entra nella stanza con il microfono abilitato e inizia a ricevere la traccia audio dell'agente. Il worker viene inviato, apre la sua sessione Speech Engine e fa da ponte per l'audio in entrambe le direzioni.

## Riferimento dei formati audio

Speech Engine supporta i seguenti formati audio. Configurali sul motore tramite `asr.user_input_audio_format` e `tts.agent_output_audio_format`.

| Formato     | Frequenza di campionamento | Codifica                  | Note                                                              |
| ----------- | -------------------------- | ------------------------- | ----------------------------------------------------------------- |
| `pcm_8000`  | 8 kHz                      | PCM LE con segno a 16 bit | Solo input ASR.                                                   |
| `pcm_16000` | 16 kHz                     | PCM LE con segno a 16 bit | Consigliato per l'input utente LiveKit.                           |
| `pcm_22050` | 22,05 kHz                  | PCM LE con segno a 16 bit |                                                                   |
| `pcm_24000` | 24 kHz                     | PCM LE con segno a 16 bit | Consigliato per l'output dell'agente LiveKit.                     |
| `pcm_44100` | 44,1 kHz                   | PCM LE con segno a 16 bit | L'output TTS richiede il piano Independent Publisher o superiore. |
| `pcm_48000` | 48 kHz                     | PCM LE con segno a 16 bit | Solo input ASR.                                                   |
| `ulaw_8000` | 8 kHz                      | μ-law                     | Usato da Twilio Media Streams.                                    |

`AudioStream` e `AudioSource` in LiveKit gestiscono il ricampionamento per te: puoi richiedere qualsiasi frequenza di campionamento a `AudioStream` e l'SDK converte dalla traccia Opus sottostante a 48 kHz.

## Considerazioni per la produzione

* **Invio esplicito**: imposta sempre `agent_name` / `agentName` su `WorkerOptions`. L'invio automatico attiva il worker per ogni stanza creata nel tuo progetto LiveKit, cosa che raramente desideri.
* **Autenticazione del server brain**: imposta un segreto condiviso su Speech Engine e verificalo nel tuo server brain, così solo Speech Engine può raggiungere il tuo endpoint:
  ```python
  await elevenlabs.speech_engine.update(
      speech_engine_id="seng_8k3m9xr4hjnfg983brhmhkd98n6",
      speech_engine={"request_headers": {"x-api-key": os.environ["SHARED_SECRET"]}},
  )
  ```
  Il server brain verifica quindi `request.headers["x-api-key"]` prima di accettare l'upgrade WebSocket.
* **Server dei token**: genera i token LiveKit e Speech Engine lato server. Non esporre mai `LIVEKIT_API_SECRET` o `ELEVENLABS_API_KEY` al browser.
* **Igiene dell'event loop**: non eseguire operazioni vincolate alla CPU nell'event loop del worker. `AudioSource.capture_frame` e l'iterazione di `AudioStream` sono sensibili ai tempi; lunghe chiamate sincrone ritarderanno o scarteranno gli eventi di interruzione. Usa `asyncio.to_thread()` (Python) o `worker_threads` (Node) per le operazioni bloccanti.
* **Arresto**: registra `ctx.add_shutdown_callback` / `ctx.addShutdownCallback` per chiudere correttamente il WebSocket ElevenLabs. Per impostazione predefinita, la stanza (e il job) viene terminata quando l'ultimo partecipante non agente esce.

## Passaggi successivi

#### [Guida rapida di Speech Engine](/docs/it/eleven-api/guides/cookbooks/speech-engine)

Crea il server brain che risponde alle trascrizioni.

#### [Integrazione Pipecat](/docs/it/eleven-api/guides/how-to/speech-engine/pipecat-integration)

Usa Pipecat come pipeline LLM dietro Speech Engine.

#### [Riferimento SDK Python](/docs/it/eleven-api/resources/libraries/speech-engine/python-sdk-reference)

Classi, metodi ed eventi per l'SDK Python di Speech Engine.

#### [Riferimento SDK JavaScript](/docs/it/eleven-api/resources/libraries/speech-engine/javascript-sdk-reference)

Classi, metodi ed eventi per l'SDK JavaScript di Speech Engine.