> This is a page from the ElevenLabs documentation. For a complete page index, fetch https://elevenlabs.io/docs/llms.txt. For the full documentation in a single file, fetch https://elevenlabs.io/docs/llms-full.txt.

# WebSocket multi-contesto

> **Avanzato**
>
> L'orchestrazione di agenti vocali tramite questa API WebSocket multi-contesto è un'attività complessa,
> consigliata agli sviluppatori esperti. Per una soluzione più gestita, scopri il nostro [prodotto piattaforma Agents](/docs/it/eleven-agents/overview), che semplifica molte di queste sfide.

## Panoramica

Per creare agenti vocali reattivi, devi poter gestire dinamicamente i flussi audio, gestire le interruzioni in modo fluido e mantenere un parlato naturale nei vari turni di conversazione. La nostra API WebSocket multi-contesto per Text to Speech (TTS) è progettata specificamente per questi scenari.

Questa API estende la nostra [funzionalità WebSocket TTS standard](/docs/it/eleven-api/guides/how-to/websockets/realtime-tts) introducendo il concetto di "contesti". Ogni contesto funziona come un flusso indipendente di generazione audio all'interno di una singola connessione WebSocket. Ciò ti consente di:

* Gestire più linee di parlato contemporaneamente, ad esempio l'agente che parla mentre prepara una risposta a un'interruzione dell'utente.
* Gestire senza interruzioni le sovrapposizioni dell'utente chiudendo un contesto di parlato esistente e avviandone uno nuovo.
* Mantenere la coerenza prosodica per gli enunciati nello stesso contesto logico.
* Ottimizzare l'uso delle risorse chiudendo selettivamente i contesti che non sono più necessari.

> **Warning**
>
> L'API WebSocket multi-contesto è ottimizzata per le applicazioni vocali e non è pensata per
> generare contemporaneamente più flussi audio non correlati. Per questo, ogni connessione è limitata
> a 5 contesti simultanei.

Questa guida ti accompagnerà nella connessione al WebSocket multi-contesto, nella gestione dei contesti e nell'applicazione delle best practice per creare agenti vocali coinvolgenti.

### Best practice

> **Note**
>
> Queste best practice sono essenziali per creare agenti vocali reattivi ed efficienti con la nostra
> API WebSocket multi-contesto.

#### Usa una singola connessione WebSocket

Stabilisci una connessione WebSocket per ogni sessione dell'utente finale. In questo modo riduci
overhead e latenza rispetto alla creazione di più connessioni. All'interno di questa singola
connessione, puoi gestire più contesti per diverse parti della conversazione.

#### Trasmetti le risposte in blocchi, genera frasi

Quando generi risposte lunghe, trasmetti il testo in blocchi più piccoli e usa il flag `flush: true`
alla fine delle frasi complete. Questo migliora la qualità dell'audio generato e la reattività.

#### Gestisci le interruzioni in modo fluido

Trasmetti il testo in un contesto finché non si verifica un'interruzione, quindi crea un nuovo contesto
e chiudi quello esistente. Questo approccio garantisce transizioni fluide quando cambia il flusso della conversazione.

#### Gestisci il ciclo di vita dei contesti

Chiudi tempestivamente i contesti inutilizzati. Il server può mantenere fino a 5 contesti simultanei per
connessione, ma dovresti chiuderli quando non sono più necessari.

#### Evita i timeout dei contesti

Per impostazione predefinita, i contesti vanno in timeout dopo 20 secondi e vengono chiusi automaticamente. Il timeout di
inattività è un parametro a livello di websocket che si applica a tutti i contesti e, se necessario, può arrivare a 180 secondi.
Invia un messaggio di testo vuoto in un contesto per reimpostare il timer del timeout.

### Gestione delle interruzioni

Quando un utente interrompe il tuo agente, dovresti [chiudere il contesto corrente](/docs/it/api-reference/text-to-speech/v-1-text-to-speech-voice-id-multi-stream-input#send.Close-Context) e [crearne uno nuovo](/docs/it/api-reference/text-to-speech/v-1-text-to-speech-voice-id-multi-stream-input#send.Initialise-Context):

```python
async def handle_interruption(websocket, old_context_id, new_context_id, new_response):
    # Close the existing context that was interrupted
    await websocket.send(json.dumps({
        "context_id": old_context_id,
        "close_context": True
    }))
    print(f"Closed interrupted context '{old_context_id}'")

    # Create a new context for the new response
    await send_text_in_context(websocket, new_response, new_context_id)
```

```javascript
function handleInterruption(websocket: WebSocket, oldContextId: string, newContextId: string, newResponse: string) {
  // Close the existing context that was interrupted
  websocket.send(JSON.stringify({
    context_id: oldContextId,
    close_context: true
  }));
  console.log(`Closed interrupted context '${oldContextId}'`);

  // Create a new context for the new response
  sendTextInContext(websocket, newResponse, newContextId);
}
```

### Mantenere attivo un contesto

I contesti vanno automaticamente in timeout dopo [20 secondi di inattività per impostazione predefinita](/docs/it/api-reference/text-to-speech/v-1-text-to-speech-voice-id-multi-stream-input#request.query.inactivity_timeout). Se devi mantenere attivo un contesto senza generare testo, ad esempio durante un ritardo di elaborazione, puoi inviare un messaggio di testo vuoto per reimpostare il timer del timeout.

```python
async def keep_context_alive(websocket, context_id):
    await websocket.send(json.dumps({
        "context_id": context_id,
        "text": ""
    }))
```

```javascript
function handleInterruption(websocket: WebSocket, contextId: string) {
  // Close the existing context that was interrupted
  websocket.send(JSON.stringify({
    context_id: oldContextId,
    text: ""
  }));
}
```

### Chiusura della connessione WebSocket

Al termine della conversazione, puoi ripulire tutti i contesti [chiudendo il socket](/docs/it/api-reference/text-to-speech/v-1-text-to-speech-voice-id-multi-stream-input#send.Close-Socket):

```python
async def end_conversation(websocket):
    # This will close all contexts and close the connection
    await websocket.send(json.dumps({
        "close_socket": True
    }))
    print("Ending conversation and closing WebSocket")`
```

```javascript
function endConversation(websocket: WebSocket) {
  // This will close all contexts and close the connection
  websocket.send(JSON.stringify({
    close_socket: true
  }));
  console.log("Ending conversation and closing WebSocket");
}
```

## Esempio completo di agente conversazionale

### Requisiti

* Un account ElevenLabs con una chiave API (scopri come [trovare la tua chiave API](/docs/it/api-reference/authentication)).
* Python o Node.js, oppure un altro runtime JavaScript, installato sul tuo computer.
* Familiarità con la comunicazione WebSocket. Ti consigliamo di leggere la nostra [guida sullo streaming WebSocket standard](/docs/it/eleven-api/guides/how-to/websockets/realtime-tts) per i concetti di base.

### Configurazione

Installa le dipendenze necessarie per il linguaggio scelto:

```python
pip install python-dotenv websockets
```

```javascript
npm install dotenv ws
for TypeScript, you might also want types:
npm install @types/dotenv @types/ws --save-dev
```

Crea un file .env nella directory del progetto per archiviare la chiave API:

**`.env`**

```python .env
ELEVENLABS_API_KEY=your_elevenlabs_api_key_here
```

### Esempio di agente vocale

> **Note**
>
> Questo codice è fornito come esempio e non è destinato all'uso in produzione

```python maxLines=100
import os
import json
import asyncio
import websockets
from dotenv import load_dotenv

load_dotenv()
ELEVENLABS_API_KEY = os.getenv("ELEVENLABS_API_KEY")
VOICE_ID = "your_voice_id"
MODEL_ID = "eleven_flash_v2_5"

WEBSOCKET_URI = f"wss://api.elevenlabs.io/v1/text-to-speech/{VOICE_ID}/multi-stream-input?model_id={MODEL_ID}"

async def send_text_in_context(websocket, text, context_id, voice_settings=None):
    """Send text to be synthesized in the specified context."""
    message = {
        "text": text,
        "context_id": context_id,
    }

    # Only include voice_settings for the first message in a context
    if voice_settings:
        message["voice_settings"] = voice_settings

    await websocket.send(json.dumps(message))

async def continue_context(websocket, text, context_id):
    """Add more text to an existing context."""
    await websocket.send(json.dumps({
        "text": text,
        "context_id": context_id
    }))

async def flush_context(websocket, context_id):
    """Force generation of any buffered audio in the context."""
    await websocket.send(json.dumps({
        "context_id": context_id,
        "flush": True
    }))

async def handle_interruption(websocket, old_context_id, new_context_id, new_response):
    """Handle user interruption by closing current context and starting a new one."""
    # Close the existing context that was interrupted
    await websocket.send(json.dumps({
        "context_id": old_context_id,
        "close_context": True
    }))

    # Create a new context for the new response
    await send_text_in_context(websocket, new_response, new_context_id)

async def end_conversation(websocket):
    """End the conversation and close the WebSocket connection."""
    await websocket.send(json.dumps({
        "close_socket": True
    }))

async def receive_messages(websocket):
    """Process incoming WebSocket messages."""
    context_audio = {}
    try:
        async for message in websocket:
            data = json.loads(message)
            context_id = data.get("contextId", "default")

            if data.get("audio"):
                print(f"Received audio for context '{context_id}'")

            if data.get("is_final"):
                print(f"Context '{context_id}' completed")
    except (websockets.exceptions.ConnectionClosed, asyncio.CancelledError):
        print("Message receiving stopped")

async def conversation_agent_demo():
    """Run a complete conversational agent demo."""
    # Connect with API key in headers
    async with websockets.connect(
        WEBSOCKET_URI,
        max_size=16 * 1024 * 1024,
        additional_headers={"xi-api-key": ELEVENLABS_API_KEY}
    ) as websocket:
        # Start receiving messages in background
        receive_task = asyncio.create_task(receive_messages(websocket))

        # Initial agent response
        await send_text_in_context(
            websocket,
            "Hello! I'm your virtual assistant. I can help you with a wide range of topics. What would you like to know about today?",
            "greeting"
        )

        # Wait a bit (simulating user listening)
        await asyncio.sleep(2)

        # Simulate user interruption
        print("USER INTERRUPTS: 'Can you tell me about the weather?'")

        # Handle the interruption by closing current context and starting new one
        await handle_interruption(
            websocket,
            "greeting",
            "weather_response",
            "I'd be happy to tell you about the weather. Currently in your area, it's 72 degrees and sunny with a slight chance of rain later this afternoon."
        )

        # Add more to the weather context
        await continue_context(
            websocket,
            " If you're planning to go outside, you might want to bring a light jacket just in case.",
            "weather_response"
        )

        # Flush at the end of this turn to ensure all audio is generated
        await flush_context(websocket, "weather_response")

        # Wait a bit (simulating user listening)
        await asyncio.sleep(3)

        # Simulate user asking another question
        print("USER: 'What about tomorrow?'")

        # Create a new context for this response
        await send_text_in_context(
            websocket,
            "Tomorrow's forecast shows temperatures around 75 degrees with partly cloudy skies. It should be a beautiful day overall!",
            "tomorrow_weather"
        )

        # Flush and close this context
        await flush_context(websocket, "tomorrow_weather")
        await websocket.send(json.dumps({
            "context_id": "tomorrow_weather",
            "close_context": True
        }))

        # End the conversation
        await asyncio.sleep(2)
        await end_conversation(websocket)

        # Cancel the receive task
        receive_task.cancel()
        try:
            await receive_task
        except asyncio.CancelledError:
            pass

if __name__ == "__main__":
    asyncio.run(conversation_agent_demo())

```

```javascript maxLines=100
// Import required modules
import dotenv from "dotenv";
import fs from "fs";
import WebSocket from "ws";

// Load environment variables
dotenv.config();
const ELEVENLABS_API_KEY = process.env.ELEVENLABS_API_KEY;
const VOICE_ID = "your_voice_id";
const MODEL_ID = "eleven_flash_v2_5";

const WEBSOCKET_URI = `wss://api.elevenlabs.io/v1/text-to-speech/${VOICE_ID}/multi-stream-input?model_id=${MODEL_ID}`;

// Function to send text in a specific context
function sendTextInContext(websocket, text, contextId, voiceSettings = null) {
  const message = {
    text: text,
    context_id: contextId,
  };

  // Only include voice_settings for the first message in a context
  if (voiceSettings) {
    message.voice_settings = voiceSettings;
  }

  websocket.send(JSON.stringify(message));
}

// Function to continue an existing context with more text
function continueContext(websocket, text, contextId) {
  websocket.send(
    JSON.stringify({
      text: text,
      context_id: contextId,
    })
  );
}

// Function to flush a context, forcing generation of buffered audio
function flushContext(websocket, contextId) {
  websocket.send(
    JSON.stringify({
      context_id: contextId,
      flush: true,
    })
  );
}

// Function to handle user interruption
function handleInterruption(websocket, oldContextId, newContextId, newResponse) {
  // Close the existing context that was interrupted
  websocket.send(
    JSON.stringify({
      context_id: oldContextId,
      close_context: true,
    })
  );

  // Create a new context for the new response
  sendTextInContext(websocket, newResponse, newContextId);
}

// Function to end the conversation and close the connection
function endConversation(websocket) {
  websocket.send(
    JSON.stringify({
      close_socket: true,
    })
  );
}

// Function to run the conversation agent demo
async function conversationAgentDemo() {
  // Connect to WebSocket with API key in headers
  const websocket = new WebSocket(WEBSOCKET_URI, {
    headers: {
      "xi-api-key": ELEVENLABS_API_KEY,
    },
    maxPayload: 16 * 1024 * 1024,
  });

  // Set up event handlers
  websocket.on("open", () => {
    // Initial agent response
    sendTextInContext(
      websocket,
      "Hello! I'm your virtual assistant. I can help you with a wide range of topics. What would you like to know about today?",
      "greeting"
    );

    // Simulate wait time (user listening)
    setTimeout(() => {
      // Simulate user interruption
      console.log("USER INTERRUPTS: 'Can you tell me about the weather?'");

      // Handle the interruption
      handleInterruption(
        websocket,
        "greeting",
        "weather_response",
        "I'd be happy to tell you about the weather. Currently in your area, it's 72 degrees and sunny with a slight chance of rain later this afternoon."
      );

      // Add more to the weather context
      setTimeout(() => {
        continueContext(
          websocket,
          " If you're planning to go outside, you might want to bring a light jacket just in case.",
          "weather_response"
        );

        // Flush at the end of this turn
        flushContext(websocket, "weather_response");

        // Simulate wait time (user listening)
        setTimeout(() => {
          // Simulate user asking another question
          console.log("USER: 'What about tomorrow?'");

          // Create a new context for this response
          sendTextInContext(
            websocket,
            "Tomorrow's forecast shows temperatures around 75 degrees with partly cloudy skies. It should be a beautiful day overall!",
            "tomorrow_weather"
          );

          // Flush and close this context
          flushContext(websocket, "tomorrow_weather");
          websocket.send(
            JSON.stringify({
              context_id: "tomorrow_weather",
              close_context: true,
            })
          );

          // End the conversation
          setTimeout(() => {
            endConversation(websocket);
          }, 2000);
        }, 3000);
      }, 500);
    }, 2000);
  });

  // Handle incoming messages
  websocket.on("message", (message) => {
    try {
      const data = JSON.parse(message);
      const contextId = data.contextId || "default";

      if (data.audio) {
        //do stuff
      }

      if (data.is_final) {
        console.log(`Context '${contextId}' completed`);
      }
    } catch (error) {
      console.error("Error parsing message:", error);
    }
  });

  // Handle WebSocket closure
  websocket.on("close", () => {
    console.log("WebSocket connection closed");
  });

  // Handle WebSocket errors
  websocket.on("error", (error) => {
    console.error("WebSocket error:", error);
  });
}

// Run the demo
conversationAgentDemo();
```

## Passaggi successivi

#### [ElevenAgents](/docs/it/eleven-agents/quickstart)

Crea agenti vocali pronti per la produzione con la piattaforma completa ElevenAgents.

#### [Capire lo streaming audio](/docs/it/eleven-api/concepts/audio-streaming)

Scopri come funziona lo streaming WebSocket e cosa influisce sulla latenza.