WebSocket

AI 에이전트와 실시간으로 대화형 음성 대화 만들기

이 문서는 ElevenLabs WebSocket API를 직접 통합하는 개발자를 위한 문서입니다. 더 편리하게 사용하려면 ElevenLabs에서 제공하는 공식 SDK를 사용하는 것을 고려해 보세요.

ElevenAgents WebSocket API를 사용하면 AI 에이전트와 실시간으로 대화형 음성 대화를 만들 수 있습니다. WebSocket 연결을 설정하면 오디오 입력을 전송하고 오디오 응답을 실시간으로 받아, 실제와 같은 대화 경험을 구현할 수 있습니다.

엔드포인트: wss://api.elevenlabs.io/v1/convai/conversation?agent_id={agent_id}

인증

에이전트 ID 사용

공개 에이전트의 경우 추가 인증 없이 WebSocket URL에서 agent_id를 직접 사용할 수 있습니다.

wss://api.elevenlabs.io/v1/convai/conversation?agent_id=<your-agent-id>

서명된 URL 사용

비공개 에이전트 또는 인증이 필요한 대화의 경우, API 키를 사용해 ElevenLabs API와 안전하게 통신하는 서명된 URL을 서버에서 발급받으세요.

cURL 사용 예시

요청:

curl -X GET "https://api.elevenlabs.io/v1/convai/conversation/get-signed-url?agent_id=<your-agent-id>" \
-H "xi-api-key: <your-api-key>"

응답:

{
"signed_url": "wss://api.elevenlabs.io/v1/convai/conversation?agent_id=<your-agent-id>&token=<token>"
}
클라이언트 측에 ElevenLabs API 키를 절대 노출하지 마세요.

WebSocket 이벤트

클라이언트-서버 이벤트

클라이언트에서 서버로 다음 이벤트를 보낼 수 있습니다.

대화 상태를 업데이트하는, 대화를 중단하지 않는 컨텍스트 정보를 전송합니다. 진행 중인 대화 흐름을 방해하지 않고 추가 컨텍스트를 제공할 수 있습니다.

{
"type": "contextual_update",
"text": "User clicked on pricing page"
}

사용 사례:

  • 사용자 상태 또는 환경설정 업데이트
  • 환경 컨텍스트 제공
  • 배경 정보 추가
  • 사용자 인터페이스 상호작용 추적

핵심 사항:

  • 현재 대화 흐름을 중단하지 않음
  • 업데이트는 대화 기록에서 도구 호출로 반영됨
  • 자연스러운 대화를 끊지 않고 컨텍스트 유지에 도움

컨텍스트 업데이트는 비동기적으로 처리되며 서버의 직접 응답이 필요하지 않습니다.

Next.js 구현 예시

이 예시는 ElevenLabs WebSocket API를 사용하여 Next.js에서 WebSocket 기반 대화형 에이전트 클라이언트를 구현하는 방법을 보여 줍니다.

이 예시에서는 마이크 입력 처리를 위해 voice-stream 패키지를 사용하지만, 오디오를 캡처하고 인코딩하는 자체 솔루션을 구현할 수 있습니다. 여기서는 ElevenLabs API를 사용하는 WebSocket 연결 및 이벤트 처리 시연에 중점을 둡니다.

1

필수 종속성 설치

먼저 필요한 패키지를 설치합니다.

npm install voice-stream

voice-stream 패키지는 마이크 액세스와 오디오 스트리밍을 처리하며, ElevenLabs API에서 요구하는 대로 오디오를 자동으로 base64 형식으로 인코딩합니다.

이 예시는 스타일링에 Tailwind CSS를 사용합니다. Next.js 프로젝트에 Tailwind를 추가하려면 다음을 실행하세요.

npm install -D tailwindcss postcss autoprefixer
npx tailwindcss init -p

그런 다음 Next.js용 공식 Tailwind CSS 설정 가이드를 따르세요.

또는 className 속성을 자체 CSS 스타일로 대체할 수 있습니다.

2

WebSocket 타입 생성

WebSocket 이벤트의 타입을 정의합니다.

app/types/websocket.ts
type BaseEvent = {
type: string;
};
type UserTranscriptEvent = BaseEvent & {
type: "user_transcript";
user_transcription_event: {
user_transcript: string;
};
};
type AgentResponseEvent = BaseEvent & {
type: "agent_response";
agent_response_event: {
agent_response: string;
};
};
type AgentResponseCorrectionEvent = BaseEvent & {
type: "agent_response_correction";
agent_response_correction_event: {
original_agent_response: string;
corrected_agent_response: string;
};
};
type AudioResponseEvent = BaseEvent & {
type: "audio";
audio_event: {
audio_base_64: string;
event_id: number;
alignment: {
chars: string[];
char_durations_ms: number[];
char_start_times_ms: number[];
};
};
};
type InterruptionEvent = BaseEvent & {
type: "interruption";
interruption_event: {
reason: string;
};
};
type PingEvent = BaseEvent & {
type: "ping";
ping_event: {
event_id: number;
ping_ms?: number;
};
};
type AgentChatResponsePartEvent = BaseEvent & {
type: "agent_chat_response_part";
text_response_part: {
type: "start" | "delta" | "stop";
text: string;
event_id: number;
response_id: string;
};
};
export type ElevenLabsWebSocketEvent =
| UserTranscriptEvent
| AgentResponseEvent
| AgentResponseCorrectionEvent
| AudioResponseEvent
| InterruptionEvent
| PingEvent
| AgentChatResponsePartEvent;
3

WebSocket 훅 생성

WebSocket 연결을 관리할 커스텀 훅을 만듭니다.

app/hooks/useAgentConversation.ts
'use client';
import { useCallback, useEffect, useRef, useState } from 'react';
import { useVoiceStream } from 'voice-stream';
import type { ElevenLabsWebSocketEvent } from '../types/websocket';
const sendMessage = (websocket: WebSocket, request: object) => {
if (websocket.readyState !== WebSocket.OPEN) {
return;
}
websocket.send(JSON.stringify(request));
};
export const useAgentConversation = () => {
const websocketRef = useRef<WebSocket>(null);
const [isConnected, setIsConnected] = useState<boolean>(false);
const { startStreaming, stopStreaming } = useVoiceStream({
onAudioChunked: (audioData) => {
if (!websocketRef.current) return;
sendMessage(websocketRef.current, {
user_audio_chunk: audioData,
});
},
});
const startConversation = useCallback(async () => {
if (isConnected) return;
const websocket = new WebSocket("wss://api.elevenlabs.io/v1/convai/conversation");
websocket.onopen = async () => {
setIsConnected(true);
sendMessage(websocket, {
type: "conversation_initiation_client_data",
});
await startStreaming();
};
websocket.onmessage = async (event) => {
const data = JSON.parse(event.data) as ElevenLabsWebSocketEvent;
// Handle ping events to keep connection alive
if (data.type === "ping") {
setTimeout(() => {
sendMessage(websocket, {
type: "pong",
event_id: data.ping_event.event_id,
});
}, data.ping_event.ping_ms);
}
if (data.type === "user_transcript") {
const { user_transcription_event } = data;
console.log("User transcript", user_transcription_event.user_transcript);
}
if (data.type === "agent_response") {
const { agent_response_event } = data;
console.log("Agent response", agent_response_event.agent_response);
}
if (data.type === "agent_response_correction") {
const { agent_response_correction_event } = data;
console.log("Agent response correction", agent_response_correction_event.corrected_agent_response);
}
if (data.type === "interruption") {
// Handle interruption
}
if (data.type === "audio") {
const { audio_event } = data;
// Implement your own audio playback system here
// Note: You'll need to handle audio queuing to prevent overlapping
// as the WebSocket sends audio events in chunks
}
if (data.type === "agent_chat_response_part") {
const { text_response_part } = data;
const { type: partType, text, response_id } = text_response_part;
// Handle the agent's response text as it is generated. Enable
// agent_chat_response_part in the agent's client_events to receive
// this during voice conversations.
console.log("Chat response part:", partType, text, response_id);
}
};
websocketRef.current = websocket;
websocket.onclose = async () => {
websocketRef.current = null;
setIsConnected(false);
stopStreaming();
};
}, [startStreaming, isConnected, stopStreaming]);
const stopConversation = useCallback(async () => {
if (!websocketRef.current) return;
websocketRef.current.close();
}, []);
useEffect(() => {
return () => {
if (websocketRef.current) {
websocketRef.current.close();
}
};
}, []);
return {
startConversation,
stopConversation,
isConnected,
};
};
4

대화 컴포넌트 생성

WebSocket 훅을 사용할 컴포넌트를 만듭니다.

app/components/Conversation.tsx
'use client';
import { useCallback } from 'react';
import { useAgentConversation } from '../hooks/useAgentConversation';
export function Conversation() {
const { startConversation, stopConversation, isConnected } = useAgentConversation();
const handleStart = useCallback(async () => {
try {
await navigator.mediaDevices.getUserMedia({ audio: true });
await startConversation();
} catch (error) {
console.error('Failed to start conversation:', error);
}
}, [startConversation]);
return (
<div className="flex flex-col items-center gap-4">
<div className="flex gap-2">
<button
onClick={handleStart}
disabled={isConnected}
className="px-4 py-2 bg-blue-500 text-white rounded disabled:bg-gray-300"
>
Start Conversation
</button>
<button
onClick={stopConversation}
disabled={!isConnected}
className="px-4 py-2 bg-red-500 text-white rounded disabled:bg-gray-300"
>
Stop Conversation
</button>
</div>
<div className="flex flex-col items-center">
<p>Status: {isConnected ? 'Connected' : 'Disconnected'}</p>
</div>
</div>
);
}

다음 단계

  1. 오디오 재생: Web Audio API 또는 라이브러리를 사용하여 자체 오디오 재생 시스템을 구현하세요. WebSocket이 오디오 이벤트를 청크 단위로 전송하므로, 겹침을 방지하기 위해 오디오 큐를 처리해야 합니다.
  2. 오류 처리: 재시도 로직과 오류 복구 메커니즘을 추가하세요.
  3. UI 피드백: 음성 활동 및 연결 상태를 위한 시각적 표시기를 추가하세요.

지연 시간 관리

원활한 대화를 위해 다음 전략을 구현하세요.

  • 적응형 버퍼링: 네트워크 상태에 따라 오디오 버퍼링을 조정합니다.
  • 지터 버퍼: 패킷 도착 시간의 변동을 완화하는 지터 버퍼를 구현합니다.
  • 핑-퐁 모니터링: ping 및 pong 이벤트를 사용해 왕복 시간을 측정하고 그에 따라 조정합니다.

보안 모범 사례

  • API 키를 정기적으로 교체하고 환경 변수에 저장하세요.
  • 악용을 방지하기 위해 속도 제한을 구현하세요.
  • 사용자에게 마이크 액세스를 요청할 때 목적을 명확하게 설명하세요.
  • 최적화된 청킹: 지연 시간과 효율성의 균형을 맞추도록 오디오 청크 길이를 조정하세요.

추가 리소스