क्लाइंट इवेंट्स

कन्वर्सेशनल ऐप्लिकेशन के दौरान क्लाइंट को मिलने वाले रीयल-टाइम इवेंट्स को समझें और हैंडल करें।

क्लाइंट इवेंट्स सर्वर से क्लाइंट को भेजे जाने वाले सिस्टम-लेवल इवेंट्स हैं, जो रीयल-टाइम कम्यूनिकेशन को आसान बनाते हैं। ये इवेंट्स क्लाइंट ऐप्लिकेशन को ऑडियो, ट्रांसक्रिप्शन, एजेंट रिस्पॉन्स और अन्य ज़रूरी जानकारी देते हैं।

क्लाइंट से सर्वर को भेजे जा सकने वाले इवेंट्स की जानकारी के लिए क्लाइंट-से-सर्वर इवेंट्स दस्तावेज़ देखें।

परिचय

क्लाइंट इवेंट्स बातचीत की रीयल-टाइम प्रकृति बनाए रखने के लिए ज़रूरी हैं। ये इनिशियलाइज़ेशन मेटाडेटा से लेकर प्रोसेस्ड ऑडियो और एजेंट रिस्पॉन्स तक सब कुछ देते हैं।

ये इवेंट्स WebSocket कम्यूनिकेशन प्रोटोकॉल का हिस्सा हैं और हमारे SDKs इन्हें अपने-आप हैंडल करते हैं। उन्नत इम्प्लीमेंटेशन और डीबगिंग के लिए इन्हें समझना ज़रूरी है।

क्लाइंट इवेंट के प्रकार

  • बातचीत शुरू करते समय अपने-आप भेजा जाता है
  • बातचीत की सेटिंग्स और पैरामीटर इनिशियलाइज़ करता है
// Example initialization metadata
{
"type": "conversation_initiation_metadata",
"conversation_initiation_metadata_event": {
"conversation_id": "conv_123",
"agent_output_audio_format": "pcm_44100", // TTS output format
"user_input_audio_format": "pcm_16000" // ASR input format
}
}
  • एजेंट के concurrency limit पर होने के दौरान कॉल क्यू में रखे गए कॉलर को ही भेजा जाता है
  • waiting, conversation_initiation_metadata के बाद और किसी भी होल्ड ऑडियो से पहले, एक बार भेजा जाता है
  • प्रतीक्षा समाप्त होने पर admitted या timed_out एक बार भेजा जाता है। timed_out के बाद कोड 4300 के साथ WebSocket बंद हो जाता है
  • क्यू में मौजूद कॉलर को हमेशा भेजा जाता है। इसे एजेंट के client_events कॉन्फ़िगरेशन में सक्षम करने की ज़रूरत नहीं है

कॉलर के क्यू में होने पर, होल्ड ऑडियो नियमित audio इवेंट के रूप में आता है। होल्ड ऑडियो को एजेंट की आवाज़ मानने के बजाय, प्रतीक्षा स्थिति दिखाने के लिए इस इवेंट का इस्तेमाल करें।

// Example queue status event structure
{
"type": "queue_status",
"queue_status_event": {
"status": "waiting" // "waiting" | "admitted" | "timed_out"
}
}
// Example queue status handler
websocket.on('queue_status', (event) => {
const { status } = event.queue_status_event;
if (status === 'waiting') {
showWaitingState();
} else if (status === 'admitted') {
hideWaitingState();
} else if (status === 'timed_out') {
showAllAgentsBusyMessage();
}
});
  • तुरंत जवाब देने वाला हेल्थ चेक इवेंट
  • SDK इसे अपने-आप संभालता है
  • WebSocket कनेक्शन बनाए रखने के लिए इस्तेमाल होता है
// Example ping event structure
{
"ping_event": {
"event_id": 123456,
"ping_ms": 50 // Optional, estimated latency in milliseconds
},
"type": "ping"
}
// Example ping handler
websocket.on('ping', () => {
websocket.send('pong');
});
  • प्लेबैक के लिए base64 एन्कोडेड ऑडियो शामिल होता है
  • ट्रैकिंग और क्रम तय करने के लिए न्यूमेरिक इवेंट ID शामिल होती है
  • वॉइस आउटपुट स्ट्रीमिंग संभालता है
  • कैरेक्टर-लेवल टाइमिंग जानकारी वाला अलाइनमेंट डेटा शामिल होता है

WebRTC कनेक्शन पर audio इवेंट नहीं भेजा जाता, क्योंकि ऑडियो को LiveKit सीधे संभालता है।

// Example audio event structure
{
"audio_event": {
"audio_base_64": "base64_encoded_audio_string",
"event_id": 12345,
"alignment": { // Character-level timing data
"chars": ["H", "e", "l", "l", "o"],
"char_durations_ms": [50, 30, 40, 40, 60],
"char_start_times_ms": [0, 50, 80, 120, 160]
}
},
"type": "audio"
}
// Example audio event handler
websocket.on('audio', (event) => {
const { audio_event } = event;
const { audio_base_64, event_id, alignment } = audio_event;
audioPlayer.play(audio_base_64);
// Use alignment data for synchronized text display
const { chars, char_start_times_ms } = alignment;
chars.forEach((char, i) => {
setTimeout(() => highlightCharacter(char, i), char_start_times_ms[i]);
});
});
  • फ़ाइनल स्पीच-टू-टेक्स्ट नतीजे शामिल होते हैं
  • यूज़र के पूरे कथन दिखाता है
  • बातचीत के इतिहास के लिए इस्तेमाल होता है
// Example transcript event structure
{
"type": "user_transcript",
"user_transcription_event": {
"user_transcript": "Hello, how can you help me today?"
}
}
// Example transcript handler
websocket.on('user_transcript', (event) => {
const { user_transcription_event } = event;
const { user_transcript } = user_transcription_event;
updateConversationHistory(user_transcript);
});
  • एजेंट का पूरा संदेश शामिल होता है
  • संदेश पूरा होने पर एक बार भेजा जाता है, इसलिए वॉइस बातचीत में यह आमतौर पर संदेश का ऑडियो स्ट्रीम होना शुरू होने के बाद आता है।
  • डिस्प्ले और इतिहास के लिए इस्तेमाल होता है

एजेंट के टेक्स्ट को बनते ही दिखाने के लिए, इस इवेंट की प्रतीक्षा करने के बजाय नीचे दिए गए agent_chat_response_part इवेंट का उपयोग करें।

// Example response event structure
{
"type": "agent_response",
"agent_response_event": {
"agent_response": "Hello, how can I assist you today?"
}
}
// Example response handler
websocket.on('agent_response', (event) => {
const { agent_response_event } = event;
const { agent_response } = agent_response_event;
displayAgentMessage(agent_response);
});
  • रुकावट के बाद का छोटा किया गया जवाब शामिल होता है
  • दिखाए गए संदेश को अपडेट करता है
  • बातचीत की सटीकता बनाए रखता है
// Example response correction event structure
{
"type": "agent_response_correction",
"agent_response_correction_event": {
"original_agent_response": "Let me tell you about the complete history...",
"corrected_agent_response": "Let me tell you about..." // Truncated after interruption
}
}
// Example response correction handler
websocket.on('agent_response_correction', (event) => {
const { agent_response_correction_event } = event;
const { corrected_agent_response } = agent_response_correction_event;
displayAgentMessage(corrected_agent_response);
});
  • कस्टम LLM रिस्पॉन्स से कोई भी मेटाडेटा शामिल होता है
  • केवल कस्टम LLM इस्तेमाल करने पर भेजा जाता है
  • इसे एजेंट के client_events कॉन्फ़िगरेशन में साफ़ तौर पर सक्षम करना ज़रूरी है

यह इवेंट कस्टम LLM इंटीग्रेशन के लिए खास है। यह आपके कस्टम LLM सर्वर को रिस्पॉन्स के साथ अतिरिक्त मेटाडेटा भेजने देता है, जिसे क्लाइंट ऐप इस्तेमाल कर सकता है।

// Example agent response metadata event structure
{
"type": "agent_response_metadata",
"agent_response_metadata_event": {
"metadata": {
// Any key-value pairs returned by your custom LLM
"key": "value"
},
"event_id": 12345
}
}
// Example metadata handler
websocket.on('agent_response_metadata', (event) => {
const { agent_response_metadata_event } = event;
const { metadata, event_id } = agent_response_metadata_event;
// Use metadata for UI updates, logging, or analytics
console.log(`Response ${event_id} metadata:`, metadata);
updateResponseDetails(metadata);
});
  • एजेंट जिस फ़ंक्शन को क्लाइंट से चलवाना चाहता है, उसके फ़ंक्शन कॉल को दिखाता है
  • टूल का नाम, टूल कॉल ID और पैरामीटर शामिल होते हैं
  • क्लाइंट-साइड पर फ़ंक्शन चलाना और नतीजा वापस सर्वर को भेजना ज़रूरी है

SDK इस्तेमाल करने पर, नतीजा वापस सर्वर को भेजने के लिए कॉलबैक उपलब्ध होते हैं।

// Example tool call event structure
{
"type": "client_tool_call",
"client_tool_call": {
"tool_name": "search_database",
"tool_call_id": "call_123456",
"parameters": {
"query": "user information",
"filters": {
"date": "2024-01-01"
}
}
}
}
// Example tool call handler
websocket.on('client_tool_call', async (event) => {
const { client_tool_call } = event;
const { tool_name, tool_call_id, parameters } = client_tool_call;
try {
const result = await executeClientTool(tool_name, parameters);
// Send success response back to continue conversation
websocket.send({
type: "client_tool_result",
tool_call_id: tool_call_id,
result: result,
is_error: false
});
} catch (error) {
// Send error response if tool execution fails
websocket.send({
type: "client_tool_result",
tool_call_id: tool_call_id,
result: error.message,
is_error: true
});
}
});
  • एजेंट द्वारा टूल फ़ंक्शन चलाने का संकेत देता है
  • टूल मेटाडेटा और चलाने की स्थिति शामिल होती है
  • बातचीत के दौरान एजेंट द्वारा टूल इस्तेमाल करने की जानकारी देता है
// Example agent tool response event structure
{
"type": "agent_tool_response",
"agent_tool_response": {
"tool_name": "skip_turn",
"tool_call_id": "skip_turn_c82ca55355c840bab193effb9a7e8101",
"tool_type": "system",
"is_error": false
}
}
// Example agent tool response handler
websocket.on('agent_tool_response', (event) => {
const { agent_tool_response } = event;
const { tool_name, tool_call_id, tool_type, is_error } = agent_tool_response;
if (is_error) {
console.error(`Agent tool ${tool_name} failed:`, tool_call_id);
} else {
console.log(`Agent executed ${tool_type} tool: ${tool_name}`);
}
});
  • agent_tool_response को मिरर करता है और टूल के पूरे रिज़ल्ट पेलोड को full_tool_result में स्ट्रिंग के रूप में भी स्ट्रीम करता है।
  • डिस्प्ले या आगे की प्रोसेसिंग के लिए क्लाइंट में टूल आउटपुट दिखाता है।
  • इसे एजेंट के client_events कॉन्फ़िगरेशन में साफ़ तौर पर सक्षम करना ज़रूरी है।

यह इवेंट क्लाइंट को टूल का पूरा नतीजा दिखाता है और इसमें संवेदनशील डेटा हो सकता है। इसे केवल तभी सक्षम करें, जब क्लाइंट पेलोड को सुरक्षित रूप से संभालने के लिए भरोसेमंद हो। 64 KB से बड़े नतीजे अपने-आप छोटे कर दिए जाते हैं।

// Example agent tool response full payload event structure
{
"type": "agent_tool_response_full_payload",
"agent_tool_response_full_payload": {
"tool_name": "lookup_order",
"tool_call_id": "lookup_order_c82ca55355c840bab193effb9a7e8101",
"tool_type": "webhook",
"is_error": false,
"full_tool_result": "{\"order_id\": \"ORD-789\", \"status\": \"shipped\"}",
"truncated": false
}
}
// Example agent tool response full payload handler (using @elevenlabs/react)
import { ConversationProvider } from '@elevenlabs/react';
function App() {
return (
<ConversationProvider
onAgentToolResponse={(response) => {
if (!('full_tool_result' in response)) return;
const { tool_name, tool_call_id, is_error, full_tool_result, truncated } = response;
if (is_error) {
console.error(`Agent tool ${tool_name} failed:`, tool_call_id);
} else {
console.log(`Tool ${tool_name} returned:`, full_tool_result);
}
if (truncated) {
console.warn(`Tool ${tool_name} result was truncated (exceeded 64 KB).`);
}
}}
>
<Agent />
</ConversationProvider>
);
}
  • वॉइस एक्टिविटी डिटेक्शन स्कोर इवेंट
  • यूज़र के बोलने की संभावना बताता है
  • वैल्यू 0 से 1 तक होती हैं; ज़्यादा वैल्यू बोलने की ज़्यादा निश्चितता बताती हैं
// Example VAD score event
{
"type": "vad_score",
"vad_score_event": {
"vad_score": 0.95
}
}
  • एजेंट द्वारा MCP टूल फ़ंक्शन चलाने का संकेत देता है
  • टूल का नाम, टूल कॉल ID और पैरामीटर शामिल होते हैं
  • चार में से किसी एक स्थिति के साथ कॉल किया जाता है: loading, awaiting_approval, success और failure।
{
"type": "mcp_tool_call",
"mcp_tool_call": {
"service_id": "xJ8kP2nQ7sL9mW4vR6tY",
"tool_call_id": "call_123456",
"tool_name": "search_database",
"tool_description": "Search the database for user information",
"parameters": {
"query": "user information",
},
"timestamp": "2024-09-30T14:23:45.123456+00:00",
"state": "loading",
"approval_timeout_secs": 10
}
}
  • एजेंट के जवाब का टेक्स्ट बनते ही start, delta और stop मैसेज के रूप में स्ट्रीम करता है
  • टेक्स्ट-ओनली मोड में हमेशा भेजा जाता है; वॉइस बातचीत में इसे एजेंट के client_events कॉन्फ़िगरेशन में साफ़ तौर पर सक्षम करना ज़रूरी है
  • एजेंट या कोई सक्रिय प्रक्रिया blocking guardrail इस्तेमाल करते समय नहीं भेजा जाता, क्योंकि किसी भी हिस्से को जारी करने से पहले उसे पूरे जवाब का मूल्यांकन करना होता है
  • response_id स्ट्रीम किए जा रहे मैसेज की पहचान करता है और बाद में उसे कमिट करने वाले agent_response के response_id से मेल खाता है
// Example start event
{
"type": "agent_chat_response_part",
"text_response_part": {
"type": "start",
"text": "",
"event_id": 12345,
"response_id": "f47ac10b-58cc-4372-a567-0e02b2c3d479"
}
}
// Example delta event with text chunk
{
"type": "agent_chat_response_part",
"text_response_part": {
"type": "delta",
"text": "Hello, how can I",
"event_id": 12345,
"response_id": "f47ac10b-58cc-4372-a567-0e02b2c3d479"
}
}
// Example stop event
{
"type": "agent_chat_response_part",
"text_response_part": {
"type": "stop",
"text": "",
"event_id": 12345,
"response_id": "f47ac10b-58cc-4372-a567-0e02b2c3d479"
}
}
// Example handler
websocket.on('agent_chat_response_part', (event) => {
const { text_response_part } = event;
const { type: partType, text, response_id } = text_response_part;
if (partType === 'start') {
initializeResponseBuffer(response_id);
} else if (partType === 'delta') {
appendToResponseBuffer(response_id, text);
} else if (partType === 'stop') {
finalizeResponse(response_id);
}
});

agent_reasoning_response_part टेक्स्ट-ओनली बातचीत के दौरान मॉडल से मिला रीजनिंग स्ट्रीम करता है। client_events में इवेंट सक्षम करें और एजेंट के लिए रीजनिंग सारांश चालू करें। सर्वर start, delta और stop मैसेज भेजता है। यह वॉइस बातचीत के दौरान या एजेंट अथवा कोई सक्रिय प्रक्रिया blocking guardrails इस्तेमाल करते समय यह इवेंट नहीं भेजता।

यह इवेंट और इससे जुड़ा SDK कॉलबैक एक्सपेरिमेंटल हैं। इनका व्यवहार और स्वरूप किसी भी रिलीज़ में बदल सकता है।

इवेंट पेलोड
{
"type": "agent_reasoning_response_part",
"reasoning_response_part": {
"type": "delta",
"text": "The user asked to cancel, so I should verify the account before continuing.",
"event_id": 123456
}
}

स्टार्ट और स्टॉप इवेंट में text वैल्यू खाली होती है।

रीजनिंग इवेंट संभालें
import { Conversation } from '@elevenlabs/client';
const conversation = await Conversation.startSession({
agentId: 'agent_7101k5zvyjhmfg983brhmhkd98n6',
textOnly: true,
onAgentReasoningResponsePart: ({ type, text, event_id }) => {
if (type === 'start') {
initializeReasoningBuffer(event_id);
} else if (type === 'delta') {
appendToReasoningBuffer(text);
} else if (type === 'stop') {
finalizeReasoning();
}
},
});
  • एजेंट का जवाब पूरा होने पर ट्रिगर होता है, जिसमें लंबित टूल कॉल भी शामिल हैं। इस इवेंट के बाद एजेंट तभी आगे आउटपुट देगा, जब यूज़र नया इनपुट देगा या टर्न टाइमआउट नया टर्न शुरू करेगा।
  • इसे एजेंट के client_events कॉन्फ़िगरेशन में साफ़ तौर पर सक्षम करना ज़रूरी है
// Example agent response complete event structure
{
"type": "agent_response_complete",
"agent_response_complete_event": {
"event_id": 12345
}
}
// Example handler
websocket.on('agent_response_complete', (event) => {
const { agent_response_complete_event } = event;
const { event_id } = agent_response_complete_event;
console.log(`Agent response ${event_id} complete`);
});
  • गार्डरेल उल्लंघन के कारण बातचीत खत्म होने पर ट्रिगर होता है। गार्डरेल के ऐसे रिट्राई पर नहीं भेजा जाता जो सफल हो जाए।
  • इवेंट खुद ही संकेत है — इसमें type फ़ील्ड के अलावा कोई पेलोड नहीं होता।
  • इसे एजेंट के client_events कॉन्फ़िगरेशन में साफ़ तौर पर सक्षम करना ज़रूरी है।
// Example guardrail triggered event structure
{
"type": "guardrail_triggered"
}
// Example guardrail triggered handler (using @elevenlabs/client)
import { Conversation } from '@elevenlabs/client';
const conversation = await Conversation.startSession({
agentId: 'agent_7101k5zvyjhmfg983brhmhkd98n6',
onGuardrailTriggered: () => {
console.warn('Guardrail triggered — conversation will end.');
},
});

इवेंट फ्लो

किसी बातचीत के दौरान इवेंट्स का एक सामान्य क्रम यह होता है:

conversation_initiation_metadata ping pong audio user_transcript audio agent_response client_tool_call client_tool_result audio agent_response agent_response_correction Connection established Playing audio User responds Client tool runs Playing audio Interruption detected Client Server

जब कोई एजेंट अपनी concurrency limit पर होता है और कॉल क्यूइंग चालू होती है, तो सर्वर conversation_initiation_metadata और पहले audio इवेंट के बीच queue_status इवेंट भेजता है। कॉलर को अनुमति मिलने तक होल्ड ऑडियो audio इवेंट्स के रूप में दिया जाता है।

बेहतरीन तरीके

  1. एरर हैंडलिंग

    • हर इवेंट टाइप के लिए सही एरर हैंडलिंग लागू करें
    • डिबगिंग के लिए महत्वपूर्ण इवेंट्स लॉग करें
    • कनेक्शन में रुकावटों को सहजता से संभालें
  2. ऑडियो प्रबंधन

    • ऑडियो चंक्स को सही तरीके से बफ़र करें
    • रुकावट होने पर सही क्लीनअप लागू करें
    • ऑडियो रिसोर्स प्रबंधन संभालें
  3. कनेक्शन प्रबंधन

    • PING इवेंट्स का तुरंत जवाब दें
    • री-कनेक्शन लॉजिक लागू करें
    • कनेक्शन की स्थिति मॉनिटर करें

समस्या निवारण

  • सही WebSocket कनेक्शन सुनिश्चित करें
  • PING/PONG रिस्पॉन्स जांचें
  • API क्रेडेंशियल्स सत्यापित करें
  • ऑडियो चंक हैंडलिंग जांचें
  • ऑडियो फ़ॉर्मैट अनुकूलता सत्यापित करें
  • मेमोरी उपयोग मॉनिटर करें
  • डिबगिंग के लिए सभी इवेंट्स लॉग करें
  • एरर बाउंड्रीज़ लागू करें
  • इवेंट हैंडलर रजिस्ट्रेशन जांचें

विस्तृत इम्प्लीमेंटेशन उदाहरणों के लिए हमारा SDK दस्तावेज़ देखें।