会話をシミュレーション

シミュレーション会話でElevenLabsエージェントをテスト・評価する方法

このガイドと、ここで使用するエンドポイントは非推奨です。代わりに、 シミュレーションのエージェントテストタイプを 使用してください。

概要

ElevenLabs Agents APIでは、AIエージェントとのテキストベースの会話をシミュレーションして評価できます。このガイドでは、会話シミュレーションエンドポイント(バッチおよびストリーミング)を使用して、エンドツーエンドのシミュレーションテストワークフローを実装する方法を紹介します。エージェントのパフォーマンスを詳細にテスト・改善し、インタラクションの目標を満たしていることを確認できます。

前提条件

シミュレーションテストワークフローの実装

1

初期評価パラメータを特定する

エージェントの会話履歴を確認し、期待どおりに機能しなかったケースを見つけます。それらの会話をもとに、エージェントとやり取りするシミュレーションユーザー向けのさまざまなプロンプトを作成します。さらに、特定のシミュレーションユーザーで確認したい結果に応じて、エージェント設定にまだ指定されていない追加の評価基準を定義します。

2

SDKで会話をシミュレーションする

ElevenLabs SDKを使用して、シミュレーションエンドポイントへのリクエストを作成します。

from dotenv import load_dotenv
from elevenlabs import (
ElevenLabs,
ConversationSimulationSpecification,
AgentConfig,
PromptAgent,
PromptEvaluationCriteria
)
load_dotenv()
api_key = os.getenv("ELEVENLABS_API_KEY")
elevenlabs = ElevenLabs(api_key=api_key)
response = elevenlabs.conversational_ai.agents.simulate_conversation(
agent_id="YOUR_AGENT_ID",
simulation_specification=ConversationSimulationSpecification(
simulated_user_config=AgentConfig(
prompt=PromptAgent(
prompt="Your goal is to be a really difficult user.",
llm="gpt-4o",
temperature=0.5
)
)
),
extra_evaluation_criteria=[
PromptEvaluationCriteria(
id="politeness_check",
name="Politeness Check",
conversation_goal_prompt="The agent was polite.",
use_knowledge_base=False
)
]
)
print(response)

これは基本的な例です。入力パラメータの一覧については、 会話をシミュレーションおよび 会話シミュレーションをストリーミングエンドポイントのAPI リファレンスを参照してください。

3

レスポンスを分析する

SDKは、会話の完全なトランスクリプトと詳細な分析を含む包括的なJSONオブジェクトを提供します。

シミュレーション会話:シミュレーションユーザーとエージェントの各ターンを、メッセージとツール使用状況の詳細とともに記録します。

Example conversation history
[
...
{
"role": "user",
"message": "Maybe a little. I'll think about it, but I'm still not convinced it's the right move.",
"tool_calls": [],
"tool_results": [],
"feedback": null,
"llm_override": null,
"time_in_call_secs": 0,
"conversation_turn_metrics": null,
"rag_retrieval_info": null,
"llm_usage": null
},
{
"role": "agent",
"message": "I understand. If you want to explore more at your own pace, I can direct you to our documentation, which has guides and API references. Would you like me to send you a link?",
"tool_calls": [],
"tool_results": [],
"feedback": null,
"llm_override": null,
"time_in_call_secs": 0,
"conversation_turn_metrics": null,
"rag_retrieval_info": null,
"llm_usage": null
},
{
"role": "user",
"message": "I guess it wouldn't hurt to take a look. Go ahead and send it over.",
"tool_calls": [],
"tool_results": [],
"feedback": null,
"llm_override": null,
"time_in_call_secs": 0,
"conversation_turn_metrics": null,
"rag_retrieval_info": null,
"llm_usage": null
},
{
"role": "agent",
"message": null,
"tool_calls": [
{
"type": "client",
"request_id": "redirectToDocs_421d21e4b4354ed9ac827d7600a2d59c",
"tool_name": "redirectToDocs",
"params_as_json": "{\"path\": \"/docs/api-reference/introduction\"}",
"tool_has_been_called": false,
"tool_details": null
}
],
"tool_results": [],
"feedback": null,
"llm_override": null,
"time_in_call_secs": 0,
"conversation_turn_metrics": null,
"rag_retrieval_info": null,
"llm_usage": null
},
{
"role": "agent",
"message": null,
"tool_calls": [],
"tool_results": [
{
"type": "client",
"request_id": "redirectToDocs_421d21e4b4354ed9ac827d7600a2d59c",
"tool_name": "redirectToDocs",
"result_value": "Tool Called.",
"is_error": false,
"tool_has_been_called": true,
"tool_latency_secs": 0
}
],
"feedback": null,
"llm_override": null,
"time_in_call_secs": 0,
"conversation_turn_metrics": null,
"rag_retrieval_info": null,
"llm_usage": null
},
{
"role": "agent",
"message": "Okay, I've sent you a link to the introduction to our API reference. It provides a good starting point for understanding our different tools and how they can be integrated. Let me know if you have any questions as you explore it.\n",
"tool_calls": [],
"tool_results": [],
"feedback": null,
"llm_override": null,
"time_in_call_secs": 0,
"conversation_turn_metrics": null,
"rag_retrieval_info": null,
"llm_usage": null
}
...
]

分析:評価基準の結果、データ収集メトリクス、会話トランスクリプトの要約に関するインサイトを提供します。

Example analysis
{
"analysis": {
"evaluation_criteria_results": {
"politeness_check": {
"criteria_id": "politeness_check",
"result": "success",
"rationale": "The agent remained polite and helpful despite the user's challenging attitude."
},
"understood_root_cause": {
"criteria_id": "understood_root_cause",
"result": "success",
"rationale": "The agent acknowledged the user's hesitation and provided relevant information."
},
"positive_interaction": {
"criteria_id": "positive_interaction",
"result": "success",
"rationale": "The user eventually asked for the documentation link, indicating engagement."
}
},
"data_collection_results": {
"issue_type": {
"data_collection_id": "issue_type",
"value": "support_issue",
"rationale": "The user asked for help with integrating ElevenLabs tools."
},
"user_intent": {
"data_collection_id": "user_intent",
"value": "The user is interested in integrating ElevenLabs tools into a project."
}
},
"call_successful": "success",
"transcript_summary": "The user expressed skepticism, but the agent provided useful information and a link to the API documentation."
}
}
4

評価基準を改善する

シミュレーション会話を十分に確認し、評価基準の有効性を評価します。評価基準がエージェントの パフォーマンスを十分に評価できていないギャップや領域を特定します。望ましい結果に沿い、 エージェントの能力を正確に測定できるよう、評価基準を適宜見直して調整してください。

5

エージェントを改善する

評価基準の正確性に確信が持てたら、シミュレーション会話から得た学びを活用して、 エージェントの能力を強化します。エージェントの応答が目標やユーザーの期待に沿うよう、 システムプロンプトを改善することを検討してください。また、エージェントのトーン調整、 特定のクエリへの対応能力の向上、応答を充実させるための追加データソースの統合など、 最適化できる機能や設定も確認します。こうした学びを体系的に適用することで、優れたユーザー 体験を提供する、より堅牢で効果的な会話型エージェントを作成できます。

6

継続的に改善する

最初のテストと改善サイクルが完了したら、幅広い想定シナリオをカバーする包括的なテスト スイートを構築するのが効果的です。このスイートでは、さまざまなシミュレーションユーザーの プロンプトと開始条件を使って、複数の会話シミュレーションを検証できます。継続的に反復し、 アプローチを改善することで、変化するユーザーニーズに対してエージェントの有効性と応答性を 維持できます。

プロのヒント

詳細なプロンプトと基準

詳細で具体的なシミュレーションユーザープロンプトと評価基準を作成すると、シミュレーションテストの効果を高められます。より多くのコンテキストと具体性を提供するほど、エージェントは複雑なやり取りをより適切に理解し、応答できます。

モックツール設定

モックツール設定を活用して、エージェントの意思決定プロセスをテストします。これにより、エージェントがツール呼び出しを行うかどうかをどのように判断し、異なるツール呼び出し結果にどう反応するかを確認できます。詳細は、APIリファレンスのtool_mock_config入力パラメータを参照してください。

部分的な会話履歴

部分的な会話履歴を使用して、特定の時点からエージェントがどのようにやり取りを処理するかを評価します。これは、ユーザーがすでに特定の方法で質問を設定している会話や、特定のツール呼び出しが成功または失敗している会話を、エージェントが管理する能力を評価する場合に特に役立ちます。詳細は、APIリファレンスのpartial_conversation_history入力パラメータを参照してください。