Eleven v4를 소개합니다역대 가장 감성적인 모델, Eleven v4를 만나보세요. 10월 12일까지 Creator+에 크레딧 3배 제공

콘텐츠로 건너뛰기

동상과 대화하기: 멀티모달 ElevenAgents 기반 앱 만들기

작성자
Joe Reeve
게시일
최종 업데이트

듣기이 글 오디오로 듣기

동상을 촬영하세요. 묘사된 인물을 식별하세요. 그런 다음 각 인물이 시대에 맞는 고유한 목소리로 말하는 실시간 음성 대화를 나눠 보세요.

ElevenLabs의 보이스 디자인과 Agent API로 이를 만들 수 있습니다. 이 글에서는 컴퓨터 비전과 음성 생성을 결합해 공공 기념물을 인터랙티브한 경험으로 바꾸는 모바일 웹 앱의 아키텍처를 살펴봅니다. 여기의 모든 내용은 아래 API와 코드 샘플로 재현할 수 있습니다.

튜토리얼 건너뛰기 - 프롬프트 하나로 만들기

아래의 전체 앱은 단일 프롬프트로 제작되었으며, 빈 NextJS 프로젝트에서 Claude Opus 4.5(high)를 사용해 Cursor에서 한 번에 성공적으로 구현되는 것을 테스트했습니다. 바로 나만의 앱을 만들고 싶다면 아래 내용을 에디터에 붙여 넣으세요:

We need to make an app that:
- is optimised for mobile
- allows the user to take a picture (of a statue, picture, monument, etc) that includes one or more people
- uses an OpenAI LLM api call to identify the statue/monument/picture, characters within it, the location, and name
- allows the user to check it's correct, and then do either a deep research or a standard search to get information about the characters and the statue's history, and it's current location
- then create an ElevenLabs agent (allowing multiple voices), that the user can then talk to as though they're talking to the characters in the statue. Each character should use voice designer api to create a matching voice.
The purpose is to be fun and educational.

https://elevenlabs.io/docs/eleven-api/guides/how-to/voices/voice-design
https://elevenlabs.io/docs/eleven-agents/quickstart 
https://elevenlabs.io/docs/api-reference/agents/create


문서 링크 대신 ElevenLabs Agent Skills를 사용할 수도 있습니다. 문서를 기반으로 만들어졌으며, 더 나은 결과를 얻을 수도 있습니다.

이 글의 나머지 부분에서는 해당 프롬프트로 무엇을 만들 수 있는지 자세히 설명합니다.

작동 방식

파이프라인은 5단계로 구성됩니다:

  1. 이미지 캡처
  2. 작품과 등장인물 식별(OpenAI)
  3. 역사 조사(OpenAI)
  4. 각 인물의 고유한 음성 생성(ElevenAPI)
  5. WebRTC를 통한 실시간 음성 대화 시작(ElevenAgents)

비전으로 동상 식별하기

사용자가 동상을 촬영하면 이미지가 OpenAI의 비전 지원 모델로 전송됩니다. 구조화된 시스템 프롬프트가 작품명, 위치, 작가, 날짜와 더불어 각 인물에 대한 상세한 음성 설명을 추출합니다. 시스템 프롬프트에는 예상되는 JSON 출력 형식이 포함됩니다:

{
  "statueName": "string - name of the statue, monument, or artwork",
  "location": "string - where it is located (city, country)",
  "artist": "string - the creator of the artwork",
  "year": "string - year completed or unveiled",
  "description": "string - brief description of the artwork and its historical significance",
  "characters": [
    {
      "name": "string - character name",
      "description": "string - who this person was and their historical significance",
      "era": "string - time period they lived in",
      "voiceDescription": "string - detailed voice description for Voice Design API (include audio quality marker, age, gender, vocal qualities, accent, pacing, and personality)"
    }
  ]
}
const response = await openai.chat.completions.create({
  model: "gpt-5.2",
  response_format: { type: "json_object" },
  messages: [
    { role: "system", content: SYSTEM_PROMPT },
    {
      role: "user",
      content: [
        {
          type: "text",
          text: "Identify this statue/monument/artwork and all characters depicted.",
        },
        {
          type: "image_url",
          image_url: {
            url: `data:image/jpeg;base64,${base64Data}`,
            detail: "high",
          },
        },
      ],
    },
  ],
  max_completion_tokens: 2500,
});

런던 웨스트민스터 브리지에 있는 부디카 동상 사진의 경우, 응답은 다음과 같습니다:

{
  "statueName": "Boudica and Her Daughters",
  "location": "Westminster Bridge, London, UK",
  "artist": "Thomas Thornycroft",
  "year": "1902",
  "description": "Bronze statue depicting Queen Boudica riding a war chariot with her two daughters, commemorating her uprising against Roman occupation of Britain.",
  "characters": [
    {
      "name": "Boudica",
      "description": "Queen of the Iceni tribe who led an uprising against Roman occupation",
      "era": "Ancient Britain, 60-61 AD",
      "voiceDescription": "Perfect audio quality. A powerful woman in her 30s with a deep, resonant voice and a thick Celtic British accent. Her tone is commanding and fierce, with a booming quality that projects authority. She speaks at a measured, deliberate pace with passionate intensity."
    },
    // Other characters in the statue
  ]
}

효과적인 음성 설명 작성하기

음성 설명의 품질은 생성되는 음성의 품질을 직접적으로 결정합니다. 보이스 디자인 프롬프팅 가이드에서 자세히 다루지만, 포함해야 할 핵심 속성은 오디오 품질 표시("Perfect audio quality."), 연령과 성별, 톤/음색(깊은, 울림 있는, 거친), 정확한 억양("British"가 아닌 "thick Celtic British accent"), 그리고 말하기 속도입니다. 더 상세한 프롬프트일수록 더 정확한 결과를 얻습니다. 예를 들어 "a tired New Yorker in her 60s with a dry sense of humor"는 언제나 "an older female voice"보다 더 좋은 결과를 냅니다.

가이드에서 유의할 점도 몇 가지 있습니다. 억양의 강도를 설명할 때는 "strong"보다 "thick"을 사용하고, "foreign"처럼 모호한 표현은 피하세요. 또한 가상의 인물이나 역사적 인물에는 실제 억양을 참고로 제안할 수 있습니다(예: "an ancient Celtic queen with a thick British accent, regal and commanding").

다음을 사용해 인물 음성 만들기: 보이스 디자인

보이스 디자인 API는 텍스트 설명을 바탕으로 새로운 합성 음성을 생성합니다. 음성 샘플이나 복제는 필요하지 않습니다. 따라서 원본 오디오가 존재하지 않는 역사적 인물에 적합합니다.

과정은 두 단계로 이루어집니다.

미리 보기 생성

const { previews } = await elevenlabs.textToVoice.design({
  modelId: "eleven_multilingual_ttv_v2",
  voiceDescription: character.voiceDescription,
  text: sampleText,
});

text 파라미터는 중요합니다. 더 길고 인물에 어울리는 샘플 텍스트(50단어 이상)는 더 안정적인 결과를 만듭니다. 일반적인 인사말 대신 인물에 맞는 대사를 사용하세요. 보이스 디자인 프롬프팅 가이드에서 자세히 알아볼 수 있습니다.

음성 저장

미리 보기를 생성한 후 하나를 선택해 영구 음성을 만드세요:

const voice = await elevenlabs.textToVoice.create({
  voiceName: `StatueScanner - ${character.name}`,
  voiceDescription: character.voiceDescription,
  generatedVoiceId: previews[0].generatedVoiceId,
});

여러 인물이 있는 동상의 경우 음성 생성은 병렬로 실행됩니다. 인물 5명의 음성도 1명의 음성을 만드는 것과 거의 같은 시간에 생성됩니다:

const results = await Promise.all(
  characters.map((character) => createVoiceForCharacter(character))
);

다중 음성 ElevenLabs Agent 구축하기

음성을 만들었다면, 다음 단계는 ElevenLabs Agent를 구성하여 실시간으로 인물별 음성을 전환할 수 있게 하는 것입니다.

const agent = await elevenlabs.conversationalAi.agents.create({
  name: `Statue Scanner - ${statueName}`,
  tags: ["statue-scanner"],
  conversationConfig: {
    agent: {
      firstMessage,
      language: "en",
      prompt: {
        prompt: systemPrompt,
        temperature: 0.7,
      },
    },
    tts: {
      voiceId: primaryCharacter.voiceId,
      modelId: "eleven_v3",
      supportedVoices: otherCharacters.map((c) => ({
        voiceId: c.voiceId,
        label: c.name,
        description: c.voiceDescription,
      })),
    },
    turn: {
      turnTimeout: 10,
    },
    conversation: {
      maxDurationSeconds: 600,
    },
  },
});

다중 음성 전환

supportedVoices 배열은 Agent에 어떤 음성을 사용할 수 있는지 알려 줍니다. Agents 플랫폼은 음성 전환을 자동으로 처리합니다. LLM의 응답이 다른 인물이 말하고 있음을 나타내면 TTS 엔진이 해당 세그먼트를 올바른 음성으로 라우팅합니다.

그룹 대화를 위한 프롬프트 엔지니어링

여러 인물이 순차적인 Q&A가 아닌 실제 그룹처럼 느껴지게 하려면 의도적으로 프롬프트를 설계해야 합니다:

const multiCharacterRules = `
MULTI-CHARACTER DYNAMICS:
You are playing ALL ${characters.length} characters simultaneously.
Make this feel like a group conversation, not an interview.

- Characters should interrupt each other:
  "Actually, if I may -" / "Wait, I must say -"

- React to what others say:
  "Well said." / "I disagree with that..." / "Always so modest..."

- Have side conversations:
  "Do you remember when -" / "Tell them about the time you -"

The goal is for users to feel like they are witnessing a real exchange
between people who happen to include them.
`;

WebRTC를 통한 실시간 음성

마지막 요소는 클라이언트 연결입니다. ElevenLabs Agents는 지연 시간이 짧은 음성 대화를 위해 WebRTC를 지원합니다. 이는 WebSocket 기반 연결보다 눈에 띄게 빠르며, 자연스러운 대화 순서 전환에 중요합니다.

서버 측: 대화 토큰 가져오기

const { token } = await client.conversationalAi.conversations.getWebrtcToken({
    agentId,
});

클라이언트 측: 세션 시작하기

import { useConversation } from "@elevenlabs/react";

const conversation = useConversation({
  onConnect: () => setIsSessionActive(true),
  onDisconnect: () => setIsSessionActive(false),
  onMessage: (message) => {
    if (message.source === "ai") {
      setMessages((prev) => [...prev, { role: "agent", text: message.message }]);
    }
  },
});

await conversation.startSession({
  agentId,
  conversationToken: token,
  connectionType: "webrtc",
});

useConversation 훅은 오디오 캡처, 스트리밍, 음성 활동 감지, 재생을 처리합니다.

웹 검색으로 리서치 심화하기

대화를 시작하기 전에 더 많은 역사적 맥락을 원하는 사용자를 위해 OpenAI의 웹 검색 도구를 활용한 고급 리서치 모드를 추가할 수 있습니다:

const response = await openai.responses.create({
  model: "gpt-5.2",
  instructions: RESEARCH_SYSTEM_PROMPT,
  tools: [{ type: "web_search_preview" }],
  input: `Research ${identification.statueName}. Search for current information
including location, visiting hours, and recent news about the artwork.`,
});

배운 점

이 프로젝트는 텍스트, 리서치, 비전, 오디오 등 AI의 다양한 모달리티를 결합하면 디지털 세계와 현실 세계를 아우르는 경험을 만들 수 있음을 보여 줍니다. 교육, 업무, 재미를 위해 더 많은 사람이 탐색해 볼 수 있는 다중 모달 Agent의 미개척 잠재력이 많습니다.

지금 시작하세요

이 프로젝트에 사용된 API인 보이스 디자인, ElevenAgents, 그리고 OpenAI는 모두 현재 사용할 수 있습니다.

작성자

Joe is on the Growth team at ElevenLabs, focused on helping developers get the most out of the company's frontier audio models. Previously, he was an Engineering Manager at Amplitude and served as CTO of the Coronavirus Tech Handbook, building a real-time collaborative editing platform used by thousands of contributors.

유사한 기사

최고 품질의 AI 오디오로 창작하세요