Pipecat 통합

Speech Engine의 LLM 브레인으로 Pipecat 파이프라인을 사용하세요.

이 가이드에서는 Speech Engine 브레인 서버 내 LLM 파이프라인으로 Pipecat을 사용하는 방법을 설명합니다. Speech Engine은 음성-텍스트 변환, 턴 관리, 텍스트 음성 변환으로 구성된 음성 루프를 처리하며, Pipecat은 프로세서(LLM 호출, RAG, 함수 호출, 가드레일, 콘텐츠 필터)로 구성 가능한 파이프라인을 통해 텍스트 생성을 처리합니다.

Pipecat은 서버 측 Python 프레임워크이므로 이 가이드는 Python만 다룹니다. 파이프라인 프로세서에 대응하는 Node 버전은 없습니다. pipecat-client-js 패키지가 있지만, 이는 Pipecat 서버와 통신하는 브라우저 클라이언트일 뿐 TypeScript에서 파이프라인을 빌드하는 방법은 아닙니다.

아키텍처

Speech Engine SDK는 외부 레이어로 실행되며, 사용자가 발화를 마칠 때마다 on_transcript 콜백이 실행됩니다. 콜백 내부에서 Pipecat 파이프라인을 빌드하고, 대화 기록을 LLMContextFrame으로 입력한 뒤 파이프라인의 텍스트 출력을 Speech Engine으로 스트리밍합니다. ElevenLabs는 텍스트를 음성으로 변환해 사용자에게 재생합니다.

User speaks (audio) on_transcript(history) LLMContextFrame(history) LLMTextFrame chunks send_response(async iterator) Agent speaks (audio) Browser ElevenLabs Brain Server (engine.serve) Pipecat Pipeline

Pipecat 파이프라인은 한 턴 동안만 실행됩니다. 새 트랜스크립트가 도착하면 다음 파이프라인이 실행되기 전에 이전 파이프라인이 취소됩니다. 이것이 Speech Engine의 인터럽트 처리가 파이프라인으로 전달되는 방식입니다.

이 패턴을 사용해야 하는 경우

브레인에 단일 LLM 호출 이상의 기능이 필요하다면 Pipecat이 유용합니다.

  • 검색 증강 생성, 함수 호출 또는 가드레일을 위한 조합 가능한 프로세서
  • 모든 단계에서 트래픽을 검사, 변환 또는 차단할 수 있는 프레임 기반 미들웨어
  • 여러 에이전트에서 공유할 수 있는 재사용 가능한 파이프라인 조각

브레인이 “트랜스크립트 입력, LLM 호출 출력” 구조라면 Speech Engine 퀵스타트가 더 간단합니다. 파이프라인 자체가 중요한 부분일 때 Pipecat을 사용하세요.

사전 요구 사항

  • Speech Engine. Speech Engine 퀵스타트를 따라 생성하세요.
  • Python 3.10+ (pipecat-ai에서 필요).
  • 브레인 서버용 공개 HTTPS 터널(예: ngrok).

종속성 설치

pip install "pipecat-ai[openai]" "elevenlabs" "python-dotenv"

pipecat-ai[openai]는 OpenAI LLM 서비스를 가져옵니다. 원한다면 다른 제공업체를 위해 추가 항목을 (anthropic, google 등)으로 바꾸세요.

Pipecat 브레인 빌드

브레인은 두 부분으로 구성됩니다. 스트리밍된 텍스트를 asyncio.Queue로 전달하는 TextSink 프로세서와, 한 턴 파이프라인을 빌드하고 청크를 비동기 이터레이터로 반환하는 run_pipecat_brain 코루틴입니다.

brain.py
import asyncio
import os
from typing import AsyncIterator
from dotenv import load_dotenv
from pipecat.frames.frames import (
Frame,
LLMContextFrame,
LLMFullResponseEndFrame,
LLMTextFrame,
)
from pipecat.pipeline.pipeline import Pipeline
from pipecat.pipeline.runner import PipelineRunner
from pipecat.pipeline.task import PipelineTask
from pipecat.processors.aggregators.llm_context import LLMContext
from pipecat.processors.frame_processor import FrameDirection, FrameProcessor
from pipecat.services.openai.llm import OpenAILLMService
load_dotenv()
SYSTEM_PROMPT = (
"You are a helpful voice assistant. Keep responses concise and conversational."
)
class TextSink(FrameProcessor):
"""Drain LLMTextFrame text into an asyncio.Queue."""
def __init__(self, queue: asyncio.Queue):
super().__init__()
self._queue = queue
async def process_frame(self, frame: Frame, direction: FrameDirection):
await super().process_frame(frame, direction)
if isinstance(frame, LLMTextFrame):
await self._queue.put(frame.text)
elif isinstance(frame, LLMFullResponseEndFrame):
await self._queue.put(None) # sentinel
await self.push_frame(frame, direction)
def build_messages(transcript: list[dict]) -> list[dict]:
messages = [{"role": "system", "content": SYSTEM_PROMPT}]
for turn in transcript:
role = "assistant" if turn["role"] == "agent" else turn["role"]
messages.append({"role": role, "content": turn["content"]})
return messages
async def run_pipecat_brain(transcript: list[dict]) -> AsyncIterator[str]:
"""Yield response text chunks from a one-turn Pipecat pipeline."""
llm = OpenAILLMService(
api_key=os.environ["OPENAI_API_KEY"],
model="gpt-4o-mini",
)
queue: asyncio.Queue[str | None] = asyncio.Queue()
sink = TextSink(queue)
task = PipelineTask(Pipeline([llm, sink]))
runner = PipelineRunner(handle_sigint=False)
async def drive():
context = LLMContext(build_messages(transcript))
await task.queue_frame(LLMContextFrame(context))
await task.stop_when_done()
run_task = asyncio.create_task(runner.run(task))
drive_task = asyncio.create_task(drive())
try:
while True:
chunk = await queue.get()
if chunk is None:
break
yield chunk
finally:
await task.cancel()
await asyncio.gather(run_task, drive_task, return_exceptions=True)

파이프라인에는 LLM 서비스와 싱크만 포함되며, STT 또는 TTS 프로세서는 포함되지 않습니다. 이는 Speech Engine이 처리하기 때문입니다. 입력은 LLMContextFrame이며 출력은 LLMTextFrame 청크입니다.

run_pipecat_brain은 비동기 제너레이터입니다. 반환되는 각 청크는 Speech Engine으로 직접 전달되므로, 전체 응답이 준비되기 전에 에이전트가 말하기를 시작합니다.

Speech Engine 서버에 연결

Speech Engine SDK의 send_response는 문자열 또는 문자열의 비동기 이터러블을 받으므로 run_pipecat_brain(transcript)를 직접 전달할 수 있습니다. Speech Engine이 제공하는 ConversationMessage 객체는 브레인에 전달하기 전에 일반 딕셔너리로 변환하세요.

server.py
import asyncio
import os
from dotenv import load_dotenv
from elevenlabs import AsyncElevenLabs
from brain import run_pipecat_brain
load_dotenv()
elevenlabs = AsyncElevenLabs(api_key=os.environ["ELEVENLABS_API_KEY"])
SPEECH_ENGINE_ID = os.environ["SPEECH_ENGINE_ID"]
async def on_transcript(transcript, session):
history = [{"role": m.role, "content": m.content} for m in transcript]
await session.send_response(run_pipecat_brain(history))
async def main():
engine = await elevenlabs.speech_engine.get(SPEECH_ENGINE_ID)
await engine.serve(
port=3001,
path="/ws",
debug=True,
on_transcript=on_transcript,
)
if __name__ == "__main__":
asyncio.run(main())

새 트랜스크립트가 도착하면 Speech Engine SDK가 이전 턴의 태스크를 취소합니다. 그러면 비동기 제너레이터와 run_pipecat_brain의 try/finally 블록을 통한 기본 PipelineTask가 취소됩니다.

서버 실행

ngrok http 3001
python server.py

퀵스타트에 나온 동일한 토큰 엔드포인트 및 클라이언트 코드를 사용해 브라우저에서 Speech Engine에 연결하세요. Pipecat 파이프라인은 서버 측에서 실행되며, 브라우저에서는 일반 Speech Engine 대화로 보입니다.

파이프라인 확장

텍스트 전용 Pipecat 파이프라인에는 LLMTextFrame 또는 LLMContextFrame에서 작동하는 모든 프레임 프로세서를 포함할 수 있습니다. 일반적인 추가 항목은 다음과 같습니다.

  • 가드레일: LLMContextFrame을 검사하고 안전하지 않은 컨텍스트를 교체하거나 차단하는, LLM 앞에 배치된 FrameProcessor입니다.
  • 함수 호출: OpenAILLMService에 도구를 등록하면 Pipecat이 도구 호출 프레임을 기본으로 처리합니다. 최종 어시스턴트 텍스트는 계속 LLMTextFrame으로 도착합니다.
  • 다단계 추론: 두 개의 OpenAILLMService 인스턴스를 연결하고, 그 사이에 두 번째 단계의 컨텍스트를 다시 작성하는 커스텀 프로세서를 둡니다.
  • 출력 필터링: 각 LLMTextFrame을 검사하고 허용되지 않는 콘텐츠를 TextSink에 도달하기 전에 제거하거나 다시 작성하는, LLM 뒤에 배치된 FrameProcessor입니다.

파이프라인 형태는 Pipeline([processor_a, llm, processor_b, sink])로 동일하게 유지되며, run_pipecat_brain도 변경되지 않습니다.

프로덕션 고려 사항

  • 취소 안전성: 파이프라인이 완전히 시작되기 전에 호출하면 PipelineTask.cancel()이 교착 상태에 빠질 수 있습니다(pipecat-ai/pipecat#4276). 위의 try/finally 패턴은 최소 하나의 프레임이 큐에 추가된 후에만 cancel()이 실행되므로 안전합니다.
  • 프롬프트 인젝션: 음성-텍스트 변환 출력은 사용자 입력입니다. 특히 후속 프로세서가 도구 호출이나 데이터베이스 쿼리에 텍스트를 사용하는 경우, LLM에 전달하기 전에 트랜스크립트를 검증하거나 정규화하세요.
  • 브레인 서버 인증: Speech Engine에 공유 시크릿을 설정하고 브레인 서버에서 이를 확인하여 /ws 엔드포인트에 대한 무단 연결을 방지하세요.
    await elevenlabs.speech_engine.update(
    speech_engine_id=SPEECH_ENGINE_ID,
    speech_engine={"request_headers": {"x-api-key": os.environ["SHARED_SECRET"]}},
    )
  • LLM 제공업체: pipecat-ai[openai]에는 OpenAILLMService가 포함됩니다. Anthropic의 경우 pipecat-ai[anthropic]을 설치하고 AnthropicLLMService를 사용하세요. 나머지 파이프라인은 변경되지 않습니다.

다음 단계