Realtime Speech to Text API
Transcribe speech live with Scribe v2 Realtime
Scribe v2 Realtime is the most accurate real-time STT with 150ms latency across 90+ languages. Available via API.
- Lovable
- Veed model
- Synthesia
- Stripe
- Perplexity
- Twilio
Built for speed and accuracy
Ultra-fast, ultra-accurate, and built for live speech. Scribe v2 Realtime delivers instant transcription for realtime use-cases.
Highest-accuracy realtime transcription
Scribe v2 Realtime achieves industry-leading transcription accuracy with ~150ms latency, even in challenging audio conditions or across diverse accents.
Designed for every scenario
Transcription that works in noisy environments, with background music, strong accents, and low-quality audio.
Speech recognition engineered for real-time performance
Built on the foundation of Scribe v1, Scribe v2 Realtime delivers ~150 ms latency with breakthrough accuracy across accents, tones, and environments.

Purpose-built for Agents and voice apps
Scribe v2 Realtime is purpose-built for developers creating conversational agents, meeting assistants, and voice applications where speed and accuracy are critical.
Predictive transcription for low latency
Scribe v2 Realtime uses predictive transcription to anticipate the most probable next words and punctuation – enabling real-time accuracy.
Voice Activity Detection
Detects when speech starts and stops, segmenting audio precisely for smooth, efficient real-time transcription.
Manual Commit Control
Gives developers control over when to finalize transcripts – ideal for custom streaming and fine-tuned accuracy.
Multiple Audio Formats
Supports PCM (8–48 kHz) and μ-law encoding for compatibility across telephony, browser, and studio setups.
Models optimized for every use-case
Scribe v2 for bulk use-cases, and Scribe v2 Realtime for low-latency use-cases

Scribe v2
Highest accuracy, designed for batch workloads.
- >95% Accuracy
- 90+ Languages
- Non-Speech Event Detection
- Entity Detection
- Keyterm Prompting

Scribe v2 Realtime
Lowest latency, for realtime workloads.
- Under 150ms Latency
- 90+ Languages
- Transcription Streaming
- Voice Activity Detection
- Automatic Language Recognition
Transcribe speech in 90+ languages and a wide range of accents
Delivering exceptional accuracy across accents, dialects, and recording conditions.
Change the languageCode to preview languages
import { useScribe } from "@elevenlabs/react";
const scribe = useScribe({
modelId: "scribe_v2_realtime",
languageCode: , // Set language
onSessionStarted: () =>
console.log("Session started"),
onPartialTranscript: (data) =>
console.log("Partial:", data.text)
});Powering the world’s leading companies and brands
Flexible pricing based on your needs
Experience best-in-class accuracy and responsiveness with pricing designed to scale from startups to enterprise teams.
$0.28 per hour & lower
on annual Business plans



