Hoppa till navigering

Realtime

Realtime speech-to-text transcription service. This WebSocket API enables streaming audio input and receiving transcription results.

Event Flow

  • Audio chunks are sent as input_audio_chunk messages
  • Transcription results are streamed back as partial_transcript (interim) and committed_transcript (stable/final for that segment)
  • Supports manual commit or VAD-based automatic commit strategies

Authentication is done either by providing a valid API key in the xi-api-key header or by providing a valid token in the token query parameter. Tokens can be generated from the single use token endpoint. Use tokens if you want to transcribe audio from the client side.

Handskakining

WSS
/v1/speech-to-text/realtime

Headers

xi-api-keystringValfri

Frågeparametrar

model_idenumObligatoriskStandard är scribe_v2_realtime

The ID of the model to use for speech-to-text transcription.

Tillåtna värden:
tokenstringValfri

Single use token for authentication. Only used when initiating a session from the client. If provided, xi-api-key is no longer required for authentication.

audio_formatanyValfri

The encoding format of the audio. Supported formats: pcm_8000, pcm_16000, pcm_22050, pcm_24000, pcm_44100, pcm_48000, ulaw_8000.

language_codestringValfri

An ISO-639-1 or ISO-639-3 language_code corresponding to the language of the audio file. Can sometimes improve transcription performance if known beforehand. Defaults to null, in this case the language is predicted automatically.

secondary_languageslist of stringsValfri

Additional ISO-639-1 or ISO-639-3 language codes that may be present in the audio. Providing them makes language identification more reliable by only focusing on a certain set of languages. Each code is validated the same way as language_code.

commit_strategyenumValfri

Commit strategy for speech segmentation. 'manual' requires explicit commits, 'vad' automatically segments speech using silence detection, 'turn_prediction' additionally commits as soon as the model predicts the speaker has finished their turn (only for models that predict turns, where it is the default).

Tillåtna värden:
vad_thresholddoubleValfri
VAD sensitivity threshold for detecting speech activity. Lower values are more sensitive to speech.
vad_silence_threshold_secsdoubleValfri
Duration of silence in seconds required to trigger a commit when VAD commit strategy is enabled. Longer values result in fewer commits but longer segments.
min_speech_duration_msintegerValfri
Minimum duration of speech in milliseconds required to be considered valid speech by VAD.
min_silence_duration_msintegerValfri
Minimum duration of silence in milliseconds required to be considered a speech break by VAD.
include_timestampsbooleanValfriStandard är false

Enable word/character-level timestamps in a delayed committed_transcript_with_timestamps message. When enabled, you'll receive an additional message with timestamps after each commit. Default: false.

include_language_detectionbooleanValfriStandard är false

Enable language detection in a delayed committed_transcript_with_timestamps message. When enabled, you'll receive an additional message with detected language_code after each commit. Default: false.

keytermslist of stringsValfri

List of keyterms to bias the model towards. Maximum 50 keyterms. Adds a 20% premium to the base transcription cost.

no_verbatimbooleanValfriStandard är false
If true, removes filler words, false starts and disfluencies from the transcript.
entity_detectionstring or list of stringsValfri

Detect entities on committed transcripts. Can be 'all', a single entity type or category, or a list of types/categories ('pii', 'phi', 'pci', 'other', 'offensive_language'). When enabled, detected entities are delivered in a separate 'committed_transcript_entities' event with their text, type, and character positions.

transcript_editstringValfri

Natural-language instruction applied to each committed transcript (max 2000 characters). The edited text is delivered in a separate 'edited_transcript' event containing the committed text and the edited text. Cannot be combined with entity_detection. Adds a 30% premium to the base transcription cost, billed for at least 10 seconds of audio per committed transcript.

filter_background_audiobooleanValfriStandard är false

Enable background speech filtering to reduce false activations from nearby conversations and ambient noise. When enabled without an explicit vad_threshold, a lower default threshold is applied. Cannot be combined with include_timestamps.

keepalive_interval_msintegerValfri500-10000

Opt-in keepalive interval in milliseconds (500-10000). While the audio being streamed contains no speech, the server sends a partial_transcript about this often, so clients that enforce a read timeout on incoming frames do not drop the connection during long pauses. The keepalive has empty text (text: ""), or repeats the latest partial text if the current segment is not committed yet. Keepalives are only sent after the transcription models have processed silent audio, so audio must keep streaming. Cadence is rounded to the audio processing granularity (about one second). Disabled by default.

enable_loggingbooleanValfriStandard är true

When enable_logging is set to false zero retention mode will be used for the request. This will mean history features are unavailable for this request, including request stitching. Zero retention mode may only be used by enterprise customers.

Skicka

inputAudioChunkobjectObligatorisk
Audio data chunk sent from client to server for transcription.

Ta emot

sessionStartedobjectObligatorisk
Sent when the transcription session is successfully started.
OR
partialTranscriptobjectObligatorisk
Interim transcription result that may change.
OR
committedTranscriptobjectObligatorisk
Committed transcription result that will not change.
OR
committedTranscriptWithTimestampsobjectObligatorisk

Committed transcription result with word-level timestamps.

OR
committedTranscriptEntitiesobjectObligatorisk

Detected entities for a committed transcript. Only sent when the entity_detection query parameter is set.

OR
Edited TranscriptobjectObligatorisk

Edited version of a committed transcript, delivered as a separate event when the transcript_edit query parameter is set. text is the committed text the instruction was applied to (use it to correlate with the committed_transcript event, since edits may arrive out of commit order). If no edits were made, edited_text is identical to text.

OR
Scribe WarningobjectObligatorisk

Non-fatal notice. Sent after session_started when enable_logging=false was requested but zero retention mode was not applied; the session continues and is still logged.

OR
scribeErrorobjectObligatorisk
Error event during transcription.
OR
scribeAuthErrorobjectObligatorisk
Authentication error during transcription session.
OR
scribeQuotaExceededErrorobjectObligatorisk
Quota exceeded error during transcription session.
OR
scribeThrottledErrorobjectObligatorisk
Throttled error during transcription session.
OR
scribeUnacceptedTermsErrorobjectObligatorisk
Unaccepted terms error during transcription session.
OR
scribeRateLimitedErrorobjectObligatorisk
Rate limited error during transcription session.
OR
scribeQueueOverflowErrorobjectObligatorisk
Queue overflow error during transcription session.
OR
scribeResourceExhaustedErrorobjectObligatorisk
Resource exhausted error during transcription session.
OR
scribeSessionTimeLimitExceededErrorobjectObligatorisk
Session time limit exceeded error during transcription session.
OR
scribeInputErrorobjectObligatorisk
Input error during transcription session.
OR
Scribe Invalid Request ErrorobjectObligatorisk

The connection parameters were rejected; the session is closed afterwards.

OR
scribeChunkSizeExceededErrorobjectObligatorisk
Chunk size exceeded error during transcription session.
OR
scribeInsufficientAudioActivityErrorobjectObligatorisk
Insufficient audio activity error during transcription session.
OR
scribeTranscriberErrorobjectObligatorisk
Transcriber error during transcription session.