Vai alla navigazione

Realtime

Realtime speech-to-text transcription service. This WebSocket API enables streaming audio input and receiving transcription results.

Event Flow

  • Audio chunks are sent as input_audio_chunk messages
  • Transcription results are streamed back as partial_transcript (interim) and committed_transcript (stable/final for that segment)
  • Supports manual commit or VAD-based automatic commit strategies

Authentication is done either by providing a valid API key in the xi-api-key header or by providing a valid token in the token query parameter. Tokens can be generated from the single use token endpoint. Use tokens if you want to transcribe audio from the client side.

Handshake

WSS
/v1/speech-to-text/realtime

Header

xi-api-keystringOpzionale

Parametri di query

model_idenumObbligatorioPredefinito a scribe_v2_realtime

The ID of the model to use for speech-to-text transcription.

Valori consentiti:
tokenstringOpzionale

Single use token for authentication. Only used when initiating a session from the client. If provided, xi-api-key is no longer required for authentication.

audio_formatanyOpzionale

The encoding format of the audio. Supported formats: pcm_8000, pcm_16000, pcm_22050, pcm_24000, pcm_44100, pcm_48000, ulaw_8000.

language_codestringOpzionale

An ISO-639-1 or ISO-639-3 language_code corresponding to the language of the audio file. Can sometimes improve transcription performance if known beforehand. Defaults to null, in this case the language is predicted automatically.

secondary_languageslist of stringsOpzionale

Additional ISO-639-1 or ISO-639-3 language codes that may be present in the audio. Providing them makes language identification more reliable by only focusing on a certain set of languages. Each code is validated the same way as language_code.

commit_strategyenumOpzionale

Commit strategy for speech segmentation. 'manual' requires explicit commits, 'vad' automatically segments speech using silence detection, 'turn_prediction' additionally commits as soon as the model predicts the speaker has finished their turn (only for models that predict turns, where it is the default).

Valori consentiti:
vad_thresholddoubleOpzionale
VAD sensitivity threshold for detecting speech activity. Lower values are more sensitive to speech.
vad_silence_threshold_secsdoubleOpzionale
Duration of silence in seconds required to trigger a commit when VAD commit strategy is enabled. Longer values result in fewer commits but longer segments.
min_speech_duration_msintegerOpzionale
Minimum duration of speech in milliseconds required to be considered valid speech by VAD.
min_silence_duration_msintegerOpzionale
Minimum duration of silence in milliseconds required to be considered a speech break by VAD.
include_timestampsbooleanOpzionalePredefinito a false

Enable word/character-level timestamps in a delayed committed_transcript_with_timestamps message. When enabled, you'll receive an additional message with timestamps after each commit. Default: false.

include_language_detectionbooleanOpzionalePredefinito a false

Enable language detection in a delayed committed_transcript_with_timestamps message. When enabled, you'll receive an additional message with detected language_code after each commit. Default: false.

keytermslist of stringsOpzionale

List of keyterms to bias the model towards. Maximum 50 keyterms. Adds a 20% premium to the base transcription cost.

no_verbatimbooleanOpzionalePredefinito a false
If true, removes filler words, false starts and disfluencies from the transcript.
entity_detectionstring or list of stringsOpzionale

Detect entities on committed transcripts. Can be 'all', a single entity type or category, or a list of types/categories ('pii', 'phi', 'pci', 'other', 'offensive_language'). When enabled, detected entities are delivered in a separate 'committed_transcript_entities' event with their text, type, and character positions.

transcript_editstringOpzionale

Natural-language instruction applied to each committed transcript (max 2000 characters). The edited text is delivered in a separate 'edited_transcript' event containing the committed text and the edited text. Cannot be combined with entity_detection. Adds a 30% premium to the base transcription cost, billed for at least 10 seconds of audio per committed transcript.

filter_background_audiobooleanOpzionalePredefinito a false

Enable background speech filtering to reduce false activations from nearby conversations and ambient noise. When enabled without an explicit vad_threshold, a lower default threshold is applied. Cannot be combined with include_timestamps.

keepalive_interval_msintegerOpzionale500-10000

Opt-in keepalive interval in milliseconds (500-10000). While the audio being streamed contains no speech, the server sends a partial_transcript about this often, so clients that enforce a read timeout on incoming frames do not drop the connection during long pauses. The keepalive has empty text (text: ""), or repeats the latest partial text if the current segment is not committed yet. Keepalives are only sent after the transcription models have processed silent audio, so audio must keep streaming. Cadence is rounded to the audio processing granularity (about one second). Disabled by default.

enable_loggingbooleanOpzionalePredefinito a true

When enable_logging is set to false zero retention mode will be used for the request. This will mean history features are unavailable for this request, including request stitching. Zero retention mode may only be used by enterprise customers.

Invia

inputAudioChunkobjectObbligatorio
Audio data chunk sent from client to server for transcription.

Ricevi

sessionStartedobjectObbligatorio
Sent when the transcription session is successfully started.
OR
partialTranscriptobjectObbligatorio
Interim transcription result that may change.
OR
committedTranscriptobjectObbligatorio
Committed transcription result that will not change.
OR
committedTranscriptWithTimestampsobjectObbligatorio

Committed transcription result with word-level timestamps.

OR
committedTranscriptEntitiesobjectObbligatorio

Detected entities for a committed transcript. Only sent when the entity_detection query parameter is set.

OR
Edited TranscriptobjectObbligatorio

Edited version of a committed transcript, delivered as a separate event when the transcript_edit query parameter is set. text is the committed text the instruction was applied to (use it to correlate with the committed_transcript event, since edits may arrive out of commit order). If no edits were made, edited_text is identical to text.

OR
Scribe WarningobjectObbligatorio

Non-fatal notice. Sent after session_started when enable_logging=false was requested but zero retention mode was not applied; the session continues and is still logged.

OR
scribeErrorobjectObbligatorio
Error event during transcription.
OR
scribeAuthErrorobjectObbligatorio
Authentication error during transcription session.
OR
scribeQuotaExceededErrorobjectObbligatorio
Quota exceeded error during transcription session.
OR
scribeThrottledErrorobjectObbligatorio
Throttled error during transcription session.
OR
scribeUnacceptedTermsErrorobjectObbligatorio
Unaccepted terms error during transcription session.
OR
scribeRateLimitedErrorobjectObbligatorio
Rate limited error during transcription session.
OR
scribeQueueOverflowErrorobjectObbligatorio
Queue overflow error during transcription session.
OR
scribeResourceExhaustedErrorobjectObbligatorio
Resource exhausted error during transcription session.
OR
scribeSessionTimeLimitExceededErrorobjectObbligatorio
Session time limit exceeded error during transcription session.
OR
scribeInputErrorobjectObbligatorio
Input error during transcription session.
OR
Scribe Invalid Request ErrorobjectObbligatorio

The connection parameters were rejected; the session is closed afterwards.

OR
scribeChunkSizeExceededErrorobjectObbligatorio
Chunk size exceeded error during transcription session.
OR
scribeInsufficientAudioActivityErrorobjectObbligatorio
Insufficient audio activity error during transcription session.
OR
scribeTranscriberErrorobjectObbligatorio
Transcriber error during transcription session.