Speech to Text API
Turn speech to text — transcribe spoken audio, or the audio track of a video, into accurate, timestamped text — with the Speech to Text API. Whether it’s audio to text or video to text, generate captions and subtitles, build searchable archives, and feed transcription into content, analytics, and moderation pipelines, with word-level timing and optional speaker labels.
Explore Speech to Text API Models
Choose from 2 powerful AI models
Turn speech to text — audio and video to text
The Speech to Text API will convert audio to text and video to text — transcribing spoken audio and the audio track of videos into timestamped text — for captions and subtitles, searchable archives, and transcription pipelines.
It is built for teams that need text out of everything they record: caption every video for accessibility and SEO, make a media library searchable, power analytics and moderation on spoken content, and generate first-draft transcripts for editors — with word-level timing and optional speaker labels.
What will it support?
| Feature | Planned support |
|---|---|
| Languages | Multi-language with auto-detect (final list at launch) |
| Timestamps | Word-level and segment-level |
| Speaker labels | Optional diarization (Speaker 1 / Speaker 2 …) |
| Output formats | Plain text, JSON, SRT and VTT captions |
| Inputs | Audio files and video files (audio track), upload or URL |
| Confidence | Per-word confidence scores in the JSON output |
What the output will look like
Illustrative example (not a live transcription):

