Grok Voice Agent API, Grok Voice TTS API, Grok Voice STT API: Pricing, Documentation

by xAI

Grok Voice Agent API delivers a powerful, low-latency solution for integrating natural, expressive, and contextually intelligent speech capabilities directly into applications. Built to handle complex, real-time conversations, this interface allows developers to convert text into hyper-realistic human audio and process spoken inputs with remarkable accuracy. By leveraging advanced deep learning architectures, it captures subtle linguistic nuances, emotional tones, and varied cadences, making interactions feel fluid and genuinely conversational rather than robotic.

Get API Key
Grok Voice API

Models Version

WELCOME BONUS

Get $5 Free Credit on First Payment

No strings attached — add funds and get $5 bonus instantly

Claim Your $5 →

Grok Voice Agent API Documentation

A speech-to-speech voice agent by xAI. Send a recording of the user speaking and get Grok's spoken reply back as a WAV file. Optionally set the agent's persona with prompt and pick one of 28 built-in voices. Asynchronous: submit returns a request_id; poll the status endpoint until the request is COMPLETED, then download the reply audio from output.media_url.

POST https://gateway.pixazo.ai/grok-voice/v1/speech-to-speech

Authentication

All requests require an API key passed via header.

HeaderTypeRequiredDescription
Ocp-Apim-Subscription-KeystringYesYour API subscription key

Speech to Speech - Grok Voice

Request Code

POST https://gateway.pixazo.ai/grok-voice/v1/speech-to-speech
Content-Type: application/json
Ocp-Apim-Subscription-Key: YOUR_SUBSCRIPTION_KEY

{
  "audio_url": "https://pub-582b7213209642b9b995c96c95a30381.r2.dev/doc-assets/audio/transcribe-sample.wav",
  "prompt": "You are a friendly assistant. Answer briefly and concretely.",
  "voice": "eve"
}
import requests

url = "https://gateway.pixazo.ai/grok-voice/v1/speech-to-speech"
headers = {
    "Content-Type": "application/json",
    "Ocp-Apim-Subscription-Key": "YOUR_SUBSCRIPTION_KEY"
}
data = {
  "audio_url": "https://pub-582b7213209642b9b995c96c95a30381.r2.dev/doc-assets/audio/transcribe-sample.wav",
  "prompt": "You are a friendly assistant. Answer briefly and concretely.",
  "voice": "eve"
}

resp = requests.post(url, json=data, headers=headers)
print(resp.json())
const res = await fetch("https://gateway.pixazo.ai/grok-voice/v1/speech-to-speech", {
  method: "POST",
  headers: {
    "Content-Type": "application/json",
    "Ocp-Apim-Subscription-Key": "YOUR_SUBSCRIPTION_KEY"
  },
  body: JSON.stringify({
  "audio_url": "https://pub-582b7213209642b9b995c96c95a30381.r2.dev/doc-assets/audio/transcribe-sample.wav",
  "prompt": "You are a friendly assistant. Answer briefly and concretely.",
  "voice": "eve"
})
});
console.log(await res.json());
curl -X POST 'https://gateway.pixazo.ai/grok-voice/v1/speech-to-speech' \
  -H 'Content-Type: application/json' \
  -H 'Ocp-Apim-Subscription-Key: YOUR_SUBSCRIPTION_KEY' \
  --data-raw '{"audio_url": "https://pub-582b7213209642b9b995c96c95a30381.r2.dev/doc-assets/audio/transcribe-sample.wav", "prompt": "You are a friendly assistant. Answer briefly and concretely.", "voice": "eve"}'

Output

{
  "request_id": "grok-voice_019dxxxx-xxxx-xxxx-xxxx-xxxxxxxxxxxx",
  "status": "QUEUED",
  "polling_url": "https://gateway.pixazo.ai/v2/requests/status/grok-voice_019dxxxx-xxxx-xxxx-xxxx-xxxxxxxxxxxx"
}

Webhook (Optional)

Instead of polling, you can receive a Webhook callback when the request reaches a terminal state. Provide a Webhook URL via header on the submit request.

HeaderRequiredDescription
X-Webhook-URLTo enableHTTPS URL to receive the Webhook callback.
X-Webhook-ModeNoterminal (default, one callback on COMPLETED/ERROR) or sync (per-poll callbacks).

Example: enable Webhook

curl -X POST 'https://gateway.pixazo.ai/grok-voice/v1/speech-to-speech' \
  -H 'Content-Type: application/json' \
  -H 'Ocp-Apim-Subscription-Key: YOUR_SUBSCRIPTION_KEY' \
  -H 'X-Webhook-URL: https://your-server.com/webhook' \
  --data-raw '{"audio_url": "https://pub-582b7213209642b9b995c96c95a30381.r2.dev/doc-assets/audio/transcribe-sample.wav", "prompt": "You are a friendly assistant. Answer briefly and concretely.", "voice": "eve"}'

Callback Payload (success)

{
  "request_id": "grok-voice_019dxxxx-xxxx-xxxx-xxxx-xxxxxxxxxxxx",
  "status": "COMPLETED",
  "model_id": "grok-voice",
  "output": {
    "media_url": [
      "https://pub-582b7213209642b9b995c96c95a30381.r2.dev/v1/{request_id}/output.wav"
    ],
    "media_type": "audio/wav"
  },
  "created_at": "2026-09-21T09:14:16.102Z",
  "completed_at": "2026-09-21T09:14:31.870Z"
}

Failure callback shape

{
  "request_id": "grok-voice_019dxxxx-xxxx-xxxx-xxxx-xxxxxxxxxxxx",
  "status": "ERROR",
  "model_id": "grok-voice",
  "error": "Description of the failure"
}

Delivery semantics

  • terminal mode: one Webhook callback when the request is COMPLETED or ERROR.
  • sync mode: a Webhook callback on each status change.
  • Callbacks are idempotent on request_id — de-duplicate on it.
  • Respond 200 within a few seconds; the Webhook endpoint must be HTTPS.

Request Parameters

ParameterRequiredTypeDefaultAllowed values / rangeDescription
audio_urlYesstringhttps URL; up to 10 minutes and 50 MBA recording of the user speaking — the thing the agent listens to and answers. Fetched by the gateway, so the URL must be reachable without authentication. Most common audio formats are accepted; the file is converted to 16-bit 24 kHz mono PCM before the model hears it.
promptNostring— (Grok's default persona)up to 20,000 charactersSystem instructions describing the agent's persona and the conversation context — who it is, how it should answer, what it knows. This is not text to be read aloud; the reply is composed from the recording. Omit it to use Grok's default persona.
voiceNostring— (Grok's default voice)one of the 28 voices listed belowBuilt-in xAI voice for the spoken reply. Case-sensitive; see the Voices section.

Audio requirements

  • At most 10 minutes of audio and at most 50 MB per request. Longer recordings fail once the file is read, reported as status: "ERROR".
  • Most common formats are accepted (WAV, MP3, M4A, OGG, FLAC and similar). The audio is converted to 16-bit 24 kHz mono PCM before it reaches the model, so you do not need to convert it yourself.
  • audio_url must be an https URL and publicly reachable.
  • Cost scales with the length of the reply the agent generates, not the recording you send; the current rate is in the pricing panel on this page.

Fields this endpoint does not accept

The provider's agent tools (tools.web_search, tools.x_search, tools.mcp_servers) are not available on this endpoint: a body containing tools is rejected with a synchronous 400. Likewise any field that is not listed in the table above, such as duration or seconds, is rejected rather than ignored.

Voices

Pass one of these 28 names as voice, exactly as written (lowercase). Omit the field to use Grok's default voice.

carina, zagan, helix, orion, luna, iris, altair, zenith, perseus, helios, lux, kepler, rigel, cosmo, celeste, ursa, sirius, lumen, castor, naksh, atlas, aurora, liora, ara, eve, leo, rex, sal

Example Request

{
  "audio_url": "https://pub-582b7213209642b9b995c96c95a30381.r2.dev/doc-assets/audio/transcribe-sample.wav",
  "prompt": "You are a friendly assistant. Answer briefly and concretely.",
  "voice": "eve"
}

Example Response

{
  "request_id": "grok-voice_019dxxxx-xxxx-xxxx-xxxx-xxxxxxxxxxxx",
  "status": "QUEUED",
  "polling_url": "https://gateway.pixazo.ai/v2/requests/status/grok-voice_019dxxxx-xxxx-xxxx-xxxx-xxxxxxxxxxxx"
}

Request Headers

HeaderRequiredDescription
Content-TypeYesapplication/json
Ocp-Apim-Subscription-KeyYesYour API subscription key.
X-Webhook-URLNoEnable Webhook callbacks (see Webhook section).

Response Handling

Status CodeMeaning
202Accepted — request queued; returns request_id and polling_url.
400Bad request — audio_url missing or not a string, prompt over 20,000 characters, voice outside the 28 listed, or a field this endpoint does not accept (such as tools). Anything checkable from the request body alone.
401Unauthorized — missing or invalid subscription key.
402Insufficient balance.
429Too many requests.
500Internal server error.

Only what can be judged from the request body itself is rejected synchronously. Everything that needs the audio to be fetched and read is reported through the status endpoint instead, as status: "ERROR". Failed requests are not billed — the hold is released in full.

ConditionHow it surfaces
audio_url missing, or the wrong typesynchronous 400
voice not one of the 28 names, or prompt over 20,000 characterssynchronous 400
tools, or any other field not in the parameter tablesynchronous 400
audio_url not https, not reachable, or not publicstatus: "ERROR"
Audio longer than 10 minutes, or a file over 50 MBstatus: "ERROR"
Audio in a format the provider cannot decodestatus: "ERROR"

Retrieving Results

Poll the status endpoint with the request_id from the submit response until status is COMPLETED (or FAILED/ERROR), then download the reply audio from output.media_url.

curl 'https://gateway.pixazo.ai/v2/requests/status/grok-voice_019dxxxx-xxxx-xxxx-xxxx-xxxxxxxxxxxx' \
  -H 'Ocp-Apim-Subscription-Key: YOUR_SUBSCRIPTION_KEY'

Completed response

{
  "request_id": "grok-voice_019dxxxx-xxxx-xxxx-xxxx-xxxxxxxxxxxx",
  "status": "COMPLETED",
  "model_id": "grok-voice",
  "output": {
    "media_url": [
      "https://pub-582b7213209642b9b995c96c95a30381.r2.dev/v1/{request_id}/output.wav"
    ],
    "media_type": "audio/wav"
  },
  "created_at": "2026-09-21T09:14:16.102Z",
  "completed_at": "2026-09-21T09:14:31.870Z"
}

The reply is returned as audio only — a 24 kHz mono WAV file. A text transcript of the reply is not returned by this endpoint; if you need one, pass the WAV to a speech-to-text API.

Response Fields

FieldTypeDescription
request_idstringUnique request identifier.
statusstringQUEUED, PROCESSING, COMPLETED, FAILED, or ERROR.
model_idstringThe model that handled the request.
output.media_urlarrayURL of the reply audio (WAV, 24 kHz mono).
output.media_typestringaudio/wav.
created_atstringRequest creation timestamp.
completed_atstringCompletion timestamp.
errorstringError message when status is FAILED/ERROR.

Status Values & Flow

QUEUEDPROCESSINGCOMPLETED (success) or FAILED/ERROR (failure).

How billing is measured

Billing is measured on the length of the reply audio the model generates, per second, rounded up to the next whole second — not on the length of the recording you send, and not on how long the request takes. A hold is placed when the request is submitted, because the reply length is not known until the model has answered; the hold is reduced to the real cost once the reply audio is ready, and released in full if the request fails. To keep replies short, say so in prompt.

The current rate is shown in the pricing panel on this page.

Grok Voice Agent API Pricing

Your request will cost $0.0009 per second of generated speech.
billed on the length of the spoken reply, not your recording — a 30-second answer is about $0.026
equivalent to $0.052 per minute of generated speech
2. Grok Voice Text to Speech API

Grok Voice Text to Speech API Documentation

https://gateway.pixazo.ai/xai-text-to-speech/v1/text-to-speech

Authentication

All requests require an API key passed via header.

Header Type Required Description
Ocp-Apim-Subscription-Key string Yes Your API subscription key

Grok Voice Text to Speech generate request - Grok Voice Text to Speech

Request Code

POST https://gateway.pixazo.ai/xai-text-to-speech/v1/text-to-speech
Content-Type: application/json
Cache-Control: no-cache
Ocp-Apim-Subscription-Key: YOUR_SUBSCRIPTION_KEY

{
  "text": "Welcome to the future of artificial intelligence and creative expression",
  "voice": "eve",
  "language": "auto"
}
import requests

url = "https://gateway.pixazo.ai/xai-text-to-speech/v1/text-to-speech"
headers = {
    "Content-Type": "application/json",
    "Cache-Control": "no-cache",
    "Ocp-Apim-Subscription-Key": "YOUR_SUBSCRIPTION_KEY"
}
data = {
    "text": "Welcome to the future of artificial intelligence and creative expression",
    "voice": "eve",
    "language": "auto"
}

response = requests.post(url, json=data, headers=headers)
print(response.json())
const url = 'https://gateway.pixazo.ai/xai-text-to-speech/v1/text-to-speech';

const data = {
  text: 'Welcome to the future of artificial intelligence and creative expression',
  voice: 'eve',
  language: 'auto'
};

fetch(url, {
  method: 'POST',
  headers: {
    'Content-Type': 'application/json',
    'Cache-Control': 'no-cache',
    'Ocp-Apim-Subscription-Key': 'YOUR_SUBSCRIPTION_KEY'
  },
  body: JSON.stringify(data)
})
.then(response => response.json())
.then(data => console.log(data))
.catch(error => console.error('Error:', error));
curl -X POST "https://gateway.pixazo.ai/xai-text-to-speech/v1/text-to-speech" \
  -H "Content-Type: application/json" \
  -H "Cache-Control: no-cache" \
  -H "Ocp-Apim-Subscription-Key: YOUR_SUBSCRIPTION_KEY" \
  --data-raw '{
    "text": "Welcome to the future of artificial intelligence and creative expression",
    "voice": "eve",
    "language": "auto"
  }'

Output

{
  "request_id": "xai-text-to-speech_019dxxxx-xxxx-xxxx-xxxx-xxxxxxxxxxxx",
  "status": "QUEUED",
  "polling_url": "https://gateway.pixazo.ai/v2/requests/status/xai-text-to-speech_019dxxxx-xxxx-xxxx-xxxx-xxxxxxxxxxxx"
}

Webhook (Optional)

Add the X-Webhook-URL header to your submit request to receive a POST callback when the job completes — no polling required.

Using curl? These are HTTP request headers — pass each with -H, e.g. -H "X-Webhook-URL: https://your-server.com/webhook/callback". Do not paste them as bare lines, and end every line of a multi-line command with \.

Webhook Headers

HeaderRequiredDefaultDescription
X-Webhook-URLYes (to enable)HTTPS endpoint on your server that will receive the POST callback. Must respond 2xx within a few seconds (process async if needed).
X-Webhook-ModeNoterminalterminal — fires once at the final status (COMPLETED/FAILED/ERROR). sync — fires on every poll cycle plus the terminal event, and caps the queue’s polling delay at 15s for tighter progress updates.

Example: enable webhook

X-Webhook-URL: https://your-server.com/webhook/callback
X-Webhook-Mode: terminal

Callback Payload

Your endpoint receives a POST application/json with the same shape as the GET /v2/requests/status/{request_id} response. Example terminal callback (mode terminal):

{
  "request_id": "xai-text-to-speech_019dxxxx-xxxx-xxxx-xxxx-xxxxxxxxxxxx",
  "status": "COMPLETED",
  "model_id": "xai-text-to-speech",
  "error": null,
  "output": {
    "media_url": [
      "https://pub-582b7213209642b9b995c96c95a30381.r2.dev/v1/xai-text-to-speech_019dxxxx-xxxx-xxxx-xxxx-xxxxxxxxxxxx/output.wav"
    ],
    "media_type": "audio/wav"
  },
  "created_at": "2026-05-22T13:17:32.110Z",
  "updated_at": "2026-05-22 13:19:23",
  "completed_at": "2026-05-22 13:19:23"
}

Failure callback shape

{
  "request_id": "xai-text-to-speech_019dxxxx-xxxx-xxxx-xxxx-xxxxxxxxxxxx",
  "status": "ERROR",
  "model_id": "xai-text-to-speech",
  "error": "Description of the error",
  "output": null,
  "created_at": "...",
  "updated_at": "...",
  "completed_at": "..."
}

Delivery semantics

  • terminal mode (default) — exactly one POST when the request reaches a terminal status. No callback during PROCESSING.
  • sync modePOST on every status poll (with delay capped at ~15s) plus a final POST at terminal status. Use when you want progress updates.
  • Idempotency — use request_id as your idempotency key. Network retries can deliver the same callback more than once; your handler must tolerate duplicates.
  • Response — respond 200 OK within a few seconds. The queue does not block on slow handlers, but persistent failures may stop further deliveries.
  • HTTPS required — plain http:// URLs are rejected.

Request Parameters - Grok Voice Text to Speech generate request

Field Type Required Default Description
text string Yes The input text to convert into speech. Supports inline speech tags for prosody control (e.g., <break time="500ms"/>).
voice string Yes The voice model to use. Valid values: eve, adam, lucy, max, sage.
language string No auto Language code for pronunciation. Use auto for automatic detection or specify a BCP-47 code (e.g., en-US, es-ES).

Minimum Request

{
  "text": "Welcome to the future of artificial intelligence and creative expression",
  "voice": "eve"
}

Full Request (all options)

{
  "text": "Welcome to the future of artificial intelligence and creative expression",
  "voice": "eve",
  "language": "auto"
}

Response

{
  "request_id": "xai-text-to-speech_019dxxxx-xxxx-xxxx-xxxx-xxxxxxxxxxxx",
  "status": "QUEUED",
  "polling_url": "https://gateway.pixazo.ai/v2/requests/status/xai-text-to-speech_019dxxxx-xxxx-xxxx-xxxx-xxxxxxxxxxxx"
}

Request Headers

Header Value
Content-Type application/json
Cache-Control no-cache
Ocp-Apim-Subscription-Key Your API subscription key

Response Handling

Common status codes for Grok Voice Text to Speech generate request.

Code Meaning
202 Accepted — Request queued
Bad Request
401 Unauthorized
403 Forbidden
404 Not Found
Too Many Requests
500 Internal Server Error

Error Responses

Queue system errors and model validation errors.

Queue System Errors

// 402 — Insufficient balance
{
  "error": "Insufficient Balance",
  "message": "Your wallet does not have enough balance."
}
// 400 — Model not found
{
  "error": "Model not found",
  "message": "Model 'xai-text-to-speech' not found or is disabled"
}

Error via Status/Webhook

{
  "request_id": "xai-text-to-speech_019dxxxx-xxxx-xxxx-xxxx-xxxxxxxxxxxx",
  "status": "ERROR",
  "model_id": "xai-text-to-speech",
  "error": "Description of the error",
  "output": null
}

Retrieving Results

Poll the universal status endpoint to check progress and retrieve results.

Endpoint

GET https://gateway.pixazo.ai/v2/requests/status/{request_id}
Ocp-Apim-Subscription-Key: YOUR_API_KEY

cURL Example

curl -H "Ocp-Apim-Subscription-Key: YOUR_API_KEY" \
  "https://gateway.pixazo.ai/v2/requests/status/xai-text-to-speech_019dxxxx-xxxx-xxxx-xxxx-xxxxxxxxxxxx"

Response (Completed)

{
  "request_id": "xai-text-to-speech_019dxxxx-xxxx-xxxx-xxxx-xxxxxxxxxxxx",
  "status": "COMPLETED",
  "model_id": "xai-text-to-speech",
  "error": null,
  "output": {
    "media_url": [
      "https://pub-582b7213209642b9b995c96c95a30381.r2.dev/v1/xai-text-to-speech_019dxxxx-xxxx/output.ext"
    ],
    "media_type": "application/octet-stream"
  },
  "created_at": "2026-03-31T10:00:00.000Z",
  "updated_at": "2026-03-31T10:00:15.000Z",
  "completed_at": "2026-03-31T10:00:15.000Z"
}

Response Fields

FieldTypeDescription
request_idstringUnique request identifier
statusstringQUEUED, PROCESSING, COMPLETED, FAILED, or ERROR
model_idstringModel that processed the request
errorstring|nullError message if failed
output.media_urlarrayURLs to generated media (R2 CDN)
output.media_typestringMIME type of the output
created_atstringWhen request was created
completed_atstring|nullWhen request completed
polling_urlstringStatus URL (initial response only)

Status Values

StatusDescription
QUEUEDRequest accepted, waiting to be processed
PROCESSINGBeing processed by the model
COMPLETEDDone — output contains the result
FAILEDFailed — check error field
ERRORSystem error — not charged

Status Flow

QUEUED → PROCESSING → COMPLETED
                    → FAILED
                    → ERROR

Typical Workflow

  1. Send a generate request to the API endpoint
  2. Save the request_id from the response
  3. Poll every 5-10 seconds: GET /v2/requests/status/{request_id}
  4. When status is "COMPLETED", download from output.media_url

Tip: Use X-Webhook-URL header to get a callback instead of polling.

Grok Voice Text to Speech API Pricing

Your request will cost $0.015 per 1,000 characters.
one of the cheapest expressive voices at $0.015 per 1,000 characters
equivalent to $15 per 1M characters
3. Grok Voice Speech to Text API

Grok Voice Speech to Text API Documentation

xAI's Grok speech-to-text. Transcribes 25 languages and reports the length of the recording alongside the transcript, so you always know what you were billed for. Asynchronous: submit returns a request_id; poll the status endpoint until the request is COMPLETED, then download the audio.

POST https://gateway.pixazo.ai/grok-stt/v1/speech-to-text

Authentication

All requests require an API key passed via header.

HeaderTypeRequiredDescription
Ocp-Apim-Subscription-KeystringYesYour API subscription key

Speech to Text - Grok Speech to Text

Request Code

POST https://gateway.pixazo.ai/grok-stt/v1/speech-to-text
Content-Type: application/json
Ocp-Apim-Subscription-Key: YOUR_SUBSCRIPTION_KEY

{
  "audio_url": "https://pub-582b7213209642b9b995c96c95a30381.r2.dev/doc-assets/audio/speech-17s.mp3"
}
import requests

url = "https://gateway.pixazo.ai/grok-stt/v1/speech-to-text"
headers = {
    "Content-Type": "application/json",
    "Ocp-Apim-Subscription-Key": "YOUR_SUBSCRIPTION_KEY"
}
data = {
  "audio_url": "https://pub-582b7213209642b9b995c96c95a30381.r2.dev/doc-assets/audio/speech-17s.mp3"
}

resp = requests.post(url, json=data, headers=headers)
print(resp.json())
const res = await fetch("https://gateway.pixazo.ai/grok-stt/v1/speech-to-text", {
  method: "POST",
  headers: {
    "Content-Type": "application/json",
    "Ocp-Apim-Subscription-Key": "YOUR_SUBSCRIPTION_KEY"
  },
  body: JSON.stringify({
  "audio_url": "https://pub-582b7213209642b9b995c96c95a30381.r2.dev/doc-assets/audio/speech-17s.mp3"
})
});
console.log(await res.json());
curl -X POST 'https://gateway.pixazo.ai/grok-stt/v1/speech-to-text' \
  -H 'Content-Type: application/json' \
  -H 'Ocp-Apim-Subscription-Key: YOUR_SUBSCRIPTION_KEY' \
  --data-raw '{"audio_url": "https://pub-582b7213209642b9b995c96c95a30381.r2.dev/doc-assets/audio/speech-17s.mp3"}'

Output

{
  "request_id": "grok-stt_019dxxxx-xxxx-xxxx-xxxx-xxxxxxxxxxxx",
  "status": "QUEUED",
  "polling_url": "https://gateway.pixazo.ai/v2/requests/status/grok-stt_019dxxxx-xxxx-xxxx-xxxx-xxxxxxxxxxxx"
}

Webhook (Optional)

Instead of polling, you can receive a Webhook callback when the request reaches a terminal state. Provide a Webhook URL via header on the submit request.

HeaderRequiredDescription
X-Webhook-URLTo enableHTTPS URL to receive the Webhook callback.
X-Webhook-ModeNoterminal (default, one callback on COMPLETED/ERROR) or sync (per-poll callbacks).

Example: enable Webhook

curl -X POST 'https://gateway.pixazo.ai/grok-stt/v1/speech-to-text' \
  -H 'Content-Type: application/json' \
  -H 'Ocp-Apim-Subscription-Key: YOUR_SUBSCRIPTION_KEY' \
  -H 'X-Webhook-URL: https://your-server.com/webhook' \
  --data-raw '{"audio_url": "https://pub-582b7213209642b9b995c96c95a30381.r2.dev/doc-assets/audio/speech-17s.mp3"}'

Callback Payload (success)

{
  "request_id": "grok-stt_019dxxxx-xxxx-xxxx-xxxx-xxxxxxxxxxxx",
  "status": "COMPLETED",
  "model_id": "grok-stt",
  "output": {
    "media_url": [
      "https://pub-582b7213209642b9b995c96c95a30381.r2.dev/v1/{request_id}/transcript.json"
    ],
    "media_type": "application/json"
  },
  "duration": 19.17,
  "created_at": "2026-08-01T09:14:16.102Z",
  "completed_at": "2026-08-01T09:14:22.870Z"
}

Failure callback shape

{
  "request_id": "grok-stt_019dxxxx-xxxx-xxxx-xxxx-xxxxxxxxxxxx",
  "status": "ERROR",
  "model_id": "grok-stt",
  "error": "Description of the failure"
}

Delivery semantics

  • terminal mode: one Webhook callback when the request is COMPLETED or ERROR.
  • sync mode: a Webhook callback on each status change.
  • Callbacks are idempotent on request_id — de-duplicate on it.
  • Respond 200 within a few seconds; the Webhook endpoint must be HTTPS.

Request Parameters

ParameterRequiredTypeDefaultAllowed values / rangeDescription
audio_urlYesstringa publicly reachable http(s) urlThe recording to transcribe. We fetch it server-side, so it must be reachable from the internet — a signed url is fine, a private one is not. audio is accepted as an alias.
languageNostringISO 639-1, e.g. enOmit this to let the model detect the language — that is the default. auto means the same. Set it to force one of the recording.
promptNostringup to 2,000 charactersBias the transcription toward expected wording — names, jargon, spellings.

Voices

Returns text, the detected language and the clip duration.

Example Request

{
  "audio_url": "https://pub-582b7213209642b9b995c96c95a30381.r2.dev/doc-assets/audio/speech-17s.mp3"
}

Example Response

{
  "request_id": "grok-stt_019dxxxx-xxxx-xxxx-xxxx-xxxxxxxxxxxx",
  "status": "QUEUED",
  "polling_url": "https://gateway.pixazo.ai/v2/requests/status/grok-stt_019dxxxx-xxxx-xxxx-xxxx-xxxxxxxxxxxx"
}

Request Headers

HeaderRequiredDescription
Content-TypeYesapplication/json
Ocp-Apim-Subscription-KeyYesYour API subscription key.
X-Webhook-URLNoEnable Webhook callbacks (see Webhook section).

Response Handling

Status CodeMeaning
202Accepted — request queued; returns request_id and polling_url.
400Bad request — a missing or out-of-range parameter. The message names the field.
401Unauthorized — missing or invalid subscription key.
402Insufficient balance.
429Too many requests.
500Internal server error.

Retrieving Results

Poll the status endpoint with the request_id from the submit response until status is COMPLETED (or ERROR), then download output.media_url.

curl 'https://gateway.pixazo.ai/v2/requests/status/grok-stt_019dxxxx-xxxx-xxxx-xxxx-xxxxxxxxxxxx' \
  -H 'Ocp-Apim-Subscription-Key: YOUR_SUBSCRIPTION_KEY'

Completed response

{
  "request_id": "grok-stt_019dxxxx-xxxx-xxxx-xxxx-xxxxxxxxxxxx",
  "status": "COMPLETED",
  "model_id": "grok-stt",
  "output": {
    "media_url": [
      "https://pub-582b7213209642b9b995c96c95a30381.r2.dev/v1/{request_id}/transcript.json"
    ],
    "media_type": "application/json"
  },
  "duration": 19.17,
  "created_at": "2026-08-01T09:14:16.102Z",
  "completed_at": "2026-08-01T09:14:22.870Z"
}

Response Fields

FieldTypeDescription
request_idstringUnique request identifier.
statusstringQUEUED, PROCESSING, COMPLETED or ERROR.
model_idstringThe model that handled the request.
output.media_urlarrayURL of the generated audio file.
output.media_typestringMIME type of the audio.
created_atstringRequest creation timestamp.
completed_atstringCompletion timestamp.
errorstringError message when status is ERROR.

Status Values & Flow

QUEUEDPROCESSINGCOMPLETED (success) or ERROR (failure).

Pricing

Billed at $0.001667 per minute of generated audio, rounded up to the next whole minute. You are charged for the audio produced, not the text you submit.

Audio producedBilled minutesCost
A 10-second clip1$0.001667
A 45-second clip1$0.001667
A 3-minute narration3$0.005001
A 10-minute narration10$0.01667

Failed requests are not billed.

Grok Voice Speech to Text API Pricing

Your request will cost $0.0017 per minute of audio transcribed.
an hour of audio is $0.10
equivalent to $0.10 per hour of audio

⚡ Performance

Live usage measured on Pixazo's gateway, split by model version. Generation time is how long a generation takes end-to-end (lower is better). Success rate is the percent of generations that complete (higher is better).

Show data for the last
Generations
400last 30d
~13 per day
Success rate
75.0%
of completed generations
Generation time
25.9savg
p95 25.9s
Requests
Aug 24max 200Sep 22
Grok Voice Speech to Text APIAvg 7/day
Grok Voice Agent APIAvg 3/day
Grok Voice Text to Speech APIAvg 3/day
Generation Time
Aug 24max 25.9sSep 22
Grok Voice Speech to Text APIAvg 
Grok Voice Agent APIAvg 25.9s
Grok Voice Text to Speech APIAvg 
Error Rate
Aug 24max 100.0%Sep 22
Grok Voice Speech to Text APIAvg 0.0%
Grok Voice Agent APIAvg 0.0%
Grok Voice Text to Speech APIAvg 100.0%

〰 Uptime

Percent of generations that succeeded over the selected period, per model version.

Avg. Success Rate (30d)
75.00%
across all generations of this model family
Uptime
Aug 24max 100%Sep 22
Grok Voice Speech to Text APIAvg 100.00%
Grok Voice Agent APIAvg 100.00%
Grok Voice Text to Speech APIAvg 0.00%