---
type: AI Model
id: grok-voice
title: Grok Voice API
provider: xAI
description: "xAI's Grok voice models: a speech-to-speech voice agent that listens to a recording and answers aloud in one of 28 voices, text to speech in five expressive voices, and speech to text across 25 languages. Every response reports the length of audio it was billed on, so you always know what you paid for."
resource: https://www.pixazo.ai/models/grok-voice
docs_url: https://www.pixazo.ai/models/grok-voice
latest_version: v1
tags:
  - text-to-speech
  - speech-to-text
  - speech-to-speech
  - audio-generation
  - xai
variants:
  - id: grok-voice
    name: Grok Voice Agent API
    version: v1
    capabilities:
      - Speech to Speech
  - id: xai-text-to-speech
    name: Grok Voice Text to Speech API
    version: v1
    capabilities:
      - Text to Speech
  - id: grok-stt
    name: Grok Voice Speech to Text API
    version: v1
    capabilities:
      - Speech to Text
timestamp: 2026-09-22T08:48:40.013Z
---

# Grok Voice API

> Provider: **xAI**
> Source: https://www.pixazo.ai/models/grok-voice

xAI's Grok voice models: a speech-to-speech voice agent that listens to a recording and answers aloud in one of 28 voices, text to speech in five expressive voices, and speech to text across 25 languages. Every response reports the length of audio it was billed on, so you always know what you paid for.

## Grok Voice Agent API

### Speech to Speech

## Grok Voice API Documentation

A speech-to-speech voice agent by xAI. Send a recording of the user speaking and get Grok's spoken reply back as a WAV file. Optionally set the agent's persona with `prompt` and pick one of 28 built-in voices. Asynchronous: submit returns a `request_id`; poll the status endpoint until the request is `COMPLETED`, then download the reply audio from `output.media_url`.

```
POST https://gateway.pixazo.ai/grok-voice/v1/speech-to-speech
```

## Authentication

All requests require an API key passed via header.

Header

Type

Required

Description

Ocp-Apim-Subscription-Key

string

Yes

Your API subscription key

## Speech to Speech - Grok Voice

## Request Code

HTTP Python JavaScript cURL

```
POST https://gateway.pixazo.ai/grok-voice/v1/speech-to-speech
Content-Type: application/json
Ocp-Apim-Subscription-Key: YOUR_SUBSCRIPTION_KEY

{
  "audio_url": "https://pub-582b7213209642b9b995c96c95a30381.r2.dev/doc-assets/audio/transcribe-sample.wav",
  "prompt": "You are a friendly assistant. Answer briefly and concretely.",
  "voice": "eve"
}
```

```
import requests

url = "https://gateway.pixazo.ai/grok-voice/v1/speech-to-speech"
headers = {
    "Content-Type": "application/json",
    "Ocp-Apim-Subscription-Key": "YOUR_SUBSCRIPTION_KEY"
}
data = {
  "audio_url": "https://pub-582b7213209642b9b995c96c95a30381.r2.dev/doc-assets/audio/transcribe-sample.wav",
  "prompt": "You are a friendly assistant. Answer briefly and concretely.",
  "voice": "eve"
}

resp = requests.post(url, json=data, headers=headers)
print(resp.json())
```

```
const res = await fetch("https://gateway.pixazo.ai/grok-voice/v1/speech-to-speech", {
  method: "POST",
  headers: {
    "Content-Type": "application/json",
    "Ocp-Apim-Subscription-Key": "YOUR_SUBSCRIPTION_KEY"
  },
  body: JSON.stringify({
  "audio_url": "https://pub-582b7213209642b9b995c96c95a30381.r2.dev/doc-assets/audio/transcribe-sample.wav",
  "prompt": "You are a friendly assistant. Answer briefly and concretely.",
  "voice": "eve"
})
});
console.log(await res.json());
```

```
curl -X POST 'https://gateway.pixazo.ai/grok-voice/v1/speech-to-speech' \
  -H 'Content-Type: application/json' \
  -H 'Ocp-Apim-Subscription-Key: YOUR_SUBSCRIPTION_KEY' \
  --data-raw '{"audio_url": "https://pub-582b7213209642b9b995c96c95a30381.r2.dev/doc-assets/audio/transcribe-sample.wav", "prompt": "You are a friendly assistant. Answer briefly and concretely.", "voice": "eve"}'
```

## Output

```
{
  "request_id": "grok-voice_019dxxxx-xxxx-xxxx-xxxx-xxxxxxxxxxxx",
  "status": "QUEUED",
  "polling_url": "https://gateway.pixazo.ai/v2/requests/status/grok-voice_019dxxxx-xxxx-xxxx-xxxx-xxxxxxxxxxxx"
}
```

[Try Now](https://api.pixazo.ai/api-details#api=grok-voice&operation=speech-to-speech)

## Webhook (Optional)

Instead of polling, you can receive a Webhook callback when the request reaches a terminal state. Provide a Webhook URL via header on the submit request.

Header

Required

Description

X-Webhook-URL

To enable

HTTPS URL to receive the Webhook callback.

X-Webhook-Mode

No

`terminal` (default, one callback on COMPLETED/ERROR) or `sync` (per-poll callbacks).

### Example: enable Webhook

```
curl -X POST 'https://gateway.pixazo.ai/grok-voice/v1/speech-to-speech' \
  -H 'Content-Type: application/json' \
  -H 'Ocp-Apim-Subscription-Key: YOUR_SUBSCRIPTION_KEY' \
  -H 'X-Webhook-URL: https://your-server.com/webhook' \
  --data-raw '{"audio_url": "https://pub-582b7213209642b9b995c96c95a30381.r2.dev/doc-assets/audio/transcribe-sample.wav", "prompt": "You are a friendly assistant. Answer briefly and concretely.", "voice": "eve"}'
```

### Callback Payload (success)

```
{
  "request_id": "grok-voice_019dxxxx-xxxx-xxxx-xxxx-xxxxxxxxxxxx",
  "status": "COMPLETED",
  "model_id": "grok-voice",
  "output": {
    "media_url": [
      "https://pub-582b7213209642b9b995c96c95a30381.r2.dev/v1/{request_id}/output.wav"
    ],
    "media_type": "audio/wav"
  },
  "created_at": "2026-09-21T09:14:16.102Z",
  "completed_at": "2026-09-21T09:14:31.870Z"
}
```

### Failure callback shape

```
{
  "request_id": "grok-voice_019dxxxx-xxxx-xxxx-xxxx-xxxxxxxxxxxx",
  "status": "ERROR",
  "model_id": "grok-voice",
  "error": "Description of the failure"
}
```

#### Delivery semantics

-   **terminal** mode: one Webhook callback when the request is COMPLETED or ERROR.
-   **sync** mode: a Webhook callback on each status change.
-   Callbacks are idempotent on `request_id` — de-duplicate on it.
-   Respond `200` within a few seconds; the Webhook endpoint must be HTTPS.

## Request Parameters

Parameter

Required

Type

Default

Allowed values / range

Description

`audio_url`

Yes

string

—

https URL; up to 10 minutes and 50 MB

A recording of the user speaking — the thing the agent listens to and answers. Fetched by the gateway, so the URL must be reachable without authentication. Most common audio formats are accepted; the file is converted to 16-bit 24 kHz mono PCM before the model hears it.

`prompt`

No

string

— (Grok's default persona)

up to 20,000 characters

System instructions describing the agent's persona and the conversation context — who it is, how it should answer, what it knows. This is **not** text to be read aloud; the reply is composed from the recording. Omit it to use Grok's default persona.

`voice`

No

string

— (Grok's default voice)

one of the 28 voices listed below

Built-in xAI voice for the spoken reply. Case-sensitive; see the Voices section.

### Audio requirements

-   At most **10 minutes** of audio and at most **50 MB** per request. Longer recordings fail once the file is read, reported as `status: "ERROR"`.
-   Most common formats are accepted (WAV, MP3, M4A, OGG, FLAC and similar). The audio is converted to 16-bit 24 kHz mono PCM before it reaches the model, so you do not need to convert it yourself.
-   `audio_url` must be an `https` URL and publicly reachable.
-   Cost scales with the length of the **reply** the agent generates, not the recording you send; the current rate is in the pricing panel on this page.

### Fields this endpoint does not accept

The provider's agent tools (`tools.web_search`, `tools.x_search`, `tools.mcp_servers`) are not available on this endpoint: a body containing `tools` is rejected with a synchronous `400`. Likewise any field that is not listed in the table above, such as `duration` or `seconds`, is rejected rather than ignored.

## Voices

Pass one of these 28 names as `voice`, exactly as written (lowercase). Omit the field to use Grok's default voice.

`carina`, `zagan`, `helix`, `orion`, `luna`, `iris`, `altair`, `zenith`, `perseus`, `helios`, `lux`, `kepler`, `rigel`, `cosmo`, `celeste`, `ursa`, `sirius`, `lumen`, `castor`, `naksh`, `atlas`, `aurora`, `liora`, `ara`, `eve`, `leo`, `rex`, `sal`

## Example Request

```
{
  "audio_url": "https://pub-582b7213209642b9b995c96c95a30381.r2.dev/doc-assets/audio/transcribe-sample.wav",
  "prompt": "You are a friendly assistant. Answer briefly and concretely.",
  "voice": "eve"
}
```

## Example Response

```
{
  "request_id": "grok-voice_019dxxxx-xxxx-xxxx-xxxx-xxxxxxxxxxxx",
  "status": "QUEUED",
  "polling_url": "https://gateway.pixazo.ai/v2/requests/status/grok-voice_019dxxxx-xxxx-xxxx-xxxx-xxxxxxxxxxxx"
}
```

## Request Headers

Header

Required

Description

Content-Type

Yes

`application/json`

Ocp-Apim-Subscription-Key

Yes

Your API subscription key.

X-Webhook-URL

No

Enable Webhook callbacks (see Webhook section).

## Response Handling

Status Code

Meaning

202

Accepted — request queued; returns `request_id` and `polling_url`.

400

Bad request — `audio_url` missing or not a string, `prompt` over 20,000 characters, `voice` outside the 28 listed, or a field this endpoint does not accept (such as `tools`). Anything checkable from the request body alone.

401

Unauthorized — missing or invalid subscription key.

402

Insufficient balance.

429

Too many requests.

500

Internal server error.

Only what can be judged from the request body itself is rejected synchronously. Everything that needs the audio to be fetched and read is reported through the status endpoint instead, as `status: "ERROR"`. Failed requests are not billed — the hold is released in full.

Condition

How it surfaces

`audio_url` missing, or the wrong type

synchronous `400`

`voice` not one of the 28 names, or `prompt` over 20,000 characters

synchronous `400`

`tools`, or any other field not in the parameter table

synchronous `400`

`audio_url` not https, not reachable, or not public

`status: "ERROR"`

Audio longer than 10 minutes, or a file over 50 MB

`status: "ERROR"`

Audio in a format the provider cannot decode

`status: "ERROR"`

## Retrieving Results

Poll the status endpoint with the `request_id` from the submit response until `status` is `COMPLETED` (or `FAILED`/`ERROR`), then download the reply audio from `output.media_url`.

```
curl 'https://gateway.pixazo.ai/v2/requests/status/grok-voice_019dxxxx-xxxx-xxxx-xxxx-xxxxxxxxxxxx' \
  -H 'Ocp-Apim-Subscription-Key: YOUR_SUBSCRIPTION_KEY'
```

### Completed response

```
{
  "request_id": "grok-voice_019dxxxx-xxxx-xxxx-xxxx-xxxxxxxxxxxx",
  "status": "COMPLETED",
  "model_id": "grok-voice",
  "output": {
    "media_url": [
      "https://pub-582b7213209642b9b995c96c95a30381.r2.dev/v1/{request_id}/output.wav"
    ],
    "media_type": "audio/wav"
  },
  "created_at": "2026-09-21T09:14:16.102Z",
  "completed_at": "2026-09-21T09:14:31.870Z"
}
```

The reply is returned as **audio only** — a 24 kHz mono WAV file. A text transcript of the reply is not returned by this endpoint; if you need one, pass the WAV to a speech-to-text API.

## Response Fields

Field

Type

Description

request\_id

string

Unique request identifier.

status

string

QUEUED, PROCESSING, COMPLETED, FAILED, or ERROR.

model\_id

string

The model that handled the request.

output.media\_url

array

URL of the reply audio (WAV, 24 kHz mono).

output.media\_type

string

`audio/wav`.

created\_at

string

Request creation timestamp.

completed\_at

string

Completion timestamp.

error

string

Error message when `status` is FAILED/ERROR.

## Status Values & Flow

`QUEUED` → `PROCESSING` → `COMPLETED` (success) or `FAILED`/`ERROR` (failure).

### How billing is measured

Billing is measured on the **length of the reply audio the model generates**, per second, rounded up to the next whole second — not on the length of the recording you send, and not on how long the request takes. A hold is placed when the request is submitted, because the reply length is not known until the model has answered; the hold is reduced to the real cost once the reply audio is ready, and released in full if the request fails. To keep replies short, say so in `prompt`.

The current rate is shown in the pricing panel on this page.

## Grok Voice Text to Speech API

### Text to Speech

## Base URL

```
https://gateway.pixazo.ai/xai-text-to-speech/v1/text-to-speech
```

## Authentication

All requests require an API key passed via header.

Header

Type

Required

Description

Ocp-Apim-Subscription-Key

string

Yes

Your API subscription key

## xAI Text to Speech generate request - xAI Text to Speech

## Request Code

HTTP Python JavaScript cURL

```
POST https://gateway.pixazo.ai/xai-text-to-speech/v1/text-to-speech
Content-Type: application/json
Cache-Control: no-cache
Ocp-Apim-Subscription-Key: YOUR_SUBSCRIPTION_KEY

{
  "text": "Welcome to the future of artificial intelligence and creative expression",
  "voice": "eve",
  "language": "auto"
}
```

```
import requests

url = "https://gateway.pixazo.ai/xai-text-to-speech/v1/text-to-speech"
headers = {
    "Content-Type": "application/json",
    "Cache-Control": "no-cache",
    "Ocp-Apim-Subscription-Key": "YOUR_SUBSCRIPTION_KEY"
}
data = {
    "text": "Welcome to the future of artificial intelligence and creative expression",
    "voice": "eve",
    "language": "auto"
}

response = requests.post(url, json=data, headers=headers)
print(response.json())
```

```
const url = 'https://gateway.pixazo.ai/xai-text-to-speech/v1/text-to-speech';

const data = {
  text: 'Welcome to the future of artificial intelligence and creative expression',
  voice: 'eve',
  language: 'auto'
};

fetch(url, {
  method: 'POST',
  headers: {
    'Content-Type': 'application/json',
    'Cache-Control': 'no-cache',
    'Ocp-Apim-Subscription-Key': 'YOUR_SUBSCRIPTION_KEY'
  },
  body: JSON.stringify(data)
})
.then(response => response.json())
.then(data => console.log(data))
.catch(error => console.error('Error:', error));
```

```
curl -X POST "https://gateway.pixazo.ai/xai-text-to-speech/v1/text-to-speech" \
  -H "Content-Type: application/json" \
  -H "Cache-Control: no-cache" \
  -H "Ocp-Apim-Subscription-Key: YOUR_SUBSCRIPTION_KEY" \
  --data-raw '{
    "text": "Welcome to the future of artificial intelligence and creative expression",
    "voice": "eve",
    "language": "auto"
  }'
```

## Output

```
{
  "request_id": "xai-text-to-speech_019dxxxx-xxxx-xxxx-xxxx-xxxxxxxxxxxx",
  "status": "QUEUED",
  "polling_url": "https://gateway.pixazo.ai/v2/requests/status/xai-text-to-speech_019dxxxx-xxxx-xxxx-xxxx-xxxxxxxxxxxx"
}
```

[Try Now](https://api.pixazo.ai/api-details#api=xai-text-to-speech&operation=xai-text-to-speech-request)

## Webhook (Optional)

Add the `X-Webhook-URL` header to your submit request to receive a `POST` callback when the job completes — no polling required.

**Using curl?** These are HTTP request headers — pass each with `-H`, e.g. `-H "X-Webhook-URL: https://your-server.com/webhook/callback"`. Do not paste them as bare lines, and end every line of a multi-line command with `\`.

### Webhook Headers

Header

Required

Default

Description

`X-Webhook-URL`

Yes (to enable)

—

HTTPS endpoint on your server that will receive the `POST` callback. Must respond `2xx` within a few seconds (process async if needed).

`X-Webhook-Mode`

No

`terminal`

`terminal` — fires once at the final status (`COMPLETED`/`FAILED`/`ERROR`). `sync` — fires on every poll cycle plus the terminal event, and caps the queue’s polling delay at **15s** for tighter progress updates.

### Example: enable webhook

```
X-Webhook-URL: https://your-server.com/webhook/callback
X-Webhook-Mode: terminal
```

### Callback Payload

Your endpoint receives a `POST application/json` with the same shape as the `GET /v2/requests/status/{request_id}` response. Example terminal callback (mode `terminal`):

```
{
  "request_id": "xai-text-to-speech_019dxxxx-xxxx-xxxx-xxxx-xxxxxxxxxxxx",
  "status": "COMPLETED",
  "model_id": "xai-text-to-speech",
  "error": null,
  "output": {
    "media_url": [
      "https://pub-582b7213209642b9b995c96c95a30381.r2.dev/v1/xai-text-to-speech_019dxxxx-xxxx-xxxx-xxxx-xxxxxxxxxxxx/output.wav"
    ],
    "media_type": "audio/wav"
  },
  "created_at": "2026-05-22T13:17:32.110Z",
  "updated_at": "2026-05-22 13:19:23",
  "completed_at": "2026-05-22 13:19:23"
}
```

### Failure callback shape

```
{
  "request_id": "xai-text-to-speech_019dxxxx-xxxx-xxxx-xxxx-xxxxxxxxxxxx",
  "status": "ERROR",
  "model_id": "xai-text-to-speech",
  "error": "Description of the error",
  "output": null,
  "created_at": "...",
  "updated_at": "...",
  "completed_at": "..."
}
```

### Delivery semantics

-   **terminal mode (default)** — exactly one `POST` when the request reaches a terminal status. No callback during `PROCESSING`.
-   **sync mode** — `POST` on every status poll (with delay capped at ~15s) plus a final `POST` at terminal status. Use when you want progress updates.
-   **Idempotency** — use `request_id` as your idempotency key. Network retries can deliver the same callback more than once; your handler must tolerate duplicates.
-   **Response** — respond `200 OK` within a few seconds. The queue does not block on slow handlers, but persistent failures may stop further deliveries.
-   **HTTPS required** — plain `http://` URLs are rejected.

## Request Parameters - xAI Text to Speech generate request

Field

Type

Required

Default

Description

text

string

Yes

—

The input text to convert into speech. Supports inline speech tags for prosody control (e.g., <break time="500ms"/>).

voice

string

Yes

—

The voice model to use. Valid values: eve, adam, lucy, max, sage.

language

string

No

auto

Language code for pronunciation. Use auto for automatic detection or specify a BCP-47 code (e.g., en-US, es-ES).

## Minimum Request

```
{
  "text": "Welcome to the future of artificial intelligence and creative expression",
  "voice": "eve"
}
```

## Full Request (all options)

```
{
  "text": "Welcome to the future of artificial intelligence and creative expression",
  "voice": "eve",
  "language": "auto"
}
```

## Response

```
{
  "request_id": "xai-text-to-speech_019dxxxx-xxxx-xxxx-xxxx-xxxxxxxxxxxx",
  "status": "QUEUED",
  "polling_url": "https://gateway.pixazo.ai/v2/requests/status/xai-text-to-speech_019dxxxx-xxxx-xxxx-xxxx-xxxxxxxxxxxx"
}
```

## Request Headers

Header

Value

Content-Type

application/json

Cache-Control

no-cache

Ocp-Apim-Subscription-Key

Your API subscription key

## Response Handling

Common status codes for xAI Text to Speech generate request.

Code

Meaning

202

Accepted — Request queued

400

Bad Request

401

Unauthorized

403

Forbidden

404

Not Found

429

Too Many Requests

500

Internal Server Error

## Error Responses

Queue system errors and model validation errors.

### Queue System Errors

```
// 402 — Insufficient balance
{
  "error": "Insufficient Balance",
  "message": "Your wallet does not have enough balance."
}
```

```
// 400 — Model not found
{
  "error": "Model not found",
  "message": "Model 'xai-text-to-speech' not found or is disabled"
}
```

### Error via Status/Webhook

```
{
  "request_id": "xai-text-to-speech_019dxxxx-xxxx-xxxx-xxxx-xxxxxxxxxxxx",
  "status": "ERROR",
  "model_id": "xai-text-to-speech",
  "error": "Description of the error",
  "output": null
}
```

## Retrieving Results

Poll the universal status endpoint to check progress and retrieve results.

### Endpoint

```
GET https://gateway.pixazo.ai/v2/requests/status/{request_id}
Ocp-Apim-Subscription-Key: YOUR_API_KEY
```

## cURL Example

```
curl -H "Ocp-Apim-Subscription-Key: YOUR_API_KEY" \
  "https://gateway.pixazo.ai/v2/requests/status/xai-text-to-speech_019dxxxx-xxxx-xxxx-xxxx-xxxxxxxxxxxx"
```

## Response (Completed)

```
{
  "request_id": "xai-text-to-speech_019dxxxx-xxxx-xxxx-xxxx-xxxxxxxxxxxx",
  "status": "COMPLETED",
  "model_id": "xai-text-to-speech",
  "error": null,
  "output": {
    "media_url": [
      "https://pub-582b7213209642b9b995c96c95a30381.r2.dev/v1/xai-text-to-speech_019dxxxx-xxxx/output.ext"
    ],
    "media_type": "application/octet-stream"
  },
  "created_at": "2026-03-31T10:00:00.000Z",
  "updated_at": "2026-03-31T10:00:15.000Z",
  "completed_at": "2026-03-31T10:00:15.000Z"
}
```

## Response Fields

Field

Type

Description

request\_id

string

Unique request identifier

status

string

QUEUED, PROCESSING, COMPLETED, FAILED, or ERROR

model\_id

string

Model that processed the request

error

string|null

Error message if failed

output.media\_url

array

URLs to generated media (R2 CDN)

output.media\_type

string

MIME type of the output

created\_at

string

When request was created

completed\_at

string|null

When request completed

polling\_url

string

Status URL (initial response only)

## Status Values

Status

Description

QUEUED

Request accepted, waiting to be processed

PROCESSING

Being processed by the model

COMPLETED

Done — output contains the result

FAILED

Failed — check error field

ERROR

System error — not charged

## Status Flow

```
QUEUED → PROCESSING → COMPLETED
                    → FAILED
                    → ERROR
```

## Typical Workflow

1.  **Send a generate request** to the API endpoint
2.  **Save the `request_id`** from the response
3.  **Poll** every 5-10 seconds: `GET /v2/requests/status/{request_id}`
4.  **When `status` is `"COMPLETED"`**, download from `output.media_url`

**Tip:** Use `X-Webhook-URL` header to get a callback instead of polling.

## Grok Voice Speech to Text API

### Speech to Text

## Grok Speech to Text API Documentation

xAI's Grok speech-to-text. Transcribes 25 languages and reports the length of the recording alongside the transcript, so you always know what you were billed for. Asynchronous: submit returns a `request_id`; poll the status endpoint until the request is `COMPLETED`, then download the audio.

```
POST https://gateway.pixazo.ai/grok-stt/v1/speech-to-text
```

## Authentication

All requests require an API key passed via header.

Header

Type

Required

Description

Ocp-Apim-Subscription-Key

string

Yes

Your API subscription key

## Speech to Text - Grok Speech to Text

## Request Code

HTTP Python JavaScript cURL

```
POST https://gateway.pixazo.ai/grok-stt/v1/speech-to-text
Content-Type: application/json
Ocp-Apim-Subscription-Key: YOUR_SUBSCRIPTION_KEY

{
  "audio_url": "https://pub-582b7213209642b9b995c96c95a30381.r2.dev/doc-assets/audio/speech-17s.mp3"
}
```

```
import requests

url = "https://gateway.pixazo.ai/grok-stt/v1/speech-to-text"
headers = {
    "Content-Type": "application/json",
    "Ocp-Apim-Subscription-Key": "YOUR_SUBSCRIPTION_KEY"
}
data = {
  "audio_url": "https://pub-582b7213209642b9b995c96c95a30381.r2.dev/doc-assets/audio/speech-17s.mp3"
}

resp = requests.post(url, json=data, headers=headers)
print(resp.json())
```

```
const res = await fetch("https://gateway.pixazo.ai/grok-stt/v1/speech-to-text", {
  method: "POST",
  headers: {
    "Content-Type": "application/json",
    "Ocp-Apim-Subscription-Key": "YOUR_SUBSCRIPTION_KEY"
  },
  body: JSON.stringify({
  "audio_url": "https://pub-582b7213209642b9b995c96c95a30381.r2.dev/doc-assets/audio/speech-17s.mp3"
})
});
console.log(await res.json());
```

```
curl -X POST 'https://gateway.pixazo.ai/grok-stt/v1/speech-to-text' \
  -H 'Content-Type: application/json' \
  -H 'Ocp-Apim-Subscription-Key: YOUR_SUBSCRIPTION_KEY' \
  --data-raw '{"audio_url": "https://pub-582b7213209642b9b995c96c95a30381.r2.dev/doc-assets/audio/speech-17s.mp3"}'
```

## Output

```
{
  "request_id": "grok-stt_019dxxxx-xxxx-xxxx-xxxx-xxxxxxxxxxxx",
  "status": "QUEUED",
  "polling_url": "https://gateway.pixazo.ai/v2/requests/status/grok-stt_019dxxxx-xxxx-xxxx-xxxx-xxxxxxxxxxxx"
}
```

[Try Now](https://api.pixazo.ai/api-details#api=grok-stt&operation=speech-to-text)

## Webhook (Optional)

Instead of polling, you can receive a Webhook callback when the request reaches a terminal state. Provide a Webhook URL via header on the submit request.

Header

Required

Description

X-Webhook-URL

To enable

HTTPS URL to receive the Webhook callback.

X-Webhook-Mode

No

`terminal` (default, one callback on COMPLETED/ERROR) or `sync` (per-poll callbacks).

### Example: enable Webhook

```
curl -X POST 'https://gateway.pixazo.ai/grok-stt/v1/speech-to-text' \
  -H 'Content-Type: application/json' \
  -H 'Ocp-Apim-Subscription-Key: YOUR_SUBSCRIPTION_KEY' \
  -H 'X-Webhook-URL: https://your-server.com/webhook' \
  --data-raw '{"audio_url": "https://pub-582b7213209642b9b995c96c95a30381.r2.dev/doc-assets/audio/speech-17s.mp3"}'
```

### Callback Payload (success)

```
{
  "request_id": "grok-stt_019dxxxx-xxxx-xxxx-xxxx-xxxxxxxxxxxx",
  "status": "COMPLETED",
  "model_id": "grok-stt",
  "output": {
    "media_url": [
      "https://pub-582b7213209642b9b995c96c95a30381.r2.dev/v1/{request_id}/transcript.json"
    ],
    "media_type": "application/json"
  },
  "duration": 19.17,
  "created_at": "2026-08-01T09:14:16.102Z",
  "completed_at": "2026-08-01T09:14:22.870Z"
}
```

### Failure callback shape

```
{
  "request_id": "grok-stt_019dxxxx-xxxx-xxxx-xxxx-xxxxxxxxxxxx",
  "status": "ERROR",
  "model_id": "grok-stt",
  "error": "Description of the failure"
}
```

#### Delivery semantics

-   **terminal** mode: one Webhook callback when the request is COMPLETED or ERROR.
-   **sync** mode: a Webhook callback on each status change.
-   Callbacks are idempotent on `request_id` — de-duplicate on it.
-   Respond `200` within a few seconds; the Webhook endpoint must be HTTPS.

## Request Parameters

Parameter

Required

Type

Default

Allowed values / range

Description

`audio_url`

Yes

string

—

a publicly reachable http(s) url

The recording to transcribe. We fetch it server-side, so it must be reachable from the internet — a signed url is fine, a private one is not. `audio` is accepted as an alias.

`language`

No

string

—

ISO 639-1, e.g. `en`

Omit this to let the model detect the language — that is the default. `auto` means the same. Set it to force one of the recording.

`prompt`

No

string

—

up to 2,000 characters

Bias the transcription toward expected wording — names, jargon, spellings.

### Voices

Returns `text`, the detected `language` and the clip `duration`.

## Example Request

```
{
  "audio_url": "https://pub-582b7213209642b9b995c96c95a30381.r2.dev/doc-assets/audio/speech-17s.mp3"
}
```

## Example Response

```
{
  "request_id": "grok-stt_019dxxxx-xxxx-xxxx-xxxx-xxxxxxxxxxxx",
  "status": "QUEUED",
  "polling_url": "https://gateway.pixazo.ai/v2/requests/status/grok-stt_019dxxxx-xxxx-xxxx-xxxx-xxxxxxxxxxxx"
}
```

## Request Headers

Header

Required

Description

Content-Type

Yes

`application/json`

Ocp-Apim-Subscription-Key

Yes

Your API subscription key.

X-Webhook-URL

No

Enable Webhook callbacks (see Webhook section).

## Response Handling

Status Code

Meaning

202

Accepted — request queued; returns `request_id` and `polling_url`.

400

Bad request — a missing or out-of-range parameter. The message names the field.

401

Unauthorized — missing or invalid subscription key.

402

Insufficient balance.

429

Too many requests.

500

Internal server error.

## Retrieving Results

Poll the status endpoint with the `request_id` from the submit response until `status` is `COMPLETED` (or `ERROR`), then download `output.media_url`.

```
curl 'https://gateway.pixazo.ai/v2/requests/status/grok-stt_019dxxxx-xxxx-xxxx-xxxx-xxxxxxxxxxxx' \
  -H 'Ocp-Apim-Subscription-Key: YOUR_SUBSCRIPTION_KEY'
```

### Completed response

```
{
  "request_id": "grok-stt_019dxxxx-xxxx-xxxx-xxxx-xxxxxxxxxxxx",
  "status": "COMPLETED",
  "model_id": "grok-stt",
  "output": {
    "media_url": [
      "https://pub-582b7213209642b9b995c96c95a30381.r2.dev/v1/{request_id}/transcript.json"
    ],
    "media_type": "application/json"
  },
  "duration": 19.17,
  "created_at": "2026-08-01T09:14:16.102Z",
  "completed_at": "2026-08-01T09:14:22.870Z"
}
```

## Response Fields

Field

Type

Description

request\_id

string

Unique request identifier.

status

string

QUEUED, PROCESSING, COMPLETED or ERROR.

model\_id

string

The model that handled the request.

output.media\_url

array

URL of the generated audio file.

output.media\_type

string

MIME type of the audio.

created\_at

string

Request creation timestamp.

completed\_at

string

Completion timestamp.

error

string

Error message when `status` is ERROR.

## Status Values & Flow

`QUEUED` → `PROCESSING` → `COMPLETED` (success) or `ERROR` (failure).

### Pricing

Billed at **$0.001667 per minute of generated audio**, rounded up to the next whole minute. You are charged for the audio produced, not the text you submit.

Audio produced

Billed minutes

Cost

A 10-second clip

1

$0.001667

A 45-second clip

1

$0.001667

A 3-minute narration

3

$0.005001

A 10-minute narration

10

$0.01667

Failed requests are not billed.
