---
type: AI Model
id: whisper
title: OpenAI Speech to Text API
provider: OpenAI
description: "OpenAI's speech recognition on the gateway: Whisper in three sizes, transcribing 99 languages with word timings and subtitles. Billed per minute of audio."
resource: https://www.pixazo.ai/models/whisper
docs_url: https://www.pixazo.ai/models/whisper
latest_version: Large v3 Turbo
tags:
  - speech-to-text
  - openai
variants:
  - id: whisper-large-v3-turbo
    name: Whisper Large v3 Turbo
    version: Large v3 Turbo
    capabilities:
      - Speech to Text
  - id: whisper-tiny-en
    name: Whisper Tiny English
    version: Tiny (EN)
    capabilities:
      - Speech to Text
  - id: whisper
    name: Whisper
    version: v1
    capabilities:
      - Speech to Text
timestamp: 2026-08-25T08:44:38.582Z
---

# OpenAI Speech to Text API

> Provider: **OpenAI**
> Source: https://www.pixazo.ai/models/whisper

OpenAI's speech recognition on the gateway: Whisper in three sizes, transcribing 99 languages with word timings and subtitles. Billed per minute of audio.

## Whisper Large v3 Turbo

### Speech to Text

## Whisper Large v3 Turbo API Documentation

OpenAI Whisper Large v3 Turbo, the fastest of the large Whisper models. Returns the transcript plus per-segment timings, word counts and a ready-made WebVTT subtitle track. Asynchronous: submit returns a `request_id`; poll the status endpoint until the request is `COMPLETED`, then download the audio.

```
POST https://gateway.pixazo.ai/whisper-large-v3-turbo/v1/speech-to-text
```

## Authentication

All requests require an API key passed via header.

Header

Type

Required

Description

Ocp-Apim-Subscription-Key

string

Yes

Your API subscription key

## Speech to Text - Whisper Large v3 Turbo

## Request Code

HTTP Python JavaScript cURL

```
POST https://gateway.pixazo.ai/whisper-large-v3-turbo/v1/speech-to-text
Content-Type: application/json
Ocp-Apim-Subscription-Key: YOUR_SUBSCRIPTION_KEY

{
  "audio_url": "https://pub-582b7213209642b9b995c96c95a30381.r2.dev/doc-assets/audio/speech-17s.mp3"
}
```

```
import requests

url = "https://gateway.pixazo.ai/whisper-large-v3-turbo/v1/speech-to-text"
headers = {
    "Content-Type": "application/json",
    "Ocp-Apim-Subscription-Key": "YOUR_SUBSCRIPTION_KEY"
}
data = {
  "audio_url": "https://pub-582b7213209642b9b995c96c95a30381.r2.dev/doc-assets/audio/speech-17s.mp3"
}

resp = requests.post(url, json=data, headers=headers)
print(resp.json())
```

```
const res = await fetch("https://gateway.pixazo.ai/whisper-large-v3-turbo/v1/speech-to-text", {
  method: "POST",
  headers: {
    "Content-Type": "application/json",
    "Ocp-Apim-Subscription-Key": "YOUR_SUBSCRIPTION_KEY"
  },
  body: JSON.stringify({
  "audio_url": "https://pub-582b7213209642b9b995c96c95a30381.r2.dev/doc-assets/audio/speech-17s.mp3"
})
});
console.log(await res.json());
```

```
curl -X POST 'https://gateway.pixazo.ai/whisper-large-v3-turbo/v1/speech-to-text' \
  -H 'Content-Type: application/json' \
  -H 'Ocp-Apim-Subscription-Key: YOUR_SUBSCRIPTION_KEY' \
  --data-raw '{"audio_url": "https://pub-582b7213209642b9b995c96c95a30381.r2.dev/doc-assets/audio/speech-17s.mp3"}'
```

## Output

```
{
  "request_id": "whisper-large-v3-turbo_019dxxxx-xxxx-xxxx-xxxx-xxxxxxxxxxxx",
  "status": "QUEUED",
  "polling_url": "https://gateway.pixazo.ai/v2/requests/status/whisper-large-v3-turbo_019dxxxx-xxxx-xxxx-xxxx-xxxxxxxxxxxx"
}
```

[Try Now](https://api.pixazo.ai/api-details#api=whisper-large-v3-turbo&operation=speech-to-text)

## Webhook (Optional)

Instead of polling, you can receive a Webhook callback when the request reaches a terminal state. Provide a Webhook URL via header on the submit request.

Header

Required

Description

X-Webhook-URL

To enable

HTTPS URL to receive the Webhook callback.

X-Webhook-Mode

No

`terminal` (default, one callback on COMPLETED/ERROR) or `sync` (per-poll callbacks).

### Example: enable Webhook

```
curl -X POST 'https://gateway.pixazo.ai/whisper-large-v3-turbo/v1/speech-to-text' \
  -H 'Content-Type: application/json' \
  -H 'Ocp-Apim-Subscription-Key: YOUR_SUBSCRIPTION_KEY' \
  -H 'X-Webhook-URL: https://your-server.com/webhook' \
  --data-raw '{"audio_url": "https://pub-582b7213209642b9b995c96c95a30381.r2.dev/doc-assets/audio/speech-17s.mp3"}'
```

### Callback Payload (success)

```
{
  "request_id": "whisper-large-v3-turbo_019dxxxx-xxxx-xxxx-xxxx-xxxxxxxxxxxx",
  "status": "COMPLETED",
  "model_id": "whisper-large-v3-turbo",
  "output": {
    "media_url": [
      "https://pub-582b7213209642b9b995c96c95a30381.r2.dev/v1/{request_id}/transcript.json"
    ],
    "media_type": "application/json"
  },
  "duration": 19.17,
  "created_at": "2026-08-01T09:14:16.102Z",
  "completed_at": "2026-08-01T09:14:22.870Z"
}
```

### Failure callback shape

```
{
  "request_id": "whisper-large-v3-turbo_019dxxxx-xxxx-xxxx-xxxx-xxxxxxxxxxxx",
  "status": "ERROR",
  "model_id": "whisper-large-v3-turbo",
  "error": "Description of the failure"
}
```

#### Delivery semantics

-   **terminal** mode: one Webhook callback when the request is COMPLETED or ERROR.
-   **sync** mode: a Webhook callback on each status change.
-   Callbacks are idempotent on `request_id` — de-duplicate on it.
-   Respond `200` within a few seconds; the Webhook endpoint must be HTTPS.

## Request Parameters

Parameter

Required

Type

Default

Allowed values / range

Description

`audio_url`

Yes

string

—

a publicly reachable http(s) url

The recording to transcribe. We fetch it server-side, so it must be reachable from the internet — a signed url is fine, a private one is not. `audio` is accepted as an alias.

`task`

No

string

`transcribe`

`transcribe`, `translate`

`transcribe` keeps the spoken language; `translate` renders English regardless of the source.

`language`

No

string

—

ISO 639-1, e.g. `en`

Omit this to let the model detect the language — that is the default. `auto` means the same. Set it to force one. Detection is reported back as `transcription_info.language` with a confidence.

`vad_filter`

No

boolean

`false`

`true`, `false`

Drop silence before transcribing. Useful on long recordings with gaps.

`initial_prompt`

No

string

—

up to 2,000 characters

Bias the model toward expected wording — names, jargon, spellings.

### Voices

Transcribes 99 languages and can translate any of them into English in the same call.

## Example Request

```
{
  "audio_url": "https://pub-582b7213209642b9b995c96c95a30381.r2.dev/doc-assets/audio/speech-17s.mp3"
}
```

## Example Response

```
{
  "request_id": "whisper-large-v3-turbo_019dxxxx-xxxx-xxxx-xxxx-xxxxxxxxxxxx",
  "status": "QUEUED",
  "polling_url": "https://gateway.pixazo.ai/v2/requests/status/whisper-large-v3-turbo_019dxxxx-xxxx-xxxx-xxxx-xxxxxxxxxxxx"
}
```

## Request Headers

Header

Required

Description

Content-Type

Yes

`application/json`

Ocp-Apim-Subscription-Key

Yes

Your API subscription key.

X-Webhook-URL

No

Enable Webhook callbacks (see Webhook section).

## Response Handling

Status Code

Meaning

202

Accepted — request queued; returns `request_id` and `polling_url`.

400

Bad request — a missing or out-of-range parameter. The message names the field.

401

Unauthorized — missing or invalid subscription key.

402

Insufficient balance.

429

Too many requests.

500

Internal server error.

## Retrieving Results

Poll the status endpoint with the `request_id` from the submit response until `status` is `COMPLETED` (or `ERROR`), then download `output.media_url`.

```
curl 'https://gateway.pixazo.ai/v2/requests/status/whisper-large-v3-turbo_019dxxxx-xxxx-xxxx-xxxx-xxxxxxxxxxxx' \
  -H 'Ocp-Apim-Subscription-Key: YOUR_SUBSCRIPTION_KEY'
```

### Completed response

```
{
  "request_id": "whisper-large-v3-turbo_019dxxxx-xxxx-xxxx-xxxx-xxxxxxxxxxxx",
  "status": "COMPLETED",
  "model_id": "whisper-large-v3-turbo",
  "output": {
    "media_url": [
      "https://pub-582b7213209642b9b995c96c95a30381.r2.dev/v1/{request_id}/transcript.json"
    ],
    "media_type": "application/json"
  },
  "duration": 19.17,
  "created_at": "2026-08-01T09:14:16.102Z",
  "completed_at": "2026-08-01T09:14:22.870Z"
}
```

## Response Fields

Field

Type

Description

request\_id

string

Unique request identifier.

status

string

QUEUED, PROCESSING, COMPLETED or ERROR.

model\_id

string

The model that handled the request.

output.media\_url

array

URL of the generated audio file.

output.media\_type

string

MIME type of the audio.

created\_at

string

Request creation timestamp.

completed\_at

string

Completion timestamp.

error

string

Error message when `status` is ERROR.

## Status Values & Flow

`QUEUED` → `PROCESSING` → `COMPLETED` (success) or `ERROR` (failure).

### Pricing

Billed at **$0.00051 per minute of generated audio**, rounded up to the next whole minute. You are charged for the audio produced, not the text you submit.

Audio produced

Billed minutes

Cost

A 10-second clip

1

$0.00051

A 45-second clip

1

$0.00051

A 3-minute narration

3

$0.00153

A 10-minute narration

10

$0.0051

Failed requests are not billed.

## Whisper Tiny English

### Speech to Text

## Whisper Tiny English API Documentation

The smallest Whisper model, English only. Noticeably faster and cheaper than the large models, at some cost in accuracy on proper nouns — a good fit for bulk transcription where a rough transcript is enough. Asynchronous: submit returns a `request_id`; poll the status endpoint until the request is `COMPLETED`, then download the audio.

```
POST https://gateway.pixazo.ai/whisper-tiny-en/v1/speech-to-text
```

## Authentication

All requests require an API key passed via header.

Header

Type

Required

Description

Ocp-Apim-Subscription-Key

string

Yes

Your API subscription key

## Speech to Text - Whisper Tiny English

## Request Code

HTTP Python JavaScript cURL

```
POST https://gateway.pixazo.ai/whisper-tiny-en/v1/speech-to-text
Content-Type: application/json
Ocp-Apim-Subscription-Key: YOUR_SUBSCRIPTION_KEY

{
  "audio_url": "https://pub-582b7213209642b9b995c96c95a30381.r2.dev/doc-assets/audio/speech-17s.mp3"
}
```

```
import requests

url = "https://gateway.pixazo.ai/whisper-tiny-en/v1/speech-to-text"
headers = {
    "Content-Type": "application/json",
    "Ocp-Apim-Subscription-Key": "YOUR_SUBSCRIPTION_KEY"
}
data = {
  "audio_url": "https://pub-582b7213209642b9b995c96c95a30381.r2.dev/doc-assets/audio/speech-17s.mp3"
}

resp = requests.post(url, json=data, headers=headers)
print(resp.json())
```

```
const res = await fetch("https://gateway.pixazo.ai/whisper-tiny-en/v1/speech-to-text", {
  method: "POST",
  headers: {
    "Content-Type": "application/json",
    "Ocp-Apim-Subscription-Key": "YOUR_SUBSCRIPTION_KEY"
  },
  body: JSON.stringify({
  "audio_url": "https://pub-582b7213209642b9b995c96c95a30381.r2.dev/doc-assets/audio/speech-17s.mp3"
})
});
console.log(await res.json());
```

```
curl -X POST 'https://gateway.pixazo.ai/whisper-tiny-en/v1/speech-to-text' \
  -H 'Content-Type: application/json' \
  -H 'Ocp-Apim-Subscription-Key: YOUR_SUBSCRIPTION_KEY' \
  --data-raw '{"audio_url": "https://pub-582b7213209642b9b995c96c95a30381.r2.dev/doc-assets/audio/speech-17s.mp3"}'
```

## Output

```
{
  "request_id": "whisper-tiny-en_019dxxxx-xxxx-xxxx-xxxx-xxxxxxxxxxxx",
  "status": "QUEUED",
  "polling_url": "https://gateway.pixazo.ai/v2/requests/status/whisper-tiny-en_019dxxxx-xxxx-xxxx-xxxx-xxxxxxxxxxxx"
}
```

[Try Now](https://api.pixazo.ai/api-details#api=whisper-tiny-en&operation=speech-to-text)

## Webhook (Optional)

Instead of polling, you can receive a Webhook callback when the request reaches a terminal state. Provide a Webhook URL via header on the submit request.

Header

Required

Description

X-Webhook-URL

To enable

HTTPS URL to receive the Webhook callback.

X-Webhook-Mode

No

`terminal` (default, one callback on COMPLETED/ERROR) or `sync` (per-poll callbacks).

### Example: enable Webhook

```
curl -X POST 'https://gateway.pixazo.ai/whisper-tiny-en/v1/speech-to-text' \
  -H 'Content-Type: application/json' \
  -H 'Ocp-Apim-Subscription-Key: YOUR_SUBSCRIPTION_KEY' \
  -H 'X-Webhook-URL: https://your-server.com/webhook' \
  --data-raw '{"audio_url": "https://pub-582b7213209642b9b995c96c95a30381.r2.dev/doc-assets/audio/speech-17s.mp3"}'
```

### Callback Payload (success)

```
{
  "request_id": "whisper-tiny-en_019dxxxx-xxxx-xxxx-xxxx-xxxxxxxxxxxx",
  "status": "COMPLETED",
  "model_id": "whisper-tiny-en",
  "output": {
    "media_url": [
      "https://pub-582b7213209642b9b995c96c95a30381.r2.dev/v1/{request_id}/transcript.json"
    ],
    "media_type": "application/json"
  },
  "duration": 19.17,
  "created_at": "2026-08-01T09:14:16.102Z",
  "completed_at": "2026-08-01T09:14:22.870Z"
}
```

### Failure callback shape

```
{
  "request_id": "whisper-tiny-en_019dxxxx-xxxx-xxxx-xxxx-xxxxxxxxxxxx",
  "status": "ERROR",
  "model_id": "whisper-tiny-en",
  "error": "Description of the failure"
}
```

#### Delivery semantics

-   **terminal** mode: one Webhook callback when the request is COMPLETED or ERROR.
-   **sync** mode: a Webhook callback on each status change.
-   Callbacks are idempotent on `request_id` — de-duplicate on it.
-   Respond `200` within a few seconds; the Webhook endpoint must be HTTPS.

## Request Parameters

Parameter

Required

Type

Default

Allowed values / range

Description

`audio_url`

Yes

string

—

a publicly reachable http(s) url

The recording to transcribe. We fetch it server-side, so it must be reachable from the internet — a signed url is fine, a private one is not. `audio` is accepted as an alias.

`task`

No

string

`transcribe`

`transcribe`, `translate`

`transcribe` keeps the spoken language; `translate` renders English regardless of the source.

`language`

No

string

—

ISO 639-1, e.g. `en`

Omit this to let the model detect the language — that is the default. `auto` means the same. Set it to force one. Detection is reported back as `transcription_info.language` with a confidence.

`vad_filter`

No

boolean

`false`

`true`, `false`

Drop silence before transcribing. Useful on long recordings with gaps.

`initial_prompt`

No

string

—

up to 2,000 characters

Bias the model toward expected wording — names, jargon, spellings.

### Voices

English only. `language` and `task=translate` have no effect on this model.

## Example Request

```
{
  "audio_url": "https://pub-582b7213209642b9b995c96c95a30381.r2.dev/doc-assets/audio/speech-17s.mp3"
}
```

## Example Response

```
{
  "request_id": "whisper-tiny-en_019dxxxx-xxxx-xxxx-xxxx-xxxxxxxxxxxx",
  "status": "QUEUED",
  "polling_url": "https://gateway.pixazo.ai/v2/requests/status/whisper-tiny-en_019dxxxx-xxxx-xxxx-xxxx-xxxxxxxxxxxx"
}
```

## Request Headers

Header

Required

Description

Content-Type

Yes

`application/json`

Ocp-Apim-Subscription-Key

Yes

Your API subscription key.

X-Webhook-URL

No

Enable Webhook callbacks (see Webhook section).

## Response Handling

Status Code

Meaning

202

Accepted — request queued; returns `request_id` and `polling_url`.

400

Bad request — a missing or out-of-range parameter. The message names the field.

401

Unauthorized — missing or invalid subscription key.

402

Insufficient balance.

429

Too many requests.

500

Internal server error.

## Retrieving Results

Poll the status endpoint with the `request_id` from the submit response until `status` is `COMPLETED` (or `ERROR`), then download `output.media_url`.

```
curl 'https://gateway.pixazo.ai/v2/requests/status/whisper-tiny-en_019dxxxx-xxxx-xxxx-xxxx-xxxxxxxxxxxx' \
  -H 'Ocp-Apim-Subscription-Key: YOUR_SUBSCRIPTION_KEY'
```

### Completed response

```
{
  "request_id": "whisper-tiny-en_019dxxxx-xxxx-xxxx-xxxx-xxxxxxxxxxxx",
  "status": "COMPLETED",
  "model_id": "whisper-tiny-en",
  "output": {
    "media_url": [
      "https://pub-582b7213209642b9b995c96c95a30381.r2.dev/v1/{request_id}/transcript.json"
    ],
    "media_type": "application/json"
  },
  "duration": 19.17,
  "created_at": "2026-08-01T09:14:16.102Z",
  "completed_at": "2026-08-01T09:14:22.870Z"
}
```

## Response Fields

Field

Type

Description

request\_id

string

Unique request identifier.

status

string

QUEUED, PROCESSING, COMPLETED or ERROR.

model\_id

string

The model that handled the request.

output.media\_url

array

URL of the generated audio file.

output.media\_type

string

MIME type of the audio.

created\_at

string

Request creation timestamp.

completed\_at

string

Completion timestamp.

error

string

Error message when `status` is ERROR.

## Status Values & Flow

`QUEUED` → `PROCESSING` → `COMPLETED` (success) or `ERROR` (failure).

### Pricing

Billed at **$0.0005 per minute of generated audio**, rounded up to the next whole minute. You are charged for the audio produced, not the text you submit.

Audio produced

Billed minutes

Cost

A 10-second clip

1

$0.0005

A 45-second clip

1

$0.0005

A 3-minute narration

3

$0.0015

A 10-minute narration

10

$0.005

Failed requests are not billed.

## Whisper

### Speech to Text

## Whisper API Documentation

The original OpenAI Whisper model. Transcribes 99 languages and can translate any of them into English, returning the transcript with word-level timings and a WebVTT subtitle track. Asynchronous: submit returns a `request_id`; poll the status endpoint until the request is `COMPLETED`, then download the audio.

```
POST https://gateway.pixazo.ai/whisper/v1/speech-to-text
```

## Authentication

All requests require an API key passed via header.

Header

Type

Required

Description

Ocp-Apim-Subscription-Key

string

Yes

Your API subscription key

## Speech to Text - Whisper

## Request Code

HTTP Python JavaScript cURL

```
POST https://gateway.pixazo.ai/whisper/v1/speech-to-text
Content-Type: application/json
Ocp-Apim-Subscription-Key: YOUR_SUBSCRIPTION_KEY

{
  "audio_url": "https://pub-582b7213209642b9b995c96c95a30381.r2.dev/doc-assets/audio/speech-17s.mp3"
}
```

```
import requests

url = "https://gateway.pixazo.ai/whisper/v1/speech-to-text"
headers = {
    "Content-Type": "application/json",
    "Ocp-Apim-Subscription-Key": "YOUR_SUBSCRIPTION_KEY"
}
data = {
  "audio_url": "https://pub-582b7213209642b9b995c96c95a30381.r2.dev/doc-assets/audio/speech-17s.mp3"
}

resp = requests.post(url, json=data, headers=headers)
print(resp.json())
```

```
const res = await fetch("https://gateway.pixazo.ai/whisper/v1/speech-to-text", {
  method: "POST",
  headers: {
    "Content-Type": "application/json",
    "Ocp-Apim-Subscription-Key": "YOUR_SUBSCRIPTION_KEY"
  },
  body: JSON.stringify({
  "audio_url": "https://pub-582b7213209642b9b995c96c95a30381.r2.dev/doc-assets/audio/speech-17s.mp3"
})
});
console.log(await res.json());
```

```
curl -X POST 'https://gateway.pixazo.ai/whisper/v1/speech-to-text' \
  -H 'Content-Type: application/json' \
  -H 'Ocp-Apim-Subscription-Key: YOUR_SUBSCRIPTION_KEY' \
  --data-raw '{"audio_url": "https://pub-582b7213209642b9b995c96c95a30381.r2.dev/doc-assets/audio/speech-17s.mp3"}'
```

## Output

```
{
  "request_id": "whisper_019dxxxx-xxxx-xxxx-xxxx-xxxxxxxxxxxx",
  "status": "QUEUED",
  "polling_url": "https://gateway.pixazo.ai/v2/requests/status/whisper_019dxxxx-xxxx-xxxx-xxxx-xxxxxxxxxxxx"
}
```

[Try Now](https://api.pixazo.ai/api-details#api=whisper&operation=speech-to-text)

## Webhook (Optional)

Instead of polling, you can receive a Webhook callback when the request reaches a terminal state. Provide a Webhook URL via header on the submit request.

Header

Required

Description

X-Webhook-URL

To enable

HTTPS URL to receive the Webhook callback.

X-Webhook-Mode

No

`terminal` (default, one callback on COMPLETED/ERROR) or `sync` (per-poll callbacks).

### Example: enable Webhook

```
curl -X POST 'https://gateway.pixazo.ai/whisper/v1/speech-to-text' \
  -H 'Content-Type: application/json' \
  -H 'Ocp-Apim-Subscription-Key: YOUR_SUBSCRIPTION_KEY' \
  -H 'X-Webhook-URL: https://your-server.com/webhook' \
  --data-raw '{"audio_url": "https://pub-582b7213209642b9b995c96c95a30381.r2.dev/doc-assets/audio/speech-17s.mp3"}'
```

### Callback Payload (success)

```
{
  "request_id": "whisper_019dxxxx-xxxx-xxxx-xxxx-xxxxxxxxxxxx",
  "status": "COMPLETED",
  "model_id": "whisper",
  "output": {
    "media_url": [
      "https://pub-582b7213209642b9b995c96c95a30381.r2.dev/v1/{request_id}/transcript.json"
    ],
    "media_type": "application/json"
  },
  "duration": 19.17,
  "created_at": "2026-08-01T09:14:16.102Z",
  "completed_at": "2026-08-01T09:14:22.870Z"
}
```

### Failure callback shape

```
{
  "request_id": "whisper_019dxxxx-xxxx-xxxx-xxxx-xxxxxxxxxxxx",
  "status": "ERROR",
  "model_id": "whisper",
  "error": "Description of the failure"
}
```

#### Delivery semantics

-   **terminal** mode: one Webhook callback when the request is COMPLETED or ERROR.
-   **sync** mode: a Webhook callback on each status change.
-   Callbacks are idempotent on `request_id` — de-duplicate on it.
-   Respond `200` within a few seconds; the Webhook endpoint must be HTTPS.

## Request Parameters

Parameter

Required

Type

Default

Allowed values / range

Description

`audio_url`

Yes

string

—

a publicly reachable http(s) url

The recording to transcribe. We fetch it server-side, so it must be reachable from the internet — a signed url is fine, a private one is not. `audio` is accepted as an alias.

`task`

No

string

`transcribe`

`transcribe`, `translate`

`transcribe` keeps the spoken language; `translate` renders English regardless of the source.

`language`

No

string

—

ISO 639-1, e.g. `en`

Omit this to let the model detect the language — that is the default. `auto` means the same. Set it to force one. Detection is reported back as `transcription_info.language` with a confidence.

`vad_filter`

No

boolean

`false`

`true`, `false`

Drop silence before transcribing. Useful on long recordings with gaps.

`initial_prompt`

No

string

—

up to 2,000 characters

Bias the model toward expected wording — names, jargon, spellings.

### Voices

Note it can hallucinate a short phrase on recordings that contain no speech — a known Whisper behaviour, not a fault of the API.

## Example Request

```
{
  "audio_url": "https://pub-582b7213209642b9b995c96c95a30381.r2.dev/doc-assets/audio/speech-17s.mp3"
}
```

## Example Response

```
{
  "request_id": "whisper_019dxxxx-xxxx-xxxx-xxxx-xxxxxxxxxxxx",
  "status": "QUEUED",
  "polling_url": "https://gateway.pixazo.ai/v2/requests/status/whisper_019dxxxx-xxxx-xxxx-xxxx-xxxxxxxxxxxx"
}
```

## Request Headers

Header

Required

Description

Content-Type

Yes

`application/json`

Ocp-Apim-Subscription-Key

Yes

Your API subscription key.

X-Webhook-URL

No

Enable Webhook callbacks (see Webhook section).

## Response Handling

Status Code

Meaning

202

Accepted — request queued; returns `request_id` and `polling_url`.

400

Bad request — a missing or out-of-range parameter. The message names the field.

401

Unauthorized — missing or invalid subscription key.

402

Insufficient balance.

429

Too many requests.

500

Internal server error.

## Retrieving Results

Poll the status endpoint with the `request_id` from the submit response until `status` is `COMPLETED` (or `ERROR`), then download `output.media_url`.

```
curl 'https://gateway.pixazo.ai/v2/requests/status/whisper_019dxxxx-xxxx-xxxx-xxxx-xxxxxxxxxxxx' \
  -H 'Ocp-Apim-Subscription-Key: YOUR_SUBSCRIPTION_KEY'
```

### Completed response

```
{
  "request_id": "whisper_019dxxxx-xxxx-xxxx-xxxx-xxxxxxxxxxxx",
  "status": "COMPLETED",
  "model_id": "whisper",
  "output": {
    "media_url": [
      "https://pub-582b7213209642b9b995c96c95a30381.r2.dev/v1/{request_id}/transcript.json"
    ],
    "media_type": "application/json"
  },
  "duration": 19.17,
  "created_at": "2026-08-01T09:14:16.102Z",
  "completed_at": "2026-08-01T09:14:22.870Z"
}
```

## Response Fields

Field

Type

Description

request\_id

string

Unique request identifier.

status

string

QUEUED, PROCESSING, COMPLETED or ERROR.

model\_id

string

The model that handled the request.

output.media\_url

array

URL of the generated audio file.

output.media\_type

string

MIME type of the audio.

created\_at

string

Request creation timestamp.

completed\_at

string

Completion timestamp.

error

string

Error message when `status` is ERROR.

## Status Values & Flow

`QUEUED` → `PROCESSING` → `COMPLETED` (success) or `ERROR` (failure).

### Pricing

Billed at **$0.0005 per minute of generated audio**, rounded up to the next whole minute. You are charged for the audio produced, not the text you submit.

Audio produced

Billed minutes

Cost

A 10-second clip

1

$0.0005

A 45-second clip

1

$0.0005

A 3-minute narration

3

$0.0015

A 10-minute narration

10

$0.005

Failed requests are not billed.
