---
type: AI Model
id: gemini-voice
title: Gemini Voice API
provider: Google
description: "Google's Gemini voice models on one endpoint pair: Gemini 3.5 Transcribe turns recorded speech into text with speaker labels and word-level timestamps across 85+ locales, and Gemini 3.1 Flash TTS turns text into speech with 30 voices, natural-language delivery control and two-speaker dialogue. Both are billed per minute of audio."
resource: https://www.pixazo.ai/models/gemini-voice
docs_url: https://www.pixazo.ai/models/gemini-voice
trending: true
latest_version: v3.5
tags:
  - trends
  - speech-to-text
  - text-to-speech
  - google
variants:
  - id: gemini-3-5-transcribe-v1
    name: Gemini 3.5 Transcribe
    version: 3.5
    capabilities:
      - Speech to Text
  - id: gemini-3-1-flash-tts-v1
    name: Gemini 3.1 Flash TTS
    version: 3.1 Flash
    capabilities:
      - Text to Speech
timestamp: 2026-08-30T12:03:01.004Z
---

# Gemini Voice API

> Provider: **Google**
> Source: https://www.pixazo.ai/models/gemini-voice

Google's Gemini voice models on one endpoint pair: Gemini 3.5 Transcribe turns recorded speech into text with speaker labels and word-level timestamps across 85+ locales, and Gemini 3.1 Flash TTS turns text into speech with 30 voices, natural-language delivery control and two-speaker dialogue. Both are billed per minute of audio.

## Gemini 3.5 Transcribe

### Speech to Text

## Gemini 3.5 Transcribe API Documentation

Google's Gemini 3.5 Transcribe converts pre-recorded speech to text, with automatic language identification across 85+ locales, speaker diarization, word-level timestamps and custom-vocabulary biasing. Accepts WAV, MP3, AIFF, AAC, OGG, FLAC, M4A, Opus, WebM and MPEG, up to 1 hour per request. Asynchronous: submit returns a `request_id`; poll the status endpoint until the request is `COMPLETED`, then fetch the JSON transcript from the returned url.

```
POST https://gateway.pixazo.ai/gemini-3-5-transcribe/v1/speech-to-text
```

## Authentication

All requests require an API key passed via header.

Header

Type

Required

Description

Ocp-Apim-Subscription-Key

string

Yes

Your API subscription key

## Speech to Text - Gemini 3.5 Transcribe

## Request Code

HTTP Python JavaScript cURL

```
POST https://gateway.pixazo.ai/gemini-3-5-transcribe/v1/speech-to-text
Content-Type: application/json
Ocp-Apim-Subscription-Key: YOUR_SUBSCRIPTION_KEY

{
  "audio_url": "https://pub-582b7213209642b9b995c96c95a30381.r2.dev/doc-assets/audio/speech-17s.mp3"
}
```

```
import requests

url = "https://gateway.pixazo.ai/gemini-3-5-transcribe/v1/speech-to-text"
headers = {
    "Content-Type": "application/json",
    "Ocp-Apim-Subscription-Key": "YOUR_SUBSCRIPTION_KEY"
}
data = {
  "audio_url": "https://pub-582b7213209642b9b995c96c95a30381.r2.dev/doc-assets/audio/speech-17s.mp3"
}

resp = requests.post(url, json=data, headers=headers)
print(resp.json())
```

```
const res = await fetch("https://gateway.pixazo.ai/gemini-3-5-transcribe/v1/speech-to-text", {
  method: "POST",
  headers: {
    "Content-Type": "application/json",
    "Ocp-Apim-Subscription-Key": "YOUR_SUBSCRIPTION_KEY"
  },
  body: JSON.stringify({
  "audio_url": "https://pub-582b7213209642b9b995c96c95a30381.r2.dev/doc-assets/audio/speech-17s.mp3"
})
});
console.log(await res.json());
```

```
curl -X POST 'https://gateway.pixazo.ai/gemini-3-5-transcribe/v1/speech-to-text' \
  -H 'Content-Type: application/json' \
  -H 'Ocp-Apim-Subscription-Key: YOUR_SUBSCRIPTION_KEY' \
  --data-raw '{"audio_url": "https://pub-582b7213209642b9b995c96c95a30381.r2.dev/doc-assets/audio/speech-17s.mp3"}'
```

## Output

```
{
  "request_id": "gemini-3-5-transcribe_01a0xxxx-xxxx-xxxx-xxxx-xxxxxxxxxxxx",
  "status": "QUEUED",
  "polling_url": "https://gateway.pixazo.ai/v2/requests/status/gemini-3-5-transcribe_01a0xxxx-xxxx-xxxx-xxxx-xxxxxxxxxxxx"
}
```

[Try Now](https://api.pixazo.ai/api-details#api=gemini-3-5-transcribe&operation=speech-to-text)

## Webhook (Optional)

Instead of polling, you can receive a Webhook callback when the request reaches a terminal state. Provide a Webhook URL via header on the submit request.

Header

Required

Description

X-Webhook-URL

To enable

HTTPS URL to receive the Webhook callback.

X-Webhook-Mode

No

`terminal` (default, one callback on COMPLETED/ERROR) or `sync` (per-poll callbacks).

### Example: enable Webhook

```
curl -X POST 'https://gateway.pixazo.ai/gemini-3-5-transcribe/v1/speech-to-text' \
  -H 'Content-Type: application/json' \
  -H 'Ocp-Apim-Subscription-Key: YOUR_SUBSCRIPTION_KEY' \
  -H 'X-Webhook-URL: https://your-server.com/webhook' \
  --data-raw '{"audio_url": "https://pub-582b7213209642b9b995c96c95a30381.r2.dev/doc-assets/audio/speech-17s.mp3"}'
```

### Callback Payload (success)

```
{
  "request_id": "gemini-3-5-transcribe_01a0xxxx-xxxx-xxxx-xxxx-xxxxxxxxxxxx",
  "status": "COMPLETED",
  "model_id": "gemini-3-5-transcribe",
  "output": {
    "media_url": [
      "https://pub-582b7213209642b9b995c96c95a30381.r2.dev/v1/{request_id}/transcript.json"
    ],
    "media_type": "application/json"
  },
  "duration": 19.17,
  "created_at": "2026-08-01T09:14:16.102Z",
  "completed_at": "2026-08-01T09:14:22.870Z"
}
```

### Failure callback shape

```
{
  "request_id": "gemini-3-5-transcribe_01a0xxxx-xxxx-xxxx-xxxx-xxxxxxxxxxxx",
  "status": "ERROR",
  "model_id": "gemini-3-5-transcribe",
  "error": "Description of the failure"
}
```

#### Delivery semantics

-   **terminal** mode: one Webhook callback when the request is COMPLETED or ERROR.
-   **sync** mode: a Webhook callback on each status change.
-   Callbacks are idempotent on `request_id` — de-duplicate on it.
-   Respond `200` within a few seconds; the Webhook endpoint must be HTTPS.

## Request Parameters

Parameter

Required

Type

Default

Allowed values / range

Description

`audio_url`

Yes

string

—

a publicly reachable http(s) url

The recording to transcribe. We fetch it server-side, so it must be reachable from the internet — a signed url is fine, a private one is not. Accepts WAV, MP3, AIFF, AAC, OGG, FLAC, M4A, Opus, WebM and MPEG.

`language_codes`

No

array

— (auto-detected)

BCP-47 codes, e.g. `["en-US"]`

Leave this out and the model identifies the language itself across 85+ locales, including speakers switching language mid-recording. Set it to pin the transcript to a known language.

`mode`

No

string

`verbatim`

`verbatim`, `smart`

`verbatim` transcribes exactly what was said, keeping filler words, repetitions and false starts. `smart` cleans that up — removing disfluencies, applying self-corrections, and formatting lists, numbers, dates and paragraph breaks. `smart` cannot be combined with `diarization` or `timestamps`.

`diarization`

No

boolean

`false`

`true`, `false`

Label who is speaking, for up to 8 speakers. Each word in the transcript carries a `speaker` tag. Requires `mode` `verbatim`.

`timestamps`

No

boolean

`false`

`true`, `false`

Word-level start and end times, in seconds. Requires `mode` `verbatim`.

`custom_vocabulary`

No

array

—

up to 1,000 terms

Bias the transcript toward terms the model would otherwise misspell — product names, people, domain jargon.

### Limits

Audio up to **1 hour** per request. Turning on `diarization` or `timestamps` lowers that ceiling to **30 minutes**. Billing is per minute of _input_ audio, rounded up to the next whole minute — so a 20-second clip and a 55-second clip both bill one minute.

## Example Request

```
{
  "audio_url": "https://pub-582b7213209642b9b995c96c95a30381.r2.dev/doc-assets/audio/speech-17s.mp3"
}
```

## Example Response

```
{
  "request_id": "gemini-3-5-transcribe_01a0xxxx-xxxx-xxxx-xxxx-xxxxxxxxxxxx",
  "status": "QUEUED",
  "polling_url": "https://gateway.pixazo.ai/v2/requests/status/gemini-3-5-transcribe_01a0xxxx-xxxx-xxxx-xxxx-xxxxxxxxxxxx"
}
```

## Request Headers

Header

Required

Description

Content-Type

Yes

`application/json`

Ocp-Apim-Subscription-Key

Yes

Your API subscription key.

X-Webhook-URL

No

Enable Webhook callbacks (see Webhook section).

## Response Handling

Status Code

Meaning

202

Accepted — request queued; returns `request_id` and `polling_url`.

400

Bad request — a missing or out-of-range parameter. The message names the field.

401

Unauthorized — missing or invalid subscription key.

402

Insufficient balance.

429

Too many requests.

500

Internal server error.

## Retrieving Results

Poll the status endpoint with the `request_id` from the submit response until `status` is `COMPLETED` (or `ERROR`), then download `output.media_url`.

```
curl 'https://gateway.pixazo.ai/v2/requests/status/gemini-3-5-transcribe_01a0xxxx-xxxx-xxxx-xxxx-xxxxxxxxxxxx' \
  -H 'Ocp-Apim-Subscription-Key: YOUR_SUBSCRIPTION_KEY'
```

### Completed response

```
{
  "request_id": "gemini-3-5-transcribe_01a0xxxx-xxxx-xxxx-xxxx-xxxxxxxxxxxx",
  "status": "COMPLETED",
  "model_id": "gemini-3-5-transcribe",
  "output": {
    "media_url": [
      "https://pub-582b7213209642b9b995c96c95a30381.r2.dev/v1/{request_id}/transcript.json"
    ],
    "media_type": "application/json"
  },
  "duration": 19.17,
  "created_at": "2026-08-01T09:14:16.102Z",
  "completed_at": "2026-08-01T09:14:22.870Z"
}
```

## Response Fields

Field

Type

Description

request\_id

string

Unique request identifier.

status

string

QUEUED, PROCESSING, COMPLETED or ERROR.

model\_id

string

The model that handled the request.

output.media\_url

array

URL of the JSON transcript file. Fetch it to read `text`, and `segments` when you asked for timestamps or diarization.

output.media\_type

string

Always `application/json` — the result is a transcript document, not audio.

created\_at

string

Request creation timestamp.

completed\_at

string

Completion timestamp.

error

string

Error message when `status` is ERROR.

## Status Values & Flow

`QUEUED` → `PROCESSING` → `COMPLETED` (success) or `ERROR` (failure).

### Pricing

Billed at **$0.006 per minute of generated audio**, rounded up to the next whole minute. You are charged for the audio produced, not the text you submit.

Audio produced

Billed minutes

Cost

A 10-second clip

1

$0.006

A 45-second clip

1

$0.006

A 3-minute narration

3

$0.018

A 10-minute narration

10

$0.06

Failed requests are not billed.

## Gemini 3.1 Flash TTS

### Text to Speech

## Gemini 3.1 Flash TTS API Documentation

Google's Gemini 3.1 Flash TTS turns text into speech, with 30 prebuilt voices, natural-language control over delivery, optional two-speaker dialogue and 90+ languages detected automatically. Asynchronous: submit returns a `request_id`; poll the status endpoint until the request is `COMPLETED`, then download the WAV from the returned url.

```
POST https://gateway.pixazo.ai/gemini-3-1-flash-tts/v1/text-to-speech
```

## Authentication

All requests require an API key passed via header.

Header

Type

Required

Description

Ocp-Apim-Subscription-Key

string

Yes

Your API subscription key

## Text to Speech - Gemini 3.1 Flash TTS

## Request Code

HTTP Python JavaScript cURL

```
POST https://gateway.pixazo.ai/gemini-3-1-flash-tts/v1/text-to-speech
Content-Type: application/json
Ocp-Apim-Subscription-Key: YOUR_SUBSCRIPTION_KEY

{   "text": "Say warmly: Welcome back — your report is ready.",   "voice": "Kore" }
```

```
import requests

url = "https://gateway.pixazo.ai/gemini-3-1-flash-tts/v1/text-to-speech"
headers = {
    "Content-Type": "application/json",
    "Ocp-Apim-Subscription-Key": "YOUR_SUBSCRIPTION_KEY"
}
data = {   "text": "Say warmly: Welcome back — your report is ready.",   "voice": "Kore" }

resp = requests.post(url, json=data, headers=headers)
print(resp.json())
```

```
const res = await fetch("https://gateway.pixazo.ai/gemini-3-1-flash-tts/v1/text-to-speech", {
  method: "POST",
  headers: {
    "Content-Type": "application/json",
    "Ocp-Apim-Subscription-Key": "YOUR_SUBSCRIPTION_KEY"
  },
  body: JSON.stringify({   "text": "Say warmly: Welcome back — your report is ready.",   "voice": "Kore" })
});
console.log(await res.json());
```

```
curl -X POST 'https://gateway.pixazo.ai/gemini-3-1-flash-tts/v1/text-to-speech' \
  -H 'Content-Type: application/json' \
  -H 'Ocp-Apim-Subscription-Key: YOUR_SUBSCRIPTION_KEY' \
  --data-raw '{   "text": "Say warmly: Welcome back — your report is ready.",   "voice": "Kore" }'
```

## Output

```
{
  "request_id": "gemini-3-1-flash-tts_01a0xxxx-xxxx-xxxx-xxxx-xxxxxxxxxxxx",
  "status": "QUEUED",
  "polling_url": "https://gateway.pixazo.ai/v2/requests/status/gemini-3-1-flash-tts_01a0xxxx-xxxx-xxxx-xxxx-xxxxxxxxxxxx"
}
```

[Try Now](https://api.pixazo.ai/api-details#api=gemini-3-1-flash-tts&operation=text-to-speech)

## Webhook (Optional)

Instead of polling, you can receive a Webhook callback when the request reaches a terminal state. Provide a Webhook URL via header on the submit request.

Header

Required

Description

X-Webhook-URL

To enable

HTTPS URL to receive the Webhook callback.

X-Webhook-Mode

No

`terminal` (default, one callback on COMPLETED/ERROR) or `sync` (per-poll callbacks).

### Example: enable Webhook

```
curl -X POST 'https://gateway.pixazo.ai/gemini-3-1-flash-tts/v1/text-to-speech' \
  -H 'Content-Type: application/json' \
  -H 'Ocp-Apim-Subscription-Key: YOUR_SUBSCRIPTION_KEY' \
  -H 'X-Webhook-URL: https://your-server.com/webhook' \
  --data-raw '{   "text": "Say warmly: Welcome back — your report is ready.",   "voice": "Kore" }'
```

### Callback Payload (success)

```
{
  "request_id": "gemini-3-1-flash-tts_01a0xxxx-xxxx-xxxx-xxxx-xxxxxxxxxxxx",
  "status": "COMPLETED",
  "model_id": "gemini-3-1-flash-tts",
  "output": {
    "media_url": [
      "https://pub-582b7213209642b9b995c96c95a30381.r2.dev/v1/{request_id}/transcript.json"
    ],
    "media_type": "application/json"
  },
  "duration": 19.17,
  "created_at": "2026-08-01T09:14:16.102Z",
  "completed_at": "2026-08-01T09:14:22.870Z"
}
```

### Failure callback shape

```
{
  "request_id": "gemini-3-1-flash-tts_01a0xxxx-xxxx-xxxx-xxxx-xxxxxxxxxxxx",
  "status": "ERROR",
  "model_id": "gemini-3-1-flash-tts",
  "error": "Description of the failure"
}
```

#### Delivery semantics

-   **terminal** mode: one Webhook callback when the request is COMPLETED or ERROR.
-   **sync** mode: a Webhook callback on each status change.
-   Callbacks are idempotent on `request_id` — de-duplicate on it.
-   Respond `200` within a few seconds; the Webhook endpoint must be HTTPS.

## Request Parameters

Parameter

Required

Type

Default

Allowed values / range

Description

`text`

Yes

string

—

up to ~8,000 tokens

What to say. This is also where you direct _how_ it is said — the model follows natural-language direction, so "Say warmly and slowly: ..." works, as do inline tags like `[whispers]` and `[laughs]`.

`voice`

No

string

`Kore`

Zephyr, Puck, Charon, Kore, Fenrir, Leda, Orus, Aoede, Callirrhoe, Autonoe, Enceladus, Iapetus, Umbriel, Algieba, Despina, Erinome, Algenib, Rasalgethi, Laomedeia, Achernar, Alnilam, Schedar, Gacrux, Pulcherrima, Achird, Zubenelgenubi, Vindemiatrix, Sadachbia, Sadaltager, Sulafat

Which of the 30 prebuilt voices to speak in. Ignored when `speakers` is given.

`speakers`

No

array

—

at most 2 entries

Two-speaker dialogue. Each entry is `{"speaker": "Joe", "voice": "Kore"}`, where `speaker` matches a name used in your `text`. Supplying this replaces `voice`.

### Limits

A session has a 32k-token context. Speech quality can drift on outputs longer than a few minutes, so split long scripts into separate requests. Language is detected from the text automatically across 90+ languages. Billing is per minute of _generated_ audio, rounded up to the next whole minute.

## Example Request

```
{   "text": "Say warmly: Welcome back — your report is ready.",   "voice": "Kore" }
```

## Example Response

```
{
  "request_id": "gemini-3-1-flash-tts_01a0xxxx-xxxx-xxxx-xxxx-xxxxxxxxxxxx",
  "status": "QUEUED",
  "polling_url": "https://gateway.pixazo.ai/v2/requests/status/gemini-3-1-flash-tts_01a0xxxx-xxxx-xxxx-xxxx-xxxxxxxxxxxx"
}
```

## Request Headers

Header

Required

Description

Content-Type

Yes

`application/json`

Ocp-Apim-Subscription-Key

Yes

Your API subscription key.

X-Webhook-URL

No

Enable Webhook callbacks (see Webhook section).

## Response Handling

Status Code

Meaning

202

Accepted — request queued; returns `request_id` and `polling_url`.

400

Bad request — a missing or out-of-range parameter. The message names the field.

401

Unauthorized — missing or invalid subscription key.

402

Insufficient balance.

429

Too many requests.

500

Internal server error.

## Retrieving Results

Poll the status endpoint with the `request_id` from the submit response until `status` is `COMPLETED` (or `ERROR`), then download `output.media_url`.

```
curl 'https://gateway.pixazo.ai/v2/requests/status/gemini-3-1-flash-tts_01a0xxxx-xxxx-xxxx-xxxx-xxxxxxxxxxxx' \
  -H 'Ocp-Apim-Subscription-Key: YOUR_SUBSCRIPTION_KEY'
```

### Completed response

```
{
  "request_id": "gemini-3-1-flash-tts_01a0xxxx-xxxx-xxxx-xxxx-xxxxxxxxxxxx",
  "status": "COMPLETED",
  "model_id": "gemini-3-1-flash-tts",
  "output": {
    "media_url": [
      "https://pub-582b7213209642b9b995c96c95a30381.r2.dev/v1/{request_id}/transcript.json"
    ],
    "media_type": "application/json"
  },
  "duration": 19.17,
  "created_at": "2026-08-01T09:14:16.102Z",
  "completed_at": "2026-08-01T09:14:22.870Z"
}
```

## Response Fields

Field

Type

Description

request\_id

string

Unique request identifier.

status

string

QUEUED, PROCESSING, COMPLETED or ERROR.

model\_id

string

The model that handled the request.

output.media\_url

array

URL of the generated WAV file (PCM, 24 kHz, mono).

output.media\_type

string

Always `audio/wav`.

created\_at

string

Request creation timestamp.

completed\_at

string

Completion timestamp.

error

string

Error message when `status` is ERROR.

## Status Values & Flow

`QUEUED` → `PROCESSING` → `COMPLETED` (success) or `ERROR` (failure).

### Pricing

Billed at **$0.006 per minute of generated audio**, rounded up to the next whole minute. You are charged for the audio produced, not the text you submit.

Audio produced

Billed minutes

Cost

A 10-second clip

1

$0.006

A 45-second clip

1

$0.006

A 3-minute narration

3

$0.018

A 10-minute narration

10

$0.06

Failed requests are not billed.
