Gemini Voice 3.5, Gemini Voice 3.1 Flash API: Pricing, Documentation
by Google
Google's Gemini voice models on one endpoint pair: Gemini 3.5 Transcribe turns recorded speech into text with speaker labels and word-level timestamps across 85+ locales, and Gemini 3.1 Flash TTS turns text into speech with 30 voices, natural-language delivery control and two-speaker dialogue. Both are billed per minute of audio.

Models Version
Get $5 Free Credit on First Payment
No strings attached — add funds and get $5 bonus instantly
Gemini 3.5 Transcribe API Documentation
Google's Gemini 3.5 Transcribe converts pre-recorded speech to text, with automatic language identification across 85+ locales, speaker diarization, word-level timestamps and custom-vocabulary biasing. Accepts WAV, MP3, AIFF, AAC, OGG, FLAC, M4A, Opus, WebM and MPEG, up to 1 hour per request. Asynchronous: submit returns a request_id; poll the status endpoint until the request is COMPLETED, then fetch the JSON transcript from the returned url.
POST https://gateway.pixazo.ai/gemini-3-5-transcribe/v1/speech-to-textAuthentication
All requests require an API key passed via header.
| Header | Type | Required | Description |
|---|---|---|---|
| Ocp-Apim-Subscription-Key | string | Yes | Your API subscription key |
Speech to Text - Gemini 3.5 Transcribe
Request Code
POST https://gateway.pixazo.ai/gemini-3-5-transcribe/v1/speech-to-text
Content-Type: application/json
Ocp-Apim-Subscription-Key: YOUR_SUBSCRIPTION_KEY
{
"audio_url": "https://pub-582b7213209642b9b995c96c95a30381.r2.dev/doc-assets/audio/speech-17s.mp3"
}import requests
url = "https://gateway.pixazo.ai/gemini-3-5-transcribe/v1/speech-to-text"
headers = {
"Content-Type": "application/json",
"Ocp-Apim-Subscription-Key": "YOUR_SUBSCRIPTION_KEY"
}
data = {
"audio_url": "https://pub-582b7213209642b9b995c96c95a30381.r2.dev/doc-assets/audio/speech-17s.mp3"
}
resp = requests.post(url, json=data, headers=headers)
print(resp.json())const res = await fetch("https://gateway.pixazo.ai/gemini-3-5-transcribe/v1/speech-to-text", {
method: "POST",
headers: {
"Content-Type": "application/json",
"Ocp-Apim-Subscription-Key": "YOUR_SUBSCRIPTION_KEY"
},
body: JSON.stringify({
"audio_url": "https://pub-582b7213209642b9b995c96c95a30381.r2.dev/doc-assets/audio/speech-17s.mp3"
})
});
console.log(await res.json());curl -X POST 'https://gateway.pixazo.ai/gemini-3-5-transcribe/v1/speech-to-text' \
-H 'Content-Type: application/json' \
-H 'Ocp-Apim-Subscription-Key: YOUR_SUBSCRIPTION_KEY' \
--data-raw '{"audio_url": "https://pub-582b7213209642b9b995c96c95a30381.r2.dev/doc-assets/audio/speech-17s.mp3"}'Output
{
"request_id": "gemini-3-5-transcribe_01a0xxxx-xxxx-xxxx-xxxx-xxxxxxxxxxxx",
"status": "QUEUED",
"polling_url": "https://gateway.pixazo.ai/v2/requests/status/gemini-3-5-transcribe_01a0xxxx-xxxx-xxxx-xxxx-xxxxxxxxxxxx"
}Webhook (Optional)
Instead of polling, you can receive a Webhook callback when the request reaches a terminal state. Provide a Webhook URL via header on the submit request.
| Header | Required | Description |
|---|---|---|
| X-Webhook-URL | To enable | HTTPS URL to receive the Webhook callback. |
| X-Webhook-Mode | No | terminal (default, one callback on COMPLETED/ERROR) or sync (per-poll callbacks). |
Example: enable Webhook
curl -X POST 'https://gateway.pixazo.ai/gemini-3-5-transcribe/v1/speech-to-text' \
-H 'Content-Type: application/json' \
-H 'Ocp-Apim-Subscription-Key: YOUR_SUBSCRIPTION_KEY' \
-H 'X-Webhook-URL: https://your-server.com/webhook' \
--data-raw '{"audio_url": "https://pub-582b7213209642b9b995c96c95a30381.r2.dev/doc-assets/audio/speech-17s.mp3"}'Callback Payload (success)
{
"request_id": "gemini-3-5-transcribe_01a0xxxx-xxxx-xxxx-xxxx-xxxxxxxxxxxx",
"status": "COMPLETED",
"model_id": "gemini-3-5-transcribe",
"output": {
"media_url": [
"https://pub-582b7213209642b9b995c96c95a30381.r2.dev/v1/{request_id}/transcript.json"
],
"media_type": "application/json"
},
"duration": 19.17,
"created_at": "2026-08-01T09:14:16.102Z",
"completed_at": "2026-08-01T09:14:22.870Z"
}Failure callback shape
{
"request_id": "gemini-3-5-transcribe_01a0xxxx-xxxx-xxxx-xxxx-xxxxxxxxxxxx",
"status": "ERROR",
"model_id": "gemini-3-5-transcribe",
"error": "Description of the failure"
}Delivery semantics
- terminal mode: one Webhook callback when the request is COMPLETED or ERROR.
- sync mode: a Webhook callback on each status change.
- Callbacks are idempotent on
request_id— de-duplicate on it. - Respond
200within a few seconds; the Webhook endpoint must be HTTPS.
Request Parameters
| Parameter | Required | Type | Default | Allowed values / range | Description |
|---|---|---|---|---|---|
audio_url | Yes | string | — | a publicly reachable http(s) url | The recording to transcribe. We fetch it server-side, so it must be reachable from the internet — a signed url is fine, a private one is not. Accepts WAV, MP3, AIFF, AAC, OGG, FLAC, M4A, Opus, WebM and MPEG. |
language_codes | No | array | — (auto-detected) | BCP-47 codes, e.g. ["en-US"] | Leave this out and the model identifies the language itself across 85+ locales, including speakers switching language mid-recording. Set it to pin the transcript to a known language. |
mode | No | string | verbatim | verbatim, smart | verbatim transcribes exactly what was said, keeping filler words, repetitions and false starts. smart cleans that up — removing disfluencies, applying self-corrections, and formatting lists, numbers, dates and paragraph breaks. smart cannot be combined with diarization or timestamps. |
diarization | No | boolean | false | true, false | Label who is speaking, for up to 8 speakers. Each word in the transcript carries a speaker tag. Requires mode verbatim. |
timestamps | No | boolean | false | true, false | Word-level start and end times, in seconds. Requires mode verbatim. |
custom_vocabulary | No | array | — | up to 1,000 terms | Bias the transcript toward terms the model would otherwise misspell — product names, people, domain jargon. |
Limits
Audio up to 1 hour per request. Turning on diarization or timestamps lowers that ceiling to 30 minutes. Billing is per minute of input audio, rounded up to the next whole minute — so a 20-second clip and a 55-second clip both bill one minute.
Example Request
{
"audio_url": "https://pub-582b7213209642b9b995c96c95a30381.r2.dev/doc-assets/audio/speech-17s.mp3"
}Example Response
{
"request_id": "gemini-3-5-transcribe_01a0xxxx-xxxx-xxxx-xxxx-xxxxxxxxxxxx",
"status": "QUEUED",
"polling_url": "https://gateway.pixazo.ai/v2/requests/status/gemini-3-5-transcribe_01a0xxxx-xxxx-xxxx-xxxx-xxxxxxxxxxxx"
}Request Headers
| Header | Required | Description |
|---|---|---|
| Content-Type | Yes | application/json |
| Ocp-Apim-Subscription-Key | Yes | Your API subscription key. |
| X-Webhook-URL | No | Enable Webhook callbacks (see Webhook section). |
Response Handling
| Status Code | Meaning |
|---|---|
| 202 | Accepted — request queued; returns request_id and polling_url. |
| 400 | Bad request — a missing or out-of-range parameter. The message names the field. |
| 401 | Unauthorized — missing or invalid subscription key. |
| 402 | Insufficient balance. |
| 429 | Too many requests. |
| 500 | Internal server error. |
Retrieving Results
Poll the status endpoint with the request_id from the submit response until status is COMPLETED (or ERROR), then download output.media_url.
curl 'https://gateway.pixazo.ai/v2/requests/status/gemini-3-5-transcribe_01a0xxxx-xxxx-xxxx-xxxx-xxxxxxxxxxxx' \
-H 'Ocp-Apim-Subscription-Key: YOUR_SUBSCRIPTION_KEY'Completed response
{
"request_id": "gemini-3-5-transcribe_01a0xxxx-xxxx-xxxx-xxxx-xxxxxxxxxxxx",
"status": "COMPLETED",
"model_id": "gemini-3-5-transcribe",
"output": {
"media_url": [
"https://pub-582b7213209642b9b995c96c95a30381.r2.dev/v1/{request_id}/transcript.json"
],
"media_type": "application/json"
},
"duration": 19.17,
"created_at": "2026-08-01T09:14:16.102Z",
"completed_at": "2026-08-01T09:14:22.870Z"
}Response Fields
| Field | Type | Description |
|---|---|---|
| request_id | string | Unique request identifier. |
| status | string | QUEUED, PROCESSING, COMPLETED or ERROR. |
| model_id | string | The model that handled the request. |
| output.media_url | array | URL of the JSON transcript file. Fetch it to read text, and segments when you asked for timestamps or diarization. |
| output.media_type | string | Always application/json — the result is a transcript document, not audio. |
| created_at | string | Request creation timestamp. |
| completed_at | string | Completion timestamp. |
| error | string | Error message when status is ERROR. |
Status Values & Flow
QUEUED → PROCESSING → COMPLETED (success) or ERROR (failure).
Pricing
Billed per minute of submitted audio, rounded up to the next whole minute — a 10-second clip and a 45-second clip both bill one minute. You are charged for the recording you send, not the length of the transcript. Failed requests are not billed. The current rate is shown on this model's page.
Gemini 3.5 Transcribe API Pricing
| Resolution | Price (USD) |
|---|---|
| Per minute of audio (rounded up to the next full minute) | $0.005 |
Gemini 3.1 Flash TTS API Documentation
Google's Gemini 3.1 Flash TTS turns text into speech, with 30 prebuilt voices, natural-language control over delivery, optional two-speaker dialogue and 90+ languages detected automatically. Asynchronous: submit returns a request_id; poll the status endpoint until the request is COMPLETED, then download the WAV from the returned url.
POST https://gateway.pixazo.ai/gemini-3-1-flash-tts/v1/text-to-speechAuthentication
All requests require an API key passed via header.
| Header | Type | Required | Description |
|---|---|---|---|
| Ocp-Apim-Subscription-Key | string | Yes | Your API subscription key |
Text to Speech - Gemini 3.1 Flash TTS
Request Code
POST https://gateway.pixazo.ai/gemini-3-1-flash-tts/v1/text-to-speech
Content-Type: application/json
Ocp-Apim-Subscription-Key: YOUR_SUBSCRIPTION_KEY
{ "text": "Say warmly: Welcome back — your report is ready.", "voice": "Kore" }import requests
url = "https://gateway.pixazo.ai/gemini-3-1-flash-tts/v1/text-to-speech"
headers = {
"Content-Type": "application/json",
"Ocp-Apim-Subscription-Key": "YOUR_SUBSCRIPTION_KEY"
}
data = { "text": "Say warmly: Welcome back — your report is ready.", "voice": "Kore" }
resp = requests.post(url, json=data, headers=headers)
print(resp.json())const res = await fetch("https://gateway.pixazo.ai/gemini-3-1-flash-tts/v1/text-to-speech", {
method: "POST",
headers: {
"Content-Type": "application/json",
"Ocp-Apim-Subscription-Key": "YOUR_SUBSCRIPTION_KEY"
},
body: JSON.stringify({ "text": "Say warmly: Welcome back — your report is ready.", "voice": "Kore" })
});
console.log(await res.json());curl -X POST 'https://gateway.pixazo.ai/gemini-3-1-flash-tts/v1/text-to-speech' \
-H 'Content-Type: application/json' \
-H 'Ocp-Apim-Subscription-Key: YOUR_SUBSCRIPTION_KEY' \
--data-raw '{ "text": "Say warmly: Welcome back — your report is ready.", "voice": "Kore" }'Output
{
"request_id": "gemini-3-1-flash-tts_01a0xxxx-xxxx-xxxx-xxxx-xxxxxxxxxxxx",
"status": "QUEUED",
"polling_url": "https://gateway.pixazo.ai/v2/requests/status/gemini-3-1-flash-tts_01a0xxxx-xxxx-xxxx-xxxx-xxxxxxxxxxxx"
}Webhook (Optional)
Instead of polling, you can receive a Webhook callback when the request reaches a terminal state. Provide a Webhook URL via header on the submit request.
| Header | Required | Description |
|---|---|---|
| X-Webhook-URL | To enable | HTTPS URL to receive the Webhook callback. |
| X-Webhook-Mode | No | terminal (default, one callback on COMPLETED/ERROR) or sync (per-poll callbacks). |
Example: enable Webhook
curl -X POST 'https://gateway.pixazo.ai/gemini-3-1-flash-tts/v1/text-to-speech' \
-H 'Content-Type: application/json' \
-H 'Ocp-Apim-Subscription-Key: YOUR_SUBSCRIPTION_KEY' \
-H 'X-Webhook-URL: https://your-server.com/webhook' \
--data-raw '{ "text": "Say warmly: Welcome back — your report is ready.", "voice": "Kore" }'Callback Payload (success)
{
"request_id": "gemini-3-1-flash-tts_01a0xxxx-xxxx-xxxx-xxxx-xxxxxxxxxxxx",
"status": "COMPLETED",
"model_id": "gemini-3-1-flash-tts",
"output": {
"media_url": [
"https://pub-582b7213209642b9b995c96c95a30381.r2.dev/v1/{request_id}/transcript.json"
],
"media_type": "application/json"
},
"duration": 19.17,
"created_at": "2026-08-01T09:14:16.102Z",
"completed_at": "2026-08-01T09:14:22.870Z"
}Failure callback shape
{
"request_id": "gemini-3-1-flash-tts_01a0xxxx-xxxx-xxxx-xxxx-xxxxxxxxxxxx",
"status": "ERROR",
"model_id": "gemini-3-1-flash-tts",
"error": "Description of the failure"
}Delivery semantics
- terminal mode: one Webhook callback when the request is COMPLETED or ERROR.
- sync mode: a Webhook callback on each status change.
- Callbacks are idempotent on
request_id— de-duplicate on it. - Respond
200within a few seconds; the Webhook endpoint must be HTTPS.
Request Parameters
| Parameter | Required | Type | Default | Allowed values / range | Description |
|---|---|---|---|---|---|
text | Yes | string | — | up to ~8,000 tokens | What to say. This is also where you direct how it is said — the model follows natural-language direction, so "Say warmly and slowly: ..." works, as do inline tags like [whispers] and [laughs]. |
voice | No | string | Kore | Zephyr, Puck, Charon, Kore, Fenrir, Leda, Orus, Aoede, Callirrhoe, Autonoe, Enceladus, Iapetus, Umbriel, Algieba, Despina, Erinome, Algenib, Rasalgethi, Laomedeia, Achernar, Alnilam, Schedar, Gacrux, Pulcherrima, Achird, Zubenelgenubi, Vindemiatrix, Sadachbia, Sadaltager, Sulafat | Which of the 30 prebuilt voices to speak in. Ignored when speakers is given. |
speakers | No | array | — | at most 2 entries | Two-speaker dialogue. Each entry is {"speaker": "Joe", "voice": "Kore"}, where speaker matches a name used in your text. Supplying this replaces voice. |
Limits
A session has a 32k-token context. Speech quality can drift on outputs longer than a few minutes, so split long scripts into separate requests. Language is detected from the text automatically across 90+ languages. Billing is per minute of generated audio, rounded up to the next whole minute.
Example Request
{ "text": "Say warmly: Welcome back — your report is ready.", "voice": "Kore" }Example Response
{
"request_id": "gemini-3-1-flash-tts_01a0xxxx-xxxx-xxxx-xxxx-xxxxxxxxxxxx",
"status": "QUEUED",
"polling_url": "https://gateway.pixazo.ai/v2/requests/status/gemini-3-1-flash-tts_01a0xxxx-xxxx-xxxx-xxxx-xxxxxxxxxxxx"
}Request Headers
| Header | Required | Description |
|---|---|---|
| Content-Type | Yes | application/json |
| Ocp-Apim-Subscription-Key | Yes | Your API subscription key. |
| X-Webhook-URL | No | Enable Webhook callbacks (see Webhook section). |
Response Handling
| Status Code | Meaning |
|---|---|
| 202 | Accepted — request queued; returns request_id and polling_url. |
| 400 | Bad request — a missing or out-of-range parameter. The message names the field. |
| 401 | Unauthorized — missing or invalid subscription key. |
| 402 | Insufficient balance. |
| 429 | Too many requests. |
| 500 | Internal server error. |
Retrieving Results
Poll the status endpoint with the request_id from the submit response until status is COMPLETED (or ERROR), then download output.media_url.
curl 'https://gateway.pixazo.ai/v2/requests/status/gemini-3-1-flash-tts_01a0xxxx-xxxx-xxxx-xxxx-xxxxxxxxxxxx' \
-H 'Ocp-Apim-Subscription-Key: YOUR_SUBSCRIPTION_KEY'Completed response
{
"request_id": "gemini-3-1-flash-tts_01a0xxxx-xxxx-xxxx-xxxx-xxxxxxxxxxxx",
"status": "COMPLETED",
"model_id": "gemini-3-1-flash-tts",
"output": {
"media_url": [
"https://pub-582b7213209642b9b995c96c95a30381.r2.dev/v1/{request_id}/transcript.json"
],
"media_type": "application/json"
},
"duration": 19.17,
"created_at": "2026-08-01T09:14:16.102Z",
"completed_at": "2026-08-01T09:14:22.870Z"
}Response Fields
| Field | Type | Description |
|---|---|---|
| request_id | string | Unique request identifier. |
| status | string | QUEUED, PROCESSING, COMPLETED or ERROR. |
| model_id | string | The model that handled the request. |
| output.media_url | array | URL of the generated WAV file (PCM, 24 kHz, mono). |
| output.media_type | string | Always audio/wav. |
| created_at | string | Request creation timestamp. |
| completed_at | string | Completion timestamp. |
| error | string | Error message when status is ERROR. |
Status Values & Flow
QUEUED → PROCESSING → COMPLETED (success) or ERROR (failure).
Pricing
Billed per minute of generated audio, rounded up to the next whole minute — a 10-second clip and a 45-second clip both bill one minute. You are charged for the audio produced, not the text you submit. Failed requests are not billed. The current rate is shown on this model's page.
Gemini 3.1 Flash TTS API Pricing
| Resolution | Price (USD) |
|---|---|
| Per minute of generated audio (rounded up to the next full minute) | $0.0384 |
⚡ Performance
Live usage measured on Pixazo's gateway, split by model version. Generation time is how long a generation takes end-to-end (lower is better). Success rate is the percent of generations that complete (higher is better).
〰 Uptime
Percent of generations that succeeded over the selected period, per model version.