Speaker Diarization API - Who Spoke When: Pricing, Documentation
by Pixazo
Speaker Diarization API answers the question transcription leaves open: not what was said, but who said it. Give it a meeting recording, an interview, a podcast or a video URL and it returns the number of distinct speakers plus a timeline of segments, each with a start time, an end time and a speaker label. Pair it with a speech-to-text model and a flat transcript becomes an attributed one, which is what call analytics, meeting summaries and subtitle tracks actually need. Results are delivered as a JSON file. Billing is per second of input audio, so a ten-minute recording costs $3.00.

Models Version
Get $5 Free Credit on First Payment
No strings attached — add funds and get $5 bonus instantly
Speaker Diarization 1.0 API Documentation
https://gateway.pixazo.ai/audio-tools/v1/diarize
Authentication
All requests require an API key passed via header.
Pricing: Billed at $0.005 per second of input audio — measured on the file you submit. A 10-minute recording costs $3.00. Accepts audio or video; for video the audio is extracted server-side and you are billed on its duration.
Rounding: billing is per second but always rounds up to a whole second — a 5.06-second source bills as 6 seconds. The shortest billable job is 1 second.
Retries: this is an asynchronous job on shared encoding capacity, so a request can occasionally come back processing_failed or take much longer than usual. These are transient and succeed on a retry, and a failed job is never charged — the wallet hold is released. If you chain these tools, retry a failed step rather than failing the whole pipeline.
| Header | Type | Required | Description |
|---|---|---|---|
| Ocp-Apim-Subscription-Key | string | Yes | Your API subscription key |
Audio Diarize generate request
Request Code
POST https://gateway.pixazo.ai/audio-tools/v1/diarize
Content-Type: application/json
Cache-Control: no-cache
Ocp-Apim-Subscription-Key: YOUR_API_KEY
{
"audio_url": "https://api-assets.pixazo.ai/media/audio-tools-example.mp3"
}
import requests
url = "https://gateway.pixazo.ai/audio-tools/v1/diarize"
headers = {
"Content-Type": "application/json",
"Cache-Control": "no-cache",
"Ocp-Apim-Subscription-Key": "YOUR_API_KEY"
}
data = {
"audio_url": "https://api-assets.pixazo.ai/media/audio-tools-example.mp3"
}
response = requests.post(url, json=data, headers=headers)
print(response.json())
const url = "https://gateway.pixazo.ai/audio-tools/v1/diarize";
const headers = {
"Content-Type": "application/json",
"Cache-Control": "no-cache",
"Ocp-Apim-Subscription-Key": "YOUR_API_KEY"
};
const data = {
"audio_url": "https://api-assets.pixazo.ai/media/audio-tools-example.mp3"
};
fetch(url, {
method: "POST",
headers: headers,
body: JSON.stringify(data)
})
.then(response => response.json())
.then(data => console.log(data));
curl -X POST "https://gateway.pixazo.ai/audio-tools/v1/diarize" \
-H "Content-Type: application/json" \
-H "Cache-Control: no-cache" \
-H "Ocp-Apim-Subscription-Key: YOUR_API_KEY" \
--data-raw '{
"audio_url": "https://api-assets.pixazo.ai/media/audio-tools-example.mp3"
}'
Output
{
"request_id": "audio-diarize_019dxxxx-xxxx-xxxx-xxxx-xxxxxxxxxxxx",
"status": "COMPLETED",
"model_id": "audio-diarize",
"error": null,
"output": {
"media_url": [
"https://pub-582b7213209642b9b995c96c95a30381.r2.dev/v1/audio-diarize_019dxxxx-xxxx-xxxx-xxxx-xxxxxxxxxxxx/output.json"
],
"media_type": "application/json"
},
"created_at": "2026-03-31T10:00:00.000Z",
"updated_at": "2026-03-31T10:00:15.000Z",
"completed_at": "2026-03-31T10:00:15.000Z"
}
Webhook (Optional)
Add the X-Webhook-URL header to your generate request to receive a POST callback instead of polling.
X-Webhook-URL: https://your-server.com/webhook/callback
Request Parameters - Audio Diarize generate request
| Parameter | Required | Type | Default | Allowed values / range | Description |
|---|---|---|---|---|---|
| audio_url | Yes | string | — | HTTP(S) URL; ≤ 500 MB | Publicly reachable HTTP(S) URL of an audio or video file; the server downloads it, so it must be fetchable from the public internet. For video the audio track is extracted server-side. Anything that is not an http:// or https:// URL, or that cannot be fetched, is rejected with 400. Maximum file size: 500 MB. |
| speaker_count | No | integer | — (auto-detected) | 1–20 | Exact number of speakers, when you already know it. Must be a whole number from 1 to 20 — any other value is rejected with 400. Omit it to let the model detect the speaker count itself; never send a guess, because a supplied value forces the audio to be split into exactly that many speakers. |
The result is a JSON file, not audio: media_url points at it and media_type is application/json. It holds ok, speakers and segments — an array of {speaker, start, end} objects with 0-based speaker numbers and float seconds, sorted by start.
speakers is null, not 0, when the run finds no speech at all — music, silence or room tone — and segments is then an empty array. That is a successful, billable result, not an error: treat null as “nobody spoke”. A count is never guessed, so you will not get a phantom single speaker out of an instrumental track.
Example Request
{
"audio_url": "https://api-assets.pixazo.ai/media/audio-tools-example.mp3"
}
Response
{
"request_id": "audio-diarize_019dxxxx-xxxx-xxxx-xxxx-xxxxxxxxxxxx",
"status": "QUEUED",
"polling_url": "https://gateway.pixazo.ai/v2/requests/status/audio-diarize_019dxxxx-xxxx-xxxx-xxxx-xxxxxxxxxxxx"
}
Request Headers
| Header | Value |
|---|---|
| Content-Type | application/json |
| Cache-Control | no-cache |
| Ocp-Apim-Subscription-Key | YOUR_API_KEY |
Response Handling
Common status codes.
| Code | Meaning |
|---|---|
| 202 | Accepted — Request queued |
| 400 | Bad Request |
| 401 | Unauthorized |
| 402 | Insufficient Balance |
| 403 | Forbidden |
| 429 | Too Many Requests |
| 500 | Internal Server Error |
Error Responses
Queue system errors and model validation errors.
Queue System Errors
// 402 — Insufficient balance
{
"error": "Insufficient Balance",
"message": "Your wallet does not have enough balance."
}
// 400 — Model not found
{
"error": "Model not found",
"message": "Model 'audio-diarize' not found or is disabled"
}
Error via Status/Webhook
{
"request_id": "audio-diarize_019dxxxx-xxxx-xxxx-xxxx-xxxxxxxxxxxx",
"status": "ERROR",
"model_id": "audio-diarize",
"error": "Description of the error",
"output": null
}
Retrieving Results
Poll the universal status endpoint to check progress and retrieve results.
Endpoint
GET https://gateway.pixazo.ai/v2/requests/status/{request_id}
Ocp-Apim-Subscription-Key: YOUR_API_KEY
cURL Example
curl -H "Ocp-Apim-Subscription-Key: YOUR_API_KEY" \
"https://gateway.pixazo.ai/v2/requests/status/audio-diarize_019dxxxx-xxxx-xxxx-xxxx-xxxxxxxxxxxx"
Response (Completed)
{
"request_id": "audio-diarize_019dxxxx-xxxx-xxxx-xxxx-xxxxxxxxxxxx",
"status": "COMPLETED",
"model_id": "audio-diarize",
"error": null,
"output": {
"media_url": [
"https://pub-582b7213209642b9b995c96c95a30381.r2.dev/v1/audio-diarize_019dxxxx-xxxx-xxxx-xxxx-xxxxxxxxxxxx/output.json"
],
"media_type": "application/json"
},
"created_at": "2026-03-31T10:00:00.000Z",
"updated_at": "2026-03-31T10:00:15.000Z",
"completed_at": "2026-03-31T10:00:15.000Z"
}
Response Fields
| Field | Type | Description |
|---|---|---|
| request_id | string | Unique request identifier |
| status | string | QUEUED, PROCESSING, COMPLETED, FAILED, or ERROR |
| model_id | string | Model that processed the request |
| error | string|null | Error message if failed |
| output.media_url | array | URLs to generated media (R2 CDN) |
| output.media_type | string | MIME type of the output |
| created_at | string | When request was created |
| completed_at | string | When request completed |
| polling_url | string | Status URL (initial response only) |
Status Values
| Status | Description |
|---|---|
| QUEUED | Request accepted, waiting to be processed |
| PROCESSING | Being processed by the model |
| COMPLETED | Done — output contains the result |
| FAILED | Failed — check error field |
| ERROR | System error — not charged |
Status Flow
QUEUED → PROCESSING → COMPLETED
→ FAILED
→ ERROR
Typical Workflow
- Send a generate request to the API endpoint
- Save the
request_idfrom the response - Poll every 5-10 seconds:
GET /v2/requests/status/{request_id} - When
statusis"COMPLETED", download fromoutput.media_url
Tip: Use X-Webhook-URL header to get a callback instead of polling.
Speaker Diarization 1.0 API Pricing
⚡ Performance
Live usage measured on Pixazo's gateway, split by model version. Generation time is how long a generation takes end-to-end (lower is better). Success rate is the percent of generations that complete (higher is better).
〰 Uptime
Percent of generations that succeeded over the selected period, per model version.