Pixazo APIAudio Tools APISpeaker Diarization
Pixazo APIAudio Tools APISpeaker Diarization

Speaker Diarization API - Who Spoke When: Pricing, Documentation

by Pixazo

Speaker Diarization API answers the question transcription leaves open: not what was said, but who said it. Give it a meeting recording, an interview, a podcast or a video URL and it returns the number of distinct speakers plus a timeline of segments, each with a start time, an end time and a speaker label. Pair it with a speech-to-text model and a flat transcript becomes an attributed one, which is what call analytics, meeting summaries and subtitle tracks actually need. Results are delivered as a JSON file. Billing is per second of input audio, so a ten-minute recording costs $3.00.

Get API Key
Speaker Diarization API

Models Version

WELCOME BONUS

Get $5 Free Credit on First Payment

No strings attached — add funds and get $5 bonus instantly

Claim Your $5 →

Speaker Diarization 1.0 API Documentation

https://gateway.pixazo.ai/audio-tools/v1/diarize

Authentication

All requests require an API key passed via header.

Pricing: Billed at $0.005 per second of input audio — measured on the file you submit. A 10-minute recording costs $3.00. Accepts audio or video; for video the audio is extracted server-side and you are billed on its duration.

Rounding: billing is per second but always rounds up to a whole second — a 5.06-second source bills as 6 seconds. The shortest billable job is 1 second.

Retries: this is an asynchronous job on shared encoding capacity, so a request can occasionally come back processing_failed or take much longer than usual. These are transient and succeed on a retry, and a failed job is never charged — the wallet hold is released. If you chain these tools, retry a failed step rather than failing the whole pipeline.

HeaderTypeRequiredDescription
Ocp-Apim-Subscription-KeystringYesYour API subscription key

Audio Diarize generate request

Request Code

POST https://gateway.pixazo.ai/audio-tools/v1/diarize
Content-Type: application/json
Cache-Control: no-cache
Ocp-Apim-Subscription-Key: YOUR_API_KEY

{
  "audio_url": "https://api-assets.pixazo.ai/media/audio-tools-example.mp3"
}
import requests

url = "https://gateway.pixazo.ai/audio-tools/v1/diarize"
headers = {
    "Content-Type": "application/json",
    "Cache-Control": "no-cache",
    "Ocp-Apim-Subscription-Key": "YOUR_API_KEY"
}
data = {
    "audio_url": "https://api-assets.pixazo.ai/media/audio-tools-example.mp3"
}

response = requests.post(url, json=data, headers=headers)
print(response.json())
const url = "https://gateway.pixazo.ai/audio-tools/v1/diarize";
const headers = {
  "Content-Type": "application/json",
  "Cache-Control": "no-cache",
  "Ocp-Apim-Subscription-Key": "YOUR_API_KEY"
};
const data = {
  "audio_url": "https://api-assets.pixazo.ai/media/audio-tools-example.mp3"
};

fetch(url, {
  method: "POST",
  headers: headers,
  body: JSON.stringify(data)
})
.then(response => response.json())
.then(data => console.log(data));
curl -X POST "https://gateway.pixazo.ai/audio-tools/v1/diarize" \
  -H "Content-Type: application/json" \
  -H "Cache-Control: no-cache" \
  -H "Ocp-Apim-Subscription-Key: YOUR_API_KEY" \
  --data-raw '{
    "audio_url": "https://api-assets.pixazo.ai/media/audio-tools-example.mp3"
  }'

Output

{
  "request_id": "audio-diarize_019dxxxx-xxxx-xxxx-xxxx-xxxxxxxxxxxx",
  "status": "COMPLETED",
  "model_id": "audio-diarize",
  "error": null,
  "output": {
    "media_url": [
      "https://pub-582b7213209642b9b995c96c95a30381.r2.dev/v1/audio-diarize_019dxxxx-xxxx-xxxx-xxxx-xxxxxxxxxxxx/output.json"
    ],
    "media_type": "application/json"
  },
  "created_at": "2026-03-31T10:00:00.000Z",
  "updated_at": "2026-03-31T10:00:15.000Z",
  "completed_at": "2026-03-31T10:00:15.000Z"
}

Webhook (Optional)

Add the X-Webhook-URL header to your generate request to receive a POST callback instead of polling.

X-Webhook-URL: https://your-server.com/webhook/callback

Request Parameters - Audio Diarize generate request

Parameter Required Type Default Allowed values / range Description
audio_urlYesstringHTTP(S) URL; ≤ 500 MBPublicly reachable HTTP(S) URL of an audio or video file; the server downloads it, so it must be fetchable from the public internet. For video the audio track is extracted server-side. Anything that is not an http:// or https:// URL, or that cannot be fetched, is rejected with 400. Maximum file size: 500 MB.
speaker_countNointeger— (auto-detected)1–20Exact number of speakers, when you already know it. Must be a whole number from 1 to 20 — any other value is rejected with 400. Omit it to let the model detect the speaker count itself; never send a guess, because a supplied value forces the audio to be split into exactly that many speakers.

The result is a JSON file, not audio: media_url points at it and media_type is application/json. It holds ok, speakers and segments — an array of {speaker, start, end} objects with 0-based speaker numbers and float seconds, sorted by start.

speakers is null, not 0, when the run finds no speech at all — music, silence or room tone — and segments is then an empty array. That is a successful, billable result, not an error: treat null as “nobody spoke”. A count is never guessed, so you will not get a phantom single speaker out of an instrumental track.

Example Request

{
  "audio_url": "https://api-assets.pixazo.ai/media/audio-tools-example.mp3"
}

Response

{
  "request_id": "audio-diarize_019dxxxx-xxxx-xxxx-xxxx-xxxxxxxxxxxx",
  "status": "QUEUED",
  "polling_url": "https://gateway.pixazo.ai/v2/requests/status/audio-diarize_019dxxxx-xxxx-xxxx-xxxx-xxxxxxxxxxxx"
}

Request Headers

Header Value
Content-Typeapplication/json
Cache-Controlno-cache
Ocp-Apim-Subscription-KeyYOUR_API_KEY

Response Handling

Common status codes.

CodeMeaning
202Accepted — Request queued
Bad Request
401Unauthorized
402Insufficient Balance
403Forbidden
Too Many Requests
500Internal Server Error

Error Responses

Queue system errors and model validation errors.

Queue System Errors

// 402 — Insufficient balance
{
  "error": "Insufficient Balance",
  "message": "Your wallet does not have enough balance."
}
// 400 — Model not found
{
  "error": "Model not found",
  "message": "Model 'audio-diarize' not found or is disabled"
}

Error via Status/Webhook

{
  "request_id": "audio-diarize_019dxxxx-xxxx-xxxx-xxxx-xxxxxxxxxxxx",
  "status": "ERROR",
  "model_id": "audio-diarize",
  "error": "Description of the error",
  "output": null
}

Retrieving Results

Poll the universal status endpoint to check progress and retrieve results.

Endpoint

GET https://gateway.pixazo.ai/v2/requests/status/{request_id}
Ocp-Apim-Subscription-Key: YOUR_API_KEY

cURL Example

curl -H "Ocp-Apim-Subscription-Key: YOUR_API_KEY" \
  "https://gateway.pixazo.ai/v2/requests/status/audio-diarize_019dxxxx-xxxx-xxxx-xxxx-xxxxxxxxxxxx"

Response (Completed)

{
  "request_id": "audio-diarize_019dxxxx-xxxx-xxxx-xxxx-xxxxxxxxxxxx",
  "status": "COMPLETED",
  "model_id": "audio-diarize",
  "error": null,
  "output": {
    "media_url": [
      "https://pub-582b7213209642b9b995c96c95a30381.r2.dev/v1/audio-diarize_019dxxxx-xxxx-xxxx-xxxx-xxxxxxxxxxxx/output.json"
    ],
    "media_type": "application/json"
  },
  "created_at": "2026-03-31T10:00:00.000Z",
  "updated_at": "2026-03-31T10:00:15.000Z",
  "completed_at": "2026-03-31T10:00:15.000Z"
}

Response Fields

FieldTypeDescription
request_idstringUnique request identifier
statusstringQUEUED, PROCESSING, COMPLETED, FAILED, or ERROR
model_idstringModel that processed the request
errorstring|nullError message if failed
output.media_urlarrayURLs to generated media (R2 CDN)
output.media_typestringMIME type of the output
created_atstringWhen request was created
completed_atstringWhen request completed
polling_urlstringStatus URL (initial response only)

Status Values

StatusDescription
QUEUEDRequest accepted, waiting to be processed
PROCESSINGBeing processed by the model
COMPLETEDDone — output contains the result
FAILEDFailed — check error field
ERRORSystem error — not charged

Status Flow

QUEUED → PROCESSING → COMPLETED
                    → FAILED
                    → ERROR

Typical Workflow

  1. Send a generate request to the API endpoint
  2. Save the request_id from the response
  3. Poll every 5-10 seconds: GET /v2/requests/status/{request_id}
  4. When status is "COMPLETED", download from output.media_url

Tip: Use X-Webhook-URL header to get a callback instead of polling.

Speaker Diarization 1.0 API Pricing

Your request will cost $0.005 per second of input audio.
10-minute recording = $3.00

⚡ Performance

Live usage measured on Pixazo's gateway, split by model version. Generation time is how long a generation takes end-to-end (lower is better). Success rate is the percent of generations that complete (higher is better).

Show data for the last
Generations
3,500last 30d
~117 per day
Success rate
94.3%
of completed generations
Generation time
24.5savg
p95 27.7s
Requests
Jul 16max 2,400Aug 14
Speaker Diarization 1.0Avg 117/day
Generation Time
Jul 16max 25.9sAug 14
Speaker Diarization 1.0Avg 24.5s
Error Rate
Jul 16max 14.3%Aug 14
Speaker Diarization 1.0Avg 5.7%

〰 Uptime

Percent of generations that succeeded over the selected period, per model version.

Avg. Success Rate (30d)
94.29%
across all generations of this model family
Uptime
Jul 16max 100%Aug 14
SuccessfulAvg 94.29%