Pixazo APIAudio Tools APIAudio Stem Separation
Pixazo APIAudio Tools APIAudio Stem Separation

Audio Stem Separation API - Vocal Remover & Stem Splitter: Pricing, Documentation

by Pixazo

Audio Stem Separation API splits a finished mix back into its parts. Two-stem mode returns vocals and no_vocals — the pairing dubbing and localisation need, because replacing the speech no longer destroys the music bed underneath it. Four-stem mode returns vocals, drums, bass and other, which is what remixing, sampling and karaoke production start from. Send an audio file or a video URL and the audio is extracted server-side, so you do not need a separate demux step. Billing is per second of input audio, so a three-minute track costs about ninety cents and you always know the price before you submit.

Get API Key
Audio Stem Separation API

Models Version

WELCOME BONUS

Get $5 Free Credit on First Payment

No strings attached — add funds and get $5 bonus instantly

Claim Your $5 →

Audio Stem Separation 1.0 API Documentation

https://gateway.pixazo.ai/audio-tools/v1/separate-stems

Authentication

All requests require an API key passed via header.

Pricing: Billed at $0.005 per second of input audio — measured on the file you submit, not on the stems returned. A 60-second track costs $0.30. Accepts audio or video; for video the audio is extracted server-side and you are billed on its duration.

Rounding: billing is per second but always rounds up to a whole second — a 5.06-second source bills as 6 seconds. The shortest billable job is 1 second.

Retries: this is an asynchronous job on shared encoding capacity, so a request can occasionally come back processing_failed or take much longer than usual. These are transient and succeed on a retry, and a failed job is never charged — the wallet hold is released. If you chain these tools, retry a failed step rather than failing the whole pipeline.

HeaderTypeRequiredDescription
Ocp-Apim-Subscription-KeystringYesYour API subscription key

Audio Separate Stems generate request

Request Code

POST https://gateway.pixazo.ai/audio-tools/v1/separate-stems
Content-Type: application/json
Cache-Control: no-cache
Ocp-Apim-Subscription-Key: YOUR_API_KEY

{
  "audio_url": "https://api-assets.pixazo.ai/media/audio-tools-example.mp3",
  "stems": "two",
  "output_format": "mp3"
}
import requests

url = "https://gateway.pixazo.ai/audio-tools/v1/separate-stems"
headers = {
    "Content-Type": "application/json",
    "Cache-Control": "no-cache",
    "Ocp-Apim-Subscription-Key": "YOUR_API_KEY"
}
data = {
    "audio_url": "https://api-assets.pixazo.ai/media/audio-tools-example.mp3",
    "stems": "two",
    "output_format": "mp3"
}

response = requests.post(url, json=data, headers=headers)
print(response.json())
const url = "https://gateway.pixazo.ai/audio-tools/v1/separate-stems";
const headers = {
  "Content-Type": "application/json",
  "Cache-Control": "no-cache",
  "Ocp-Apim-Subscription-Key": "YOUR_API_KEY"
};
const data = {
  "audio_url": "https://api-assets.pixazo.ai/media/audio-tools-example.mp3",
  "stems": "two",
  "output_format": "mp3"
};

fetch(url, {
  method: "POST",
  headers: headers,
  body: JSON.stringify(data)
})
.then(response => response.json())
.then(data => console.log(data));
curl -X POST "https://gateway.pixazo.ai/audio-tools/v1/separate-stems" \
  -H "Content-Type: application/json" \
  -H "Cache-Control: no-cache" \
  -H "Ocp-Apim-Subscription-Key: YOUR_API_KEY" \
  --data-raw '{
    "audio_url": "https://api-assets.pixazo.ai/media/audio-tools-example.mp3",
    "stems": "two",
    "output_format": "mp3"
  }'

Output

{
  "request_id": "audio-separate-stems_019dxxxx-xxxx-xxxx-xxxx-xxxxxxxxxxxx",
  "status": "QUEUED",
  "polling_url": "https://gateway.pixazo.ai/v2/requests/status/audio-separate-stems_019dxxxx-xxxx-xxxx-xxxx-xxxxxxxxxxxx"
}

Webhook (Optional)

Add the X-Webhook-URL header to your generate request to receive a POST callback instead of polling.

X-Webhook-URL: https://your-server.com/webhook/callback

Request Parameters - Audio Separate Stems generate request

Parameter Required Type Default Allowed values / range Description
audio_urlYesstringHTTP(S) URL; ≤ 500 MBPublicly reachable HTTP(S) URL of an audio or video file. For video the audio track is extracted server-side. Maximum file size: 500 MB.
stemsNoenum"two""two", "four"How many stems to split the track into. "two" returns vocals + no_vocals (dialogue vs the music-and-effects bed — the dubbing case, and roughly 2x faster); "four" returns vocals, drums, bass and other.
output_formatNoenum"mp3""mp3", "wav"Encoding of every returned stem. "mp3" is 192 kbps; "wav" is uncompressed (lossless, larger files).

Example Request

{
  "audio_url": "https://api-assets.pixazo.ai/media/audio-tools-example.mp3",
  "stems": "two",
  "output_format": "mp3"
}

Response

{
  "request_id": "audio-separate-stems_019dxxxx-xxxx-xxxx-xxxx-xxxxxxxxxxxx",
  "status": "QUEUED",
  "polling_url": "https://gateway.pixazo.ai/v2/requests/status/audio-separate-stems_019dxxxx-xxxx-xxxx-xxxx-xxxxxxxxxxxx"
}

Request Headers

Header Value
Content-Typeapplication/json
Cache-Controlno-cache
Ocp-Apim-Subscription-KeyYOUR_API_KEY

Response Handling

Common status codes.

CodeMeaning
202Accepted — Request queued
Bad Request
401Unauthorized
402Insufficient Balance
403Forbidden
Too Many Requests
500Internal Server Error

Error Responses

Queue system errors and model validation errors.

Queue System Errors

// 402 — Insufficient balance
{
  "error": "Insufficient Balance",
  "message": "Your wallet does not have enough balance."
}
// 400 — Model not found
{
  "error": "Model not found",
  "message": "Model 'audio-separate-stems' not found or is disabled"
}

Error via Status/Webhook

{
  "request_id": "audio-separate-stems_019dxxxx-xxxx-xxxx-xxxx-xxxxxxxxxxxx",
  "status": "ERROR",
  "model_id": "audio-separate-stems",
  "error": "Description of the error",
  "output": null
}

Retrieving Results

Poll the universal status endpoint to check progress and retrieve results.

Endpoint

GET https://gateway.pixazo.ai/v2/requests/status/{request_id}
Ocp-Apim-Subscription-Key: YOUR_API_KEY

cURL Example

curl -H "Ocp-Apim-Subscription-Key: YOUR_API_KEY" \
  "https://gateway.pixazo.ai/v2/requests/status/audio-separate-stems_019dxxxx-xxxx-xxxx-xxxx-xxxxxxxxxxxx"

Response (Completed)

{
  "request_id": "audio-separate-stems_019dxxxx-xxxx-xxxx-xxxx-xxxxxxxxxxxx",
  "status": "COMPLETED",
  "model_id": "audio-separate-stems",
  "error": null,
  "output": {
    "media_url": [
      "https://pub-582b7213209642b9b995c96c95a30381.r2.dev/v1/audio-separate-stems_019dxxxx/vocals.mp3",
      "https://pub-582b7213209642b9b995c96c95a30381.r2.dev/v1/audio-separate-stems_019dxxxx/no_vocals.mp3"
    ],
    "media_type": "audio/mpeg"
  },
  "created_at": "2026-03-31T10:00:00.000Z",
  "updated_at": "2026-03-31T10:00:15.000Z",
  "completed_at": "2026-03-31T10:00:15.000Z"
}

Response Fields

FieldTypeDescription
request_idstringUnique request identifier
statusstringQUEUED, PROCESSING, COMPLETED, FAILED, or ERROR
model_idstringModel that processed the request
errorstring|nullError message if failed
output.media_urlarrayURLs to generated media (R2 CDN)
output.media_typestringMIME type of the output
created_atstringWhen request was created
completed_atstringWhen request completed
polling_urlstringStatus URL (initial response only)

Status Values

StatusDescription
QUEUEDRequest accepted, waiting to be processed
PROCESSINGBeing processed by the model
COMPLETEDDone — output contains the result
FAILEDFailed — check error field
ERRORSystem error — not charged

Status Flow

QUEUED → PROCESSING → COMPLETED
                    → FAILED
                    → ERROR

Typical Workflow

  1. Send a generate request to the API endpoint
  2. Save the request_id from the response
  3. Poll every 5-10 seconds: GET /v2/requests/status/{request_id}
  4. When status is "COMPLETED", download from output.media_url

Tip: Use X-Webhook-URL header to get a callback instead of polling.

Audio Stem Separation 1.0 API Pricing

Your request will cost $0.005 per second of input audio.
60-second track = $0.30

⚡ Performance

Live usage measured on Pixazo's gateway, split by model version. Generation time is how long a generation takes end-to-end (lower is better). Success rate is the percent of generations that complete (higher is better).

Show data for the last
Generations
3,100last 30d
~103 per day
Success rate
93.5%
of completed generations
Generation time
27.6savg
p95 35.9s
Requests
Jul 28max 1,700Aug 26
Audio Stem Separation 1.0Avg 103/day
Generation Time
Jul 28max 51.9sAug 26
Audio Stem Separation 1.0Avg 27.6s
Error Rate
Jul 28max 16.7%Aug 26
Audio Stem Separation 1.0Avg 6.5%

〰 Uptime

Percent of generations that succeeded over the selected period, per model version.

Avg. Success Rate (30d)
93.55%
across all generations of this model family
Uptime
Jul 28max 100%Aug 26
SuccessfulAvg 93.55%