Audio Stem Separation API - Vocal Remover & Stem Splitter: Pricing, Documentation
by Pixazo
Audio Stem Separation API splits a finished mix back into its parts. Two-stem mode returns vocals and no_vocals — the pairing dubbing and localisation need, because replacing the speech no longer destroys the music bed underneath it. Four-stem mode returns vocals, drums, bass and other, which is what remixing, sampling and karaoke production start from. Send an audio file or a video URL and the audio is extracted server-side, so you do not need a separate demux step. Billing is per second of input audio, so a three-minute track costs about ninety cents and you always know the price before you submit.

Models Version
Get $5 Free Credit on First Payment
No strings attached — add funds and get $5 bonus instantly
Audio Stem Separation 1.0 API Documentation
https://gateway.pixazo.ai/audio-tools/v1/separate-stems
Authentication
All requests require an API key passed via header.
Pricing: Billed at $0.005 per second of input audio — measured on the file you submit, not on the stems returned. A 60-second track costs $0.30. Accepts audio or video; for video the audio is extracted server-side and you are billed on its duration.
Rounding: billing is per second but always rounds up to a whole second — a 5.06-second source bills as 6 seconds. The shortest billable job is 1 second.
Retries: this is an asynchronous job on shared encoding capacity, so a request can occasionally come back processing_failed or take much longer than usual. These are transient and succeed on a retry, and a failed job is never charged — the wallet hold is released. If you chain these tools, retry a failed step rather than failing the whole pipeline.
| Header | Type | Required | Description |
|---|---|---|---|
| Ocp-Apim-Subscription-Key | string | Yes | Your API subscription key |
Audio Separate Stems generate request
Request Code
POST https://gateway.pixazo.ai/audio-tools/v1/separate-stems
Content-Type: application/json
Cache-Control: no-cache
Ocp-Apim-Subscription-Key: YOUR_API_KEY
{
"audio_url": "https://api-assets.pixazo.ai/media/audio-tools-example.mp3",
"stems": "two",
"output_format": "mp3"
}
import requests
url = "https://gateway.pixazo.ai/audio-tools/v1/separate-stems"
headers = {
"Content-Type": "application/json",
"Cache-Control": "no-cache",
"Ocp-Apim-Subscription-Key": "YOUR_API_KEY"
}
data = {
"audio_url": "https://api-assets.pixazo.ai/media/audio-tools-example.mp3",
"stems": "two",
"output_format": "mp3"
}
response = requests.post(url, json=data, headers=headers)
print(response.json())
const url = "https://gateway.pixazo.ai/audio-tools/v1/separate-stems";
const headers = {
"Content-Type": "application/json",
"Cache-Control": "no-cache",
"Ocp-Apim-Subscription-Key": "YOUR_API_KEY"
};
const data = {
"audio_url": "https://api-assets.pixazo.ai/media/audio-tools-example.mp3",
"stems": "two",
"output_format": "mp3"
};
fetch(url, {
method: "POST",
headers: headers,
body: JSON.stringify(data)
})
.then(response => response.json())
.then(data => console.log(data));
curl -X POST "https://gateway.pixazo.ai/audio-tools/v1/separate-stems" \
-H "Content-Type: application/json" \
-H "Cache-Control: no-cache" \
-H "Ocp-Apim-Subscription-Key: YOUR_API_KEY" \
--data-raw '{
"audio_url": "https://api-assets.pixazo.ai/media/audio-tools-example.mp3",
"stems": "two",
"output_format": "mp3"
}'
Output
{
"request_id": "audio-separate-stems_019dxxxx-xxxx-xxxx-xxxx-xxxxxxxxxxxx",
"status": "QUEUED",
"polling_url": "https://gateway.pixazo.ai/v2/requests/status/audio-separate-stems_019dxxxx-xxxx-xxxx-xxxx-xxxxxxxxxxxx"
}
Webhook (Optional)
Add the X-Webhook-URL header to your generate request to receive a POST callback instead of polling.
X-Webhook-URL: https://your-server.com/webhook/callback
Request Parameters - Audio Separate Stems generate request
| Parameter | Required | Type | Default | Allowed values / range | Description |
|---|---|---|---|---|---|
| audio_url | Yes | string | — | HTTP(S) URL; ≤ 500 MB | Publicly reachable HTTP(S) URL of an audio or video file. For video the audio track is extracted server-side. Maximum file size: 500 MB. |
| stems | No | enum | "two" | "two", "four" | How many stems to split the track into. "two" returns vocals + no_vocals (dialogue vs the music-and-effects bed — the dubbing case, and roughly 2x faster); "four" returns vocals, drums, bass and other. |
| output_format | No | enum | "mp3" | "mp3", "wav" | Encoding of every returned stem. "mp3" is 192 kbps; "wav" is uncompressed (lossless, larger files). |
Example Request
{
"audio_url": "https://api-assets.pixazo.ai/media/audio-tools-example.mp3",
"stems": "two",
"output_format": "mp3"
}
Response
{
"request_id": "audio-separate-stems_019dxxxx-xxxx-xxxx-xxxx-xxxxxxxxxxxx",
"status": "QUEUED",
"polling_url": "https://gateway.pixazo.ai/v2/requests/status/audio-separate-stems_019dxxxx-xxxx-xxxx-xxxx-xxxxxxxxxxxx"
}
Request Headers
| Header | Value |
|---|---|
| Content-Type | application/json |
| Cache-Control | no-cache |
| Ocp-Apim-Subscription-Key | YOUR_API_KEY |
Response Handling
Common status codes.
| Code | Meaning |
|---|---|
| 202 | Accepted — Request queued |
| 400 | Bad Request |
| 401 | Unauthorized |
| 402 | Insufficient Balance |
| 403 | Forbidden |
| 429 | Too Many Requests |
| 500 | Internal Server Error |
Error Responses
Queue system errors and model validation errors.
Queue System Errors
// 402 — Insufficient balance
{
"error": "Insufficient Balance",
"message": "Your wallet does not have enough balance."
}
// 400 — Model not found
{
"error": "Model not found",
"message": "Model 'audio-separate-stems' not found or is disabled"
}
Error via Status/Webhook
{
"request_id": "audio-separate-stems_019dxxxx-xxxx-xxxx-xxxx-xxxxxxxxxxxx",
"status": "ERROR",
"model_id": "audio-separate-stems",
"error": "Description of the error",
"output": null
}
Retrieving Results
Poll the universal status endpoint to check progress and retrieve results.
Endpoint
GET https://gateway.pixazo.ai/v2/requests/status/{request_id}
Ocp-Apim-Subscription-Key: YOUR_API_KEY
cURL Example
curl -H "Ocp-Apim-Subscription-Key: YOUR_API_KEY" \
"https://gateway.pixazo.ai/v2/requests/status/audio-separate-stems_019dxxxx-xxxx-xxxx-xxxx-xxxxxxxxxxxx"
Response (Completed)
{
"request_id": "audio-separate-stems_019dxxxx-xxxx-xxxx-xxxx-xxxxxxxxxxxx",
"status": "COMPLETED",
"model_id": "audio-separate-stems",
"error": null,
"output": {
"media_url": [
"https://pub-582b7213209642b9b995c96c95a30381.r2.dev/v1/audio-separate-stems_019dxxxx/vocals.mp3",
"https://pub-582b7213209642b9b995c96c95a30381.r2.dev/v1/audio-separate-stems_019dxxxx/no_vocals.mp3"
],
"media_type": "audio/mpeg"
},
"created_at": "2026-03-31T10:00:00.000Z",
"updated_at": "2026-03-31T10:00:15.000Z",
"completed_at": "2026-03-31T10:00:15.000Z"
}
Response Fields
| Field | Type | Description |
|---|---|---|
| request_id | string | Unique request identifier |
| status | string | QUEUED, PROCESSING, COMPLETED, FAILED, or ERROR |
| model_id | string | Model that processed the request |
| error | string|null | Error message if failed |
| output.media_url | array | URLs to generated media (R2 CDN) |
| output.media_type | string | MIME type of the output |
| created_at | string | When request was created |
| completed_at | string | When request completed |
| polling_url | string | Status URL (initial response only) |
Status Values
| Status | Description |
|---|---|
| QUEUED | Request accepted, waiting to be processed |
| PROCESSING | Being processed by the model |
| COMPLETED | Done — output contains the result |
| FAILED | Failed — check error field |
| ERROR | System error — not charged |
Status Flow
QUEUED → PROCESSING → COMPLETED
→ FAILED
→ ERROR
Typical Workflow
- Send a generate request to the API endpoint
- Save the
request_idfrom the response - Poll every 5-10 seconds:
GET /v2/requests/status/{request_id} - When
statusis"COMPLETED", download fromoutput.media_url
Tip: Use X-Webhook-URL header to get a callback instead of polling.
Audio Stem Separation 1.0 API Pricing
⚡ Performance
Live usage measured on Pixazo's gateway, split by model version. Generation time is how long a generation takes end-to-end (lower is better). Success rate is the percent of generations that complete (higher is better).
〰 Uptime
Percent of generations that succeeded over the selected period, per model version.