How AI Audio Generation Models Are Ranked: Inside the Pixazo Leaderboard
Audio is uniquely hard to rank, because you cannot glance at a waveform and know whether a song slaps or a voice sounds human. Judging sound means listening, and human listeners get tired, distracted and biased within minutes. So the Pixazo AI audio leaderboard replaces the crowd with something steadier: a panel of free, open audio scorers that grade every clip the same way, every run. This is how it works, shown with the same graphs the live board uses.
Why a single ear is not enough
Two problems break naive audio benchmarking. First, quality is multi-dimensional: a speech clip can sound natural yet mangle a word, and a song can have a great melody over a muddy mix. Second, human preference drifts, so the same listener can rank the same two clips differently an hour apart. A benchmark built on one ear, or on a handful of crowd votes, inherits both problems. The fix is to score each clip with several specialised models whose judgements are fixed and repeatable, and to average over many prompts so no single clip decides a rank. Because those scorers return the same numbers on the same clips every run, a move on the board means the models changed, not the test.
Reproducibility is the real goal. A crowd vote can never be run again the same way, so when a crowd-sourced board moves you cannot tell whether the models improved or the voters changed. By delegating scoring to open models that are deterministic on a given clip, the benchmark makes movement meaningful, and it lets the board be re-run cleanly whenever a new audio model launches on Pixazo.
Suggested Read: Best AI Audio Generation Models in 2026: A Comparison Guide
The judges: an ensemble, weighted
Different jobs need different judges, so each track is graded by an ensemble tuned to it, and each judge carries a weight set by how much it contributes to the final call. Here is the music board’s three-judge ensemble:
MuQ-Eval estimates music quality the way a mean-opinion-score panel would; CLAP measures how well the audio matches the text prompt; and Audiobox-PQ rates production quality. Speech swaps in UTMOS for naturalness and Whisper transcription to check the words came out right; song uses four judges, adding Whisper to grade the lyrics. Using an ensemble rather than a single metric means one scorer’s blind spot is covered by the others.
The number of judges is matched to the job, not fixed. Music needs three because it has to sound good, follow the prompt, and be cleanly produced. Song needs four because it must also get the lyrics right, so Whisper is added to transcribe the vocal and check it against the words. Speech leans on UTMOS and Whisper because a narrator that sounds lovely but drops words is useless. The benchmark uses as many scorers as a track genuinely needs and no more, which keeps each board fast to re-run.
Suggested Read: Best AI Music Generators in 2026: 10 Tools Ranked
Why the ensemble matters: the judges disagree
This is the clearest argument against trusting any single scorer. Plot how the music judges rank the same models against each other and they frequently pull in opposite directions:
| MuQ-Eval | CLAP | Audiobox | |
| MuQ-Eval | 1.00 | -0.60 | 0.84 |
| CLAP | -0.60 | 1.00 | -0.72 |
| Audiobox | 0.84 | -0.72 | 1.00 |
+1 (lime) = the judges rank models alike; -1 (red) = they disagree.
MuQ-Eval and Audiobox-PQ agree strongly (0.84), but CLAP moves against both (-0.60 and -0.72): the tracks that follow a prompt most literally are often not the ones that sound best produced. If the board trusted only CLAP it would crown obedient but dull tracks; if it trusted only MuQ-Eval it would reward great-sounding clips that ignored the brief. Weighting all three, and requiring a model to satisfy them together, is what makes the ranking hard to game.
This is also why a model cannot climb by optimising for one scorer. Because the judges disagree, gaming CLAP alone would drag down MuQ-Eval and Audiobox, so the only way up the board is to genuinely improve on all the axes at once. The ensemble turns a set of imperfect, arguable metrics into a robust verdict that no single trick can spoof.
Suggested Read: Best Text To Speech APIs in 2026
Which judge favours which model
Zoom in from “do the judges agree” to “who does each judge actually prefer” and the disagreement becomes concrete. This grid is each scorer’s normalised preference for every model, where 1 is that judge’s favourite and 0 its least:
| Model Judge | MuQ-Eval | CLAP | Audiobox-PQ |
| Lyria 3 | 0.88 | 1.00 | 0.63 |
| Lyria 2 | 1.00 | 0.00 | 1.00 |
| Lyria 3 Pro | 0.29 | 0.83 | 0.67 |
| Stable Audio | 0.19 | 0.88 | 0.39 |
| ElevenLabs Music | 0.00 | 0.96 | 0.00 |
Each judge’s normalised preference for a model (1 = this judge’s favourite, 0 = its least). Rows disagree = the judges see different winners.
Read down a column and you see one judge’s taste; read across a row and you see how divided the judges are about a single model. Lyria 2 is the clearest case: MuQ-Eval and Audiobox-PQ both rate it top (1.00), yet CLAP puts it dead last (0.00), because it sounds superb but drifts from the prompt. The same three axes drawn as a radar make each model’s shape obvious at a glance:
A balanced model fills the triangle evenly; a lopsided one, like Lyria 2 with its long pull toward MuQ-Eval and Audiobox but almost nothing on CLAP, shows exactly where it is strong and where it cuts corners. The weighted ensemble is what reconciles these competing tastes into a single defensible rank, instead of letting any one judge crown its personal favourite.
From clips to a ranking: arena Elo
Scores alone are not a ranking, so the benchmark converts them into head-to-head outcomes. On the same prompt groups, each model’s clip is compared against its rivals, the ensemble decides the winner of each matchup, and those wins become a per-group Elo with a Bradley-Terry cross-check. Averaging the per-group ratings gives the mean arena Elo, and every rating carries a bootstrap 95% confidence interval so you can see when a gap is real. Ratings do not appear fully formed, either; they settle as more matchups accumulate:
Illustrative: ratings start spread and settle toward their final Elo as more matchups accumulate.
Early on the ratings swing wildly, then converge as each model plays more matches, which is why the benchmark runs a fixed, sizeable prompt set rather than a handful of clips. Computing Elo per group and then averaging stops a model from buying a high rank with a few standout clips, and the confidence interval keeps you honest about which gaps actually matter: on the music board Lyria 3 and Lyria 2 end up nine Elo apart with fully overlapping ranges, which is the board’s way of saying they are effectively tied. The settled standings for every track are shown on the companion comparison guide.
The Bradley-Terry cross-check adds a second layer of confidence. Bradley-Terry is a statistical model built precisely for turning many pairwise win-and-loss results into one rating per competitor, and running it alongside the raw ensemble tally catches cases where a model beats weak rivals often but folds against strong ones. When both methods agree on the order, the ranking is trustworthy; when they disagree, the field is genuinely tight rather than cleanly separated, which is itself useful to know.
Suggested Read: Best Audio Generation APIs in 2026
The value view: quality per dollar
Sound is generated at scale, so raw quality is only half the decision. The live board also records the public per-clip price on Pixazo’s gateway, for a typical clip of roughly forty-five seconds of music or song or fifteen seconds of speech, and a value view divides mean Elo by that price. That surfaces the models that deliver the most quality per dollar, which is often not the outright quality leader. For a long audiobook or a batch of ad reads, the value column can matter more than the top of the board. A value view is only fair if the price is measured consistently, so the board fixes the clip length it prices and uses the public gateway price rather than any promotional rate, which keeps the comparison apples to apples. The value ranking can look very different from the quality ranking, and that difference is the point: the model that wins a single showcase piece is rarely the one you want billing you for ten thousand clips, so a near-leader that costs a fraction as much can be the smarter production choice. Reporting both lets a studio decide for itself rather than defaulting to the name at the top.
Suggested Read: Top 17 AI Voice Generators (Text-To-Speech) in 2026
The guardrails
Three rules keep the benchmark honest. Every model, including Pixazo’s own Tracks, runs the identical prompt groups and is scored by the identical ensembles, at its real position with no adjustment, which is why Tracks sits second on the song board rather than first. Scoring is delegated to free, reproducible models rather than crowd voting, so no single listener’s ear or fatigue can tilt a rank. And Pixazo hosts these models but does not own them: each provider is named beside its model, and the clearest proof of fairness is that the board publishes its own model losing to a competitor.
It is worth being explicit about why that last point matters. A benchmark built to sell a product would never show its own model in second place, so the fact that Tracks sits behind ACE-Step XL on the song board, at its real Elo with its real confidence interval, is the strongest signal that these numbers are read off the ensembles rather than written by hand. The same machinery that ranks Tracks ranks every rival, with no thumb on the scale.
See the board for yourself
Open the live leaderboard to switch between the music, song and speech tracks, sort by Elo or by value, and open any model for its per-track profile. Every number in this article is pulled straight from that board, and because the whole run is reproducible, it will read the same when you check it as it did when this was written.
Frequently asked questions
Is the ranking judged by humans or AI?
By AI. Pixazo’s research team writes the prompt groups, generates every clip and reviews the results, but the scoring itself is delegated to free, reproducible open audio models (MuQ-Eval, CLAP, Audiobox, UTMOS and Whisper), not to crowd voting.
Why use an ensemble of scorers instead of one?
Because audio quality is multi-dimensional and the scorers often disagree. On music, CLAP correlates negatively with MuQ-Eval and Audiobox, so relying on any one of them would reward the wrong thing. Weighting several covers each one’s blind spot.
What does the 95% CI on the board mean?
It is a bootstrap confidence interval over the prompt groups. When two models’ intervals overlap, the Elo gap between them is not statistically meaningful and they should be read as tied.
Why score three separate tracks?
Because music, song and speech are different skills that need different judges. Each track has its own board and its own ensemble, tuned to that job.
Is Pixazo’s Tracks model given any advantage?
No. It runs the same prompts and the same scoring as every rival and holds second place on the song board, behind ACE-Step XL. If it were favoured, it would not be losing.

Deepak Joshi
Author · Pixazo
Deepak writes about generative AI models, APIs, and the workflows teams use to ship them. Reviewed by Abhinav Girdhar.