The AI Model Leaderboard
Independent benchmarks of the leading generative models across image, video, audio and 3D. Every model is run by our research team and scored by automated ensembles under one fixed rubric — a controlled, reproducible complement to crowd-voted human-preference arenas.
Explore a benchmark
Four independent leaderboards. Open any category for its full interactive ranking, all tracks, and the per-category method.
How decisive is each race?
The Elo gap between the #1 and #3 model on each board. A ~100-point gap means the leader wins about 64% of head-to-head matches — so a short bar is a photo-finish and a long bar is a clear front-runner. Gaps are comparable across benchmarks even though the raw scores are not.
How the models are ranked
The same method underpins every board — automated, reproducible, and calibrated within each benchmark.
Run by our research team
Every model generates the same prompt groups. Pixazo's team curates the prompts, runs each model, and reviews the outputs — the benchmark is re-run whenever a major model launches.
Scored by automated ensembles
Each output is judged by a fixed ensemble of open scorers, not by public voting — so a re-run gives the same result and no single rater's bias can tilt a rank.
Per-group Elo + Bradley–Terry
Outputs are matched within each prompt group, converted to Elo, and cross-checked with a Bradley–Terry fit. Overlapping intervals are ties — read the tier, not the exact order.
Calibrated within a benchmark
An image Elo and an audio Elo are not comparable — each benchmark has its own scale and judges. Compare models inside a category, a controlled complement to crowd-voted arenas.
| Category | Scoring method | Models | Judge ensemble |
|---|---|---|---|
| AI Image Generation | 6-model vision-language ensemble | 15 | PickScore · HPSv2 · ImageReward · VQAScore · VLM · CLIP |
| AI Video Generation | Motion-aware VLM panel | 10 | Qwen2.5-VL on the clips · RAFT optical flow |
| AI Audio Generation | Audio scorer ensembles | 14 | MuQ-Eval · CLAP · UTMOS · Whisper WER · Audiobox |
| AI 3D Model Generation | Multi-view VLM rubric | 7 | Qwen2.5-VL-32B rubric · trimesh/pymeshlab mesh hygiene |
Who leads across the board
Models benchmarked by provider, counted straight from each board's own ranking — 37 models across the four boards' primary tracks.
| Benchmark track | Current leader | Elo |
|---|---|---|
| Image · Text-to-Image | gpt-image-2 (medium) OpenAI | 1324 |
| Image · Image Editing | gpt-image-2 (medium) OpenAI | 1310 |
| Video · Text-to-Video | Seedance 2.0 ByteDance | 1212 |
| Video · Image-to-Video | Gemini Omni Flash Google | 1183 |
| Video · Video-to-Video | Happy Horse 1.0 Alibaba | 1200 |
| Audio · Text-to-Music | Lyria 3 Google | 1036 |
| Audio · Text-to-Song | ACE-Step XL ACE Studio | 1029 |
| Audio · Text-to-Speech | ElevenLabs v3 ElevenLabs | 1078 |
| 3D · Text-to-3D | Hunyuan 3D Pro Tencent | 1451 |
| 3D · Image-to-3D | Hunyuan 3D Pro Tencent | 1451 |