Pixazo Research · Generative-AI Benchmarks

The AI Model Leaderboard

Independent benchmarks of the leading generative models across image, video, audio and 3D. Every model is run by our research team and scored by automated ensembles under one fixed rubric — a controlled, reproducible complement to crowd-voted human-preference arenas.

4Categories
48Models ranked
10Benchmark tracks
EloBradley–Terry ranked

Explore a benchmark

Four independent leaderboards. Open any category for its full interactive ranking, all tracks, and the per-category method.

1
GPT Image 2.5
OpenAI
1330Elo
2
gpt-image-2 (medium)
OpenAI
1324Elo
3
Grok Imagine Image 2.0
xAI
1322Elo
Tight race — near-tied at the topView full leaderboard →
1
GPT Image 2.5
OpenAI
1330Elo
2
gpt-image-2 (medium)
OpenAI
1310Elo
3
MAI-Image 2.6
Microsoft
1308Elo
Tight race — near-tied at the topView full leaderboard →
1
Gemini Omni Flash
Google
1212Elo
2
Wan 3.0
Alibaba
1210Elo
3
flux-3-video
Black Forest Labs
1208Elo
#1 leads by 121 EloView full leaderboard →
1
Gemini Omni Flash
Google
1183Elo
2
Seedance 2.0
ByteDance
1180Elo
3
Grok Imagine 1.5
xAI
1067Elo
#1 leads by 116 EloView full leaderboard →
1
Happy Horse 1.0
Alibaba
1200Elo
2
Gemini Omni Flash
Google
1134Elo
3
Kling 3.0 Omni
Kuaishou
1000Elo
#1 leads by 200 EloView full leaderboard →
1
Lyria 3
Google
1036Elo
2
Lyria 2
Google
1027Elo
3
Lyria 3 Pro
Google
985Elo
#1 leads by 51 EloView full leaderboard →
1
ACE-Step XL
ACE Studio
1029Elo
2
Tracks
Pixazo
988Elo
3
ACE-Step
ACE Studio
984Elo
#1 leads by 45 EloView full leaderboard →
1
ElevenLabs v3
ElevenLabs
1078Elo
2
Qwen3-TTS
Alibaba
1046Elo
3
VibeVoice
Microsoft
999Elo
#1 leads by 79 EloView full leaderboard →
1
Hunyuan 3D Pro
Tencent
1451Elo
2
Hyper3D Rodin 2
Deemos
1409Elo
3
Meshy 6 Preview
Meshy
1289Elo
#1 leads by 162 EloView full leaderboard →
1
Hunyuan 3D Pro
Tencent
1451Elo
2
Hyper3D Rodin 2
Deemos
1409Elo
3
Meshy 6 Preview
Meshy
1289Elo
#1 leads by 162 EloView full leaderboard →

How decisive is each race?

The Elo gap between the #1 and #3 model on each board. A ~100-point gap means the leader wins about 64% of head-to-head matches — so a short bar is a photo-finish and a long bar is a clear front-runner. Gaps are comparable across benchmarks even though the raw scores are not.

How the models are ranked

The same method underpins every board — automated, reproducible, and calibrated within each benchmark.

Run by our research team

Every model generates the same prompt groups. Pixazo's team curates the prompts, runs each model, and reviews the outputs — the benchmark is re-run whenever a major model launches.

Scored by automated ensembles

Each output is judged by a fixed ensemble of open scorers, not by public voting — so a re-run gives the same result and no single rater's bias can tilt a rank.

Per-group Elo + Bradley–Terry

Outputs are matched within each prompt group, converted to Elo, and cross-checked with a Bradley–Terry fit. Overlapping intervals are ties — read the tier, not the exact order.

Calibrated within a benchmark

An image Elo and an audio Elo are not comparable — each benchmark has its own scale and judges. Compare models inside a category, a controlled complement to crowd-voted arenas.

CategoryScoring methodModelsJudge ensemble
AI Image Generation6-model vision-language ensemble15PickScore · HPSv2 · ImageReward · VQAScore · VLM · CLIP
AI Video GenerationMotion-aware VLM panel10Qwen2.5-VL on the clips · RAFT optical flow
AI Audio GenerationAudio scorer ensembles14MuQ-Eval · CLAP · UTMOS · Whisper WER · Audiobox
AI 3D Model GenerationMulti-view VLM rubric7Qwen2.5-VL-32B rubric · trimesh/pymeshlab mesh hygiene

Who leads across the board

Models benchmarked by provider, counted straight from each board's own ranking — 37 models across the four boards' primary tracks.

Google9
Alibaba3
Black Forest Labs2
ByteDance2
Deemos2
Kuaishou2
Microsoft2
OpenAI2
Current track leaders
Benchmark trackCurrent leaderElo
Image · Text-to-ImageGPT Image 2.5 OpenAI1330
Image · Image EditingGPT Image 2.5 OpenAI1330
Video · Text-to-VideoGemini Omni Flash Google1212
Video · Image-to-VideoGemini Omni Flash Google1183
Video · Video-to-VideoHappy Horse 1.0 Alibaba1200
Audio · Text-to-MusicLyria 3 Google1036
Audio · Text-to-SongACE-Step XL ACE Studio1029
Audio · Text-to-SpeechElevenLabs v3 ElevenLabs1078
3D · Text-to-3DHunyuan 3D Pro Tencent1451
3D · Image-to-3DHunyuan 3D Pro Tencent1451

Frequently asked questions

How are AI models ranked on this leaderboard?
Every model is run by Pixazo's research team on the same prompt groups, then scored by a fixed ensemble of open scorers under one rubric — a vision-language ensemble for image, a motion-aware panel for video, audio ensembles for music/song/speech, and a multi-view rubric plus mesh-hygiene checks for 3D. Ensemble wins become per-group Elo with a Bradley–Terry cross-check.
What does the Elo score mean?
Elo is a relative rating from pairwise comparisons. ~1000 is the per-benchmark baseline and a 100-point gap means the higher model wins about 64% of head-to-head matches within that benchmark. Scores are calibrated within a category, so compare models inside a category — not across them.
Are these rankings independent?
Yes. The scoring ensembles are automated and identical for every model, and the benchmark is re-run whenever a major model launches. Pixazo's own models compete under the same rubric — Tracks sits second on the text-to-song board and Pixal 3D mid-table — shown at their real positions with no adjustment.
Which AI image generator is best in 2026?
OpenAI's GPT Image 2.5 leads text-to-image at 1330 mean Elo, with gpt-image-2 (medium) and Microsoft's MAI-Image 2.6 within a few points — the top of the board is a tight pack rather than a runaway.
Which AI video model is best in 2026?
Google's Gemini Omni Flash leads text-to-video at 1212 mean Elo, followed by Alibaba's Wan 3.0 and Black Forest Labs' flux-3-video.
Which AI music generator is best?
Google's Lyria 3 leads the text-to-music track at 1036 mean Elo, ahead of Lyria 2 and Lyria 3 Pro. Song and speech are ranked on their own tracks.
Which AI 3D model generator is best in 2026?
Tencent's Hunyuan 3D Pro leads at 1451 mean Elo, followed by Deemos' Hyper3D Rodin 2 and Meshy 6 Preview, across the text-to-3D and image-to-3D tracks.
Independent rankingsNamed human editorReviewed for accuracyNo paid placement
Deepak Joshi
Authored by ·Content Marketing Specialist, Pixazo·LinkedIn

10+ years in digital tools and AI products. Maintains the Pixazo Model Index and personally re-tests frontier model launches against standardized prompts.

Reviewed by · Founder & CEO, Pixazo · LinkedIn
Powered by Pixazo — your unified AI creative platform. Last updated July 2026.