Snapshot 0 · 2026-06-30
One table. Eight categories. Which model leads on coding vs. reasoning vs. vision, all in one place.
Arena.ai rates models separately for each task type. Most leaderboard displays show one category at a time; this view merges all eight so you can see the full picture. 0 models captured today. Scores are Arena Elo (higher = better). A dash means the model has no votes in that category. Votes column shows overall vote count.
| Model | Org | overall | coding | math | reasoning | science | hard | instruction | vision | Votes |
|---|
The snapshot job hits arena.ai/leaderboard once a day per category (eight fetches total), merges the results into per-model rows, and commits a dated JSON file to data/leaderboards/. Only models with an overall score and at least three category scores are included. This page bundles the latest snapshot at deploy time; daily refresh means a redeploy.
Caveats: scores are human-preference Elo based on head-to-head comparisons, not benchmark prompts, which makes them harder to game but also slower to update for newly released models. Categories use independent voter pools so a model's coding Elo is not comparable in absolute terms to its vision Elo, only relative to other models in the same category. Low vote counts (under ~500) make scores less stable.