7 AI Model Leaderboards: Latest Rankings and What They Measure
No single score can tell you which AI model is best. Some models excel at conversation, some at coding, and others at completing long workflows across multiple apps. These seven leaderboards measure different parts of that picture.
Rankings were collected on September 2, 2026. Leaderboards change quickly. Compare scores only within the same benchmark, version, and configuration.
LMArena maintains separate Text and WebDev rankings, so seven sources produce eight ranking tables. Artificial Analysis also has a multilingual leaderboard, covered below.
1. Artificial Analysis: overall capability
Artificial Analysis combines reasoning, knowledge, coding, instruction following, and agentic evaluations into one Intelligence Index. It is useful for a quick view of a model's overall ceiling, but it does not represent price, speed, or performance on your exact workflow.
Shown: all 28 valid configurations.
| Rank | Model | Provider | Source score |
|---|---|---|---|
| #1 | Anthropic | 65.7 | |
| #2 | Anthropic | 64.8 | |
| #3 | Anthropic | 63.1 | |
| #4 | Anthropic | 62.5 | |
| #5 | Anthropic | 62.5 | |
| #6 | Anthropic | 62.1 | |
| #7 | OpenAI | 60.9 | |
| #8 | SpaceXAI | 60.9 | |
| #9 | Kimi | 59.7 | |
| #10 | Z AI | 59.5 | |
| #11 | Alibaba | 57.7 | |
| #12 | Z AI | 57.5 | |
| #13 | Meta | 56.8 | |
| #14 | OpenAI | 56.6 | |
| #15 | 56.0 | ||
| #16 | DeepSeek | 53.2 | |
| #17 | OpenAI | 52.3 | |
| #18 | Alibaba | 52.0 | |
| #19 | MiniMax | 45.4 | |
| #20 | Thinking Machines | 42.3 | |
| #21 | NVIDIA | 38.3 | |
| #22 | 37.4 | ||
| #23 | Meta | 35.1 | |
| #24 | Mistral | 30.4 | |
| #25 | Anthropic | 29.9 | |
| #26 | OpenAI | 24.1 | |
| #27 | NVIDIA | 23.6 | |
| #28 | Cohere | 22.8 |
Its multilingual leaderboard is more useful when language performance matters. The current Chinese top three are Claude Opus 4.6 Max, Gemini 3.1 Pro Preview, and Gemini 3 Pro Preview High, all scoring 94.
View the Intelligence Index · View the multilingual leaderboard
2. LiveBench: an objective test that keeps changing
LiveBench covers mathematics, reasoning, coding, data analysis, language, and instruction following. It regularly refreshes its questions and uses objective grading, reducing the advantage of models that may have seen old benchmark questions during training.
It is useful for comparing core problem-solving ability. The current public snapshot contains 23 tasks across seven categories and was released on June 25, 2026.
Shown: Top 50 of 51 configurations.
| Rank | Model | Provider | Source score |
|---|---|---|---|
| #1 | Anthropic | 83.414 | |
| #2 | Anthropic | 82.971 | |
| #3 | OpenAI | 81.054 | |
| #4 | OpenAI | 80.188 | |
| #5 | Anthropic | 80.085 | |
| #6 | Other | 79.508 | |
| #7 | Moonshot AI | 79.193 | |
| #8 | 78.826 | ||
| #9 | Alibaba | 78.461 | |
| #10 | xAI | 78.042 | |
| #11 | OpenAI | 77.972 | |
| #12 | Meta | 77.955 | |
| #13 | OpenAI | 77.936 | |
| #14 | DeepSeek | 77.436 | |
| #15 | 76.952 | ||
| #16 | DeepSeek | 76.758 | |
| #17 | Anthropic | 76.529 | |
| #18 | Anthropic | 76.224 | |
| #19 | Alibaba | 76.194 | |
| #20 | Z.ai | 76.138 | |
| #21 | Anthropic | 76.040 | |
| #22 | xAI | 75.773 | |
| #23 | Meta | 75.300 | |
| #24 | Alibaba | 75.269 | |
| #25 | 74.638 | ||
| #26 | OpenAI | 74.635 | |
| #27 | Anthropic | 74.520 | |
| #28 | DeepSeek | 74.171 | |
| #29 | OpenAI | 73.975 | |
| #30 | 73.588 | ||
| #31 | OpenAI | 73.559 | |
| #32 | Z.ai | 73.157 | |
| #33 | Alibaba | 73.137 | |
| #34 | Anthropic | 72.990 | |
| #35 | Anthropic | 72.582 | |
| #36 | Other | 71.922 | |
| #37 | Z.ai | 71.591 | |
| #38 | DeepSeek | 71.572 | |
| #39 | Moonshot AI | 70.539 | |
| #40 | OpenAI | 69.576 | |
| #41 | Other | 69.235 | |
| #42 | Alibaba | 68.906 | |
| #43 | Moonshot AI | 68.412 | |
| #44 | xAI | 67.781 | |
| #45 | MiniMax | 67.257 | |
| #46 | OpenAI | 66.374 | |
| #47 | DeepSeek | 65.481 | |
| #48 | NVIDIA | 64.155 | |
| #49 | Alibaba | 64.027 | |
| #50 | 63.937 |
3. LMArena: blind votes from real users
LMArena does not use a conventional exam. Users see two anonymous answers and vote for the better one. This is closer to real-world preference, although results are affected by the user population, response style, and number of votes.
- Text covers chat, writing, mathematics, coding, and other open-ended prompts.
- WebDev compares websites and front-end applications produced by models.
Shown: Text Top 50; all 20 WebDev labs.
| Board | Rank | Model | Provider | Arena score |
|---|---|---|---|---|
| Text | #1 | Anthropic | 1508±5 | |
| Text | #2 | Anthropic | 1505±4 | |
| Text | #3 | Anthropic | 1502±4 | |
| Text | #4 | Meta | 1499±10 | |
| Text | #5 | Anthropic | 1498±3 | |
| Text | #6 | Anthropic | 1494±4 | |
| Text | #7 | Anthropic | 1492±5 | |
| Text | #8 | Meta | 1491±5 | |
| Text | #9 | 1491±8Preliminary | ||
| Text | #10 | Moonshot | 1489±5 | |
| Text | #11 | Meta | 1488±6 | |
| Text | #12 | Anthropic | 1488±6 | |
| Text | #13 | 1487±3 | ||
| Text | #14 | 1486±4 | ||
| Text | #15 | OpenAI | 1483±5 | |
| Text | #16 | OpenAI | 1482±4 | |
| Text | #17 | Z.ai | 1482±7 | |
| Text | #18 | Anthropic | 1482±4 | |
| Text | #19 | 1480±5 | ||
| Text | #20 | Alibaba | 1479±6 | |
| Text | #21 | 1479±4 | ||
| Text | #22 | OpenAI | 1477±4 | |
| Text | #23 | OpenAI | 1477±4 | |
| Text | #24 | OpenAI | 1476±4 | |
| Text | #25 | 1475±5 | ||
| Text | #26 | SpaceXAI | 1475±5 | |
| Text | #27 | 1474±4 | ||
| Text | #28 | Alibaba | 1474±10 | |
| Text | #29 | OpenAI | 1474±5 | |
| Text | #30 | Z.ai | 1473±9 | |
| Text | #31 | Anthropic | 1473±4 | |
| Text | #32 | Anthropic | 1473±4 | |
| Text | #33 | Anthropic | 1472±4 | |
| Text | #34 | SpaceXAI | 1472±4 | |
| Text | #35 | Z.ai | 1472±5 | |
| Text | #36 | SpaceXAI | 1470±4 | |
| Text | #37 | SpaceXAI | 1470±5 | |
| Text | #38 | Anthropic | 1469±3 | |
| Text | #39 | Xiaomi | 1468±4 | |
| Text | #40 | Baidu | 1468±5 | |
| Text | #41 | OpenAI | 1466±5 | |
| Text | #42 | Z.ai | 1466±4 | |
| Text | #43 | OpenAI | 1466±4 | |
| Text | #44 | Alibaba | 1466±5 | |
| Text | #45 | SpaceXAI | 1465±3 | |
| Text | #46 | Anthropic | 1462±5 | |
| Text | #47 | SpaceXAI | 1461±10Preliminary | |
| Text | #48 | Moonshot | 1461±5 | |
| Text | #49 | Alibaba | 1460±8 | |
| Text | #50 | DeepSeek | 1459±8 | |
| WebDev | #1 | Alibaba | 1691+19/-19Preliminary | |
| WebDev | #2 | Anthropic | 1687+8/-8 | |
| WebDev | #3 | Moonshot AI | 1674+11/-11 | |
| WebDev | #4 | Other | 1629+17/-17Preliminary | |
| WebDev | #5 | Tencent | 1627+17/-17Preliminary | |
| WebDev | #6 | OpenAI | 1616+7/-7 | |
| WebDev | #7 | Other | 1609+13/-13 | |
| WebDev | #8 | 1587+12/-12Preliminary | ||
| WebDev | #9 | DeepSeek | 1583+11/-11 | |
| WebDev | #10 | Other | 1540+8/-8 | |
| WebDev | #11 | Other | 1522+8/-8 | |
| WebDev | #12 | Other | 1487+7/-7 | |
| WebDev | #13 | Other | 1476+6/-6 | |
| WebDev | #14 | Other | 1408+8/-8 | |
| WebDev | #15 | Other | 1371+17/-17 | |
| WebDev | #16 | Other | 1347+10/-10 | |
| WebDev | #17 | Other | 1265+15/-15 | |
| WebDev | #18 | Other | 1254+20/-20 | |
| WebDev | #19 | IBM | 1192+19/-19 | |
| WebDev | #20 | Inception AI | 1166+25/-25 |
WebDev leader Qwen 3.8 Max 0902 is still marked Preliminary, so its position may move as more votes arrive. Confidence intervals for the top Text models also overlap heavily; a difference of a few points is not a decisive win.
View the Text leaderboard · View the WebDev leaderboard
4. APEX-Agents: long professional workflows
APEX-Agents tests whether an AI can work like an investment banking analyst, management consultant, or corporate lawyer. Tasks require research, analysis, document creation, multi-app operation, and planning over long workflows.
The benchmark contains 33 work environments and 480 tasks. It is more useful than a standard question-answer benchmark when you need an agent to finish complex professional work.
Shown: all 28 configurations under the default Mean Score + Loop view.
| Rank | Model | Provider | Mean Score |
|---|---|---|---|
| #1 | Anthropic | 62.0% ± 3.5% | |
| #2 | Anthropic | 60.6% ± 3.6% | |
| #3 | Anthropic | 60.0% ± 3.5% | |
| #4 | Anthropic | 59.2% ± 3.7% | |
| #5 | Meta | 58.1% ± 3.4% | |
| #6 | xAI | 57.5% ± 3.5% | |
| #7 | OpenAI | 56.7% ± 3.3% | |
| #8 | OpenAI | 56.4% ± 3.4% | |
| #9 | Anthropic | 56.2% ± 3.5% | |
| #10 | OpenAI | 55.5% ± 3.6% | |
| #11 | Kimi | 55.4% ± 3.3% | |
| #12 | OpenAI | 54.3% ± 3.3% | |
| #13 | Zhipu | 52.2% ± 3.5% | |
| #14 | DeepSeek | 51.6% ± 3.4% | |
| #15 | OpenAI | 51.3% ± 3.6% | |
| #16 | Anthropic | 50.1% ± 3.5% | |
| #17 | Anthropic | 48.7% ± 3.5% | |
| #18 | 48.7% ± 3.3% | ||
| #19 | Anthropic | 48.5% ± 3.2% | |
| #20 | OpenAI | 47.8% ± 3.7% | |
| #21 | xAI | 47.1% ± 3.4% | |
| #22 | 46.7% ± 3.4% | ||
| #23 | OpenAI | 44.6% ± 3.4% | |
| #24 | OpenAI | 42.5% ± 3.2% | |
| #25 | Kimi | 41.9% ± 3.0% | |
| #26 | Applied Compute | 40.1% ± 3.2% | |
| #27 | Anthropic | 37.1% ± 3.2% | |
| #28 | Zhipu | 21.2% ± 2.2% |
5. Vals Finance Agent: financial analyst tasks
This leaderboard focuses on financial analysis, including researching companies and SEC filings, forming conclusions, and producing evidence-based answers. It is useful for selecting a financial research assistant, not for judging general chat or coding ability.
This table uses the Finance Agent v1.1 Accuracy data currently published on the page, updated June 4, 2026. A leading score near 64% still shows that complex financial research is far from fully reliable.
Shown: Top 50 of 51 models.
| Rank | Model | Provider | Accuracy |
|---|---|---|---|
| #1 | Anthropic | 64.373% | |
| #2 | Anthropic | 63.331% | |
| #3 | Meta | 60.595% | |
| #4 | DeepSeek | 60.389% | |
| #5 | Anthropic | 60.046% | |
| #6 | OpenAI | 59.963% | |
| #7 | 59.717% | ||
| #8 | Anthropic | 58.810% | |
| #9 | OpenAI | 58.535% | |
| #10 | Z.ai | 57.655% | |
| #11 | OpenAI | 57.152% | |
| #12 | Moonshot AI | 57.056% | |
| #13 | OpenAI | 55.309% | |
| #14 | 55.154% | ||
| #15 | Alibaba | 54.627% | |
| #16 | Anthropic | 54.500% | |
| #17 | Alibaba | 54.475% | |
| #18 | Grok | 53.812% | |
| #19 | Grok | 53.506% | |
| #20 | OpenAI | 53.405% | |
| #21 | Z.ai | 53.182% | |
| #22 | Alibaba | 52.785% | |
| #23 | Grok | 52.448% | |
| #24 | Grok | 52.295% | |
| #25 | OpenAI | 52.151% | |
| #26 | OpenAI | 51.928% | |
| #27 | 50.788% | ||
| #28 | Moonshot AI | 50.622% | |
| #29 | MiniMax | 48.402% | |
| #30 | OpenAI | 47.801% | |
| #31 | 47.598% | ||
| #32 | Anthropic | 46.931% | |
| #33 | 46.123% | ||
| #34 | Mistralai | 46.113% | |
| #35 | Grok | 46.084% | |
| #36 | Z.ai | 45.977% | |
| #37 | Alibaba | 45.639% | |
| #38 | Grok | 44.362% | |
| #39 | Alibaba | 44.295% | |
| #40 | 41.589% | ||
| #41 | MiniMax | 38.579% | |
| #42 | Moonshot AI | 36.647% | |
| #43 | Z.ai | 36.480% | |
| #44 | MiniMax | 33.350% | |
| #45 | Fireworks | 21.541% | |
| #46 | Mistralai | 18.049% | |
| #47 | OpenAI | 8.064% | |
| #48 | Cohere | 4.226% | |
| #49 | Fireworks | 2.345% | |
| #50 | AI21 Labs | 0.370% |
6. DeepSWE v1.1: fixing real repositories
DeepSWE gives agents original, long-running software engineering tasks. A model must understand a repository, edit the code, submit a patch, and pass tests inside a clean verification container.
Version 1.1 contains 113 tasks, and the current snapshot was updated on August 26, 2026. Models run through the same agent framework, making comparison easier than with scattered vendor-reported results.
Shown: all 19 models in the Best view.
| Rank | Model | Provider | Pass@1 |
|---|---|---|---|
| #1 | Anthropic | 74% ± 4% | |
| #2 | OpenAI | 73% ± 3% | |
| #3 | Anthropic | 70% ± 4% | |
| #4 | ZAI.png | 69% ± 3% | |
| #5 | Moonshot.png | 69% ± 5% | |
| #6 | OpenAI | 67% ± 4% | |
| #7 | OpenAI | 67% ± 6% | |
| #8 | xAI | 67% ± 2% | |
| #9 | 65% ± 2% | ||
| #10 | ZAI.png | 63% ± 4% | |
| #11 | DeepSeek | 63% ± 6% | |
| #12 | Anthropic | 59% ± 2% | |
| #13 | Qwen.png | 57% ± 3% | |
| #14 | Meta | 55% ± 2% | |
| #15 | Anthropic | 54% ± 4% | |
| #16 | DeepSeek | 53% ± 4% | |
| #17 | 47% ± 4% | ||
| #18 | ZAI.png | 44% ± 2% | |
| #19 | 36% ± 4% |
7. TapTap Maker: making games that actually run
TapTap Maker asks AI agents to write Lua game code for the UrhoX engine. A real engine executes and scores the result; another language model does not act as the judge.
The current ranking uses snapshot 2026-08-27-v5. The top scores are close, and the confidence intervals for second and third place are wide, so the leaders are better viewed as one top tier.
Shown: all 29 configurations.
| Rank | Model | Provider | Source score |
|---|---|---|---|
| #1 | Anthropic | 89.02% | |
| #2 | Anthropic | 88.89% | |
| #3 | xAI | 87.30% | |
| #4 | OpenAI | 86.51% | |
| #5 | Anthropic | 86.51% | |
| #6 | 84.13% | ||
| #7 | xAI | 84.13% | |
| #8 | Z.ai | 83.60% | |
| #9 | Anthropic | 82.76% | |
| #10 | Alibaba | 82.54% | |
| #11 | OpenAI | 81.75% | |
| #12 | Alibaba | 81.48% | |
| #13 | Z.ai | 81.48% | |
| #14 | Moonshot AI | 79.72% | |
| #15 | xAI | 78.31% | |
| #16 | OpenAI | 77.78% | |
| #17 | DeepSeek | 77.78% | |
| #18 | 76.19% | ||
| #19 | Anthropic | 76.19% | |
| #20 | DeepSeek | 74.60% | |
| #21 | Z.ai | 74.34% | |
| #22 | Anthropic | 73.28% | |
| #23 | Tencent | 73.02% | |
| #24 | ByteDance | 72.75% | |
| #25 | Alibaba | 68.25% | |
| #26 | 64.70% | ||
| #27 | 63.89% | ||
| #28 | MiniMax | 60.05% | |
| #29 | Anthropic | 43.65% |
Which leaderboard should you use?
- For overall capability, use Artificial Analysis and LiveBench.
- For real-user preference, use LMArena Text.
- For web and front-end development, use LMArena WebDev.
- For long professional workflows, use APEX-Agents.
- For financial research, use Vals Finance Agent.
- For software engineering, use DeepSWE.
- For game development, use TapTap Maker.
The practical approach is not to chase one overall ranking. Start with your task, consult the relevant leaderboard, then run a small test using your own work. A leaderboard narrows the field; your real workflow makes the final choice.