7 AI Model Leaderboards: Latest Rankings and What They Measure

2026-09-02 · AI · Benchmarks · 中文

No single score can tell you which AI model is best. Some models excel at conversation, some at coding, and others at completing long workflows across multiple apps. These seven leaderboards measure different parts of that picture.

Rankings were collected on September 2, 2026. Leaderboards change quickly. Compare scores only within the same benchmark, version, and configuration.

LMArena maintains separate Text and WebDev rankings, so seven sources produce eight ranking tables. Artificial Analysis also has a multilingual leaderboard, covered below.

1. Artificial Analysis: overall capability

Artificial Analysis combines reasoning, knowledge, coding, instruction following, and agentic evaluations into one Intelligence Index. It is useful for a quick view of a model's overall ceiling, but it does not represent price, speed, or performance on your exact workflow.

Shown: all 28 valid configurations.

Rank Model Provider Source score
#1 Claude Fable 5.1 (max with fallback) Anthropic 65.7
#2 Claude Fable 5.1 (xhigh with fallback) Anthropic 64.8
#3 Claude Opus 5 (max) Anthropic 63.1
#4 Claude Opus 5 (xhigh) Anthropic 62.5
#5 Claude Fable 5.1 (high with fallback) Anthropic 62.5
#6 Claude Fable 5 (with fallback) Anthropic 62.1
#7 GPT-5.6 Sol (max) OpenAI 60.9
#8 Grok 4.6 (high) SpaceXAI 60.9
#9 Kimi K3 (max) Kimi 59.7
#10 GLM-5.3 (max) Z AI 59.5
#11 Qwen3.8 2.4T A95B Alibaba 57.7
#12 GLM-5.3-Flash Z AI 57.5
#13 Muse Spark 1.2 (xhigh) Meta 56.8
#14 GPT-5.6 Terra (max) OpenAI 56.6
#15 Gemini 3.7 Flash (high) Google 56.0
#16 DeepSeek V4 Pro 0813 (max) DeepSeek 53.2
#17 GPT-5.6 Luna (max) OpenAI 52.3
#18 Qwen3.8 27B (xhigh) Alibaba 52.0
#19 MiniMax-M3 MiniMax 45.4
#20 Inkling Thinking Machines 42.3
#21 Nemotron 3 Ultra NVIDIA 38.3
#22 Gemini 3.5 Flash-Lite Google 37.4
#23 Muse Glimmer (high) Meta 35.1
#24 Mistral Medium 3.5 Mistral 30.4
#25 Claude 4.5 Haiku Anthropic 29.9
#26 gpt-oss-120b (high) OpenAI 24.1
#27 Nemotron 3.5 Lightning NVIDIA 23.6
#28 Command A+ Cohere 22.8

Its multilingual leaderboard is more useful when language performance matters. The current Chinese top three are Claude Opus 4.6 Max, Gemini 3.1 Pro Preview, and Gemini 3 Pro Preview High, all scoring 94.

View the Intelligence Index · View the multilingual leaderboard

2. LiveBench: an objective test that keeps changing

LiveBench covers mathematics, reasoning, coding, data analysis, language, and instruction following. It regularly refreshes its questions and uses objective grading, reducing the advantage of models that may have seen old benchmark questions during training.

It is useful for comparing core problem-solving ability. The current public snapshot contains 23 tasks across seven categories and was released on June 25, 2026.

Shown: Top 50 of 51 configurations.

Rank Model Provider Source score
#1 claude-fable-5-1-max-effort Anthropic 83.414
#2 claude-fable-5-max-effort Anthropic 82.971
#3 gpt-5.6-sol-max OpenAI 81.054
#4 gpt-5.5-xhigh OpenAI 80.188
#5 claude-opus-5-max-effort Anthropic 80.085
#6 smaug-agentic Other 79.508
#7 kimi-k3 Moonshot AI 79.193
#8 gemini-3.7-flash-high Google 78.826
#9 qwen3.8-max Alibaba 78.461
#10 grok-4.6 xAI 78.042
#11 gpt-5.4-xhigh OpenAI 77.972
#12 muse-spark-1.2-xhigh Meta 77.955
#13 gpt-5.6-terra-max OpenAI 77.936
#14 deepseek-v4-pro-0813 DeepSeek 77.436
#15 gemini-3.1-pro-preview-high Google 76.952
#16 deepseek-v4-flash-vision-exp DeepSeek 76.758
#17 claude-opus-4-7-xhigh-effort Anthropic 76.529
#18 claude-opus-4-8-max-effort Anthropic 76.224
#19 qwen3.8-flash-next Alibaba 76.194
#20 glm-5.3 Z.ai 76.138
#21 claude-sonnet-5-xhigh-effort Anthropic 76.040
#22 grok-4.5 xAI 75.773
#23 muse-spark-1.1-xhigh Meta 75.300
#24 qwen3.8-27b Alibaba 75.269
#25 gemini-3.5-flash-high Google 74.638
#26 gpt-5.2-2025-12-11-high OpenAI 74.635
#27 claude-opus-4-6-thinking-auto-high-effort Anthropic 74.520
#28 deepseek-v4-flash-0731 DeepSeek 74.171
#29 gpt-5.2-codex OpenAI 73.975
#30 gemini-3.6-flash-high Google 73.588
#31 gpt-5.6-luna-max OpenAI 73.559
#32 glm-5.2 Z.ai 73.157
#33 qwen3.7-max Alibaba 73.137
#34 claude-sonnet-4-6-thinking-auto-medium-effort Anthropic 72.990
#35 claude-opus-4-5-20251101-thinking-64k-high-effort Anthropic 72.582
#36 inkling-xhigh Other 71.922
#37 glm-5.3-flash Z.ai 71.591
#38 deepseek-v4-pro DeepSeek 71.572
#39 kimi-k2.6-thinking Moonshot AI 70.539
#40 gpt-5.4-nano-xhigh OpenAI 69.576
#41 ox-alpha-max Other 69.235
#42 qwen3.6-plus Alibaba 68.906
#43 kimi-k2.7-code Moonshot AI 68.412
#44 grok-build-0.1 xAI 67.781
#45 minimax-m3 MiniMax 67.257
#46 gpt-5.4-mini-xhigh OpenAI 66.374
#47 deepseek-v4-flash DeepSeek 65.481
#48 nemotron-3-ultra-550b-a55b NVIDIA 64.155
#49 qwen3.6-27b Alibaba 64.027
#50 gemini-3.5-flash-lite-high Google 63.937

View LiveBench

3. LMArena: blind votes from real users

LMArena does not use a conventional exam. Users see two anonymous answers and vote for the better one. This is closer to real-world preference, although results are affected by the user population, response style, and number of votes.

  • Text covers chat, writing, mathematics, coding, and other open-ended prompts.
  • WebDev compares websites and front-end applications produced by models.

Shown: Text Top 50; all 20 WebDev labs.

Board Rank Model Provider Arena score
Text #1 claude-fable-5 Anthropic 1508±5
Text #2 claude-opus-4-6-high Anthropic 1505±4
Text #3 claude-opus-4-7-high Anthropic 1502±4
Text #4 muse-spark-1.2 (xHigh) Meta 1499±10
Text #5 claude-opus-4-6 Anthropic 1498±3
Text #6 claude-opus-4-7 Anthropic 1494±4
Text #7 claude-opus-5-high Anthropic 1492±5
Text #8 muse-spark-1.1 Meta 1491±5
Text #9 gemini-3.7-flash-high Google 1491±8Preliminary
Text #10 kimi-k3-max Moonshot 1489±5
Text #11 muse-spark Meta 1488±6
Text #12 claude-opus-5-max Anthropic 1488±6
Text #13 gemini-3.1-pro-preview Google 1487±3
Text #14 gemini-3-pro Google 1486±4
Text #15 gpt-5.6-sol-xhigh OpenAI 1483±5
Text #16 gpt-5.5-high OpenAI 1482±4
Text #17 glm-5.3-max Z.ai 1482±7
Text #18 claude-opus-4-8-high Anthropic 1482±4
Text #19 gemini-3.6-flash-high Google 1480±5
Text #20 qwen3.8-max Alibaba 1479±6
Text #21 gemini-3.5-flash-high Google 1479±4
Text #22 gpt-5.5 OpenAI 1477±4
Text #23 gpt-5.4-high OpenAI 1477±4
Text #24 gpt-5.2-chat-latest-20260210 OpenAI 1476±4
Text #25 gemini-3.5-flash-medium Google 1475±5
Text #26 grok-4.20-beta1 SpaceXAI 1475±5
Text #27 gemini-3-flash Google 1474±4
Text #28 qwen3.7-max-preview Alibaba 1474±10
Text #29 gpt-5.5-instant OpenAI 1474±5
Text #30 glm-5.3-flash Z.ai 1473±9
Text #31 claude-opus-4-8 Anthropic 1473±4
Text #32 claude-opus-4-5-20251101-high-32k Anthropic 1473±4
Text #33 claude-sonnet-4-6 Anthropic 1472±4
Text #34 grok-4.20-beta-0309-reasoning SpaceXAI 1472±4
Text #35 glm-5.2-max Z.ai 1472±5
Text #36 grok-4.20-multi-agent-beta-0309 SpaceXAI 1470±4
Text #37 grok-4.5 SpaceXAI 1470±5
Text #38 claude-opus-4-5-20251101 Anthropic 1469±3
Text #39 mimo-v2.5-pro Xiaomi 1468±4
Text #40 ernie-5.1 Baidu 1468±5
Text #41 gpt-5.6-terra-xhigh OpenAI 1466±5
Text #42 glm-5.1 Z.ai 1466±4
Text #43 gpt-5.4 OpenAI 1466±4
Text #44 qwen3.5-max-preview Alibaba 1466±5
Text #45 grok-4.1-thinking SpaceXAI 1465±3
Text #46 claude-sonnet-5-high Anthropic 1462±5
Text #47 grok-4.6-high SpaceXAI 1461±10Preliminary
Text #48 kimi-k2.6 Moonshot 1461±5
Text #49 qwen3.6-max-preview Alibaba 1460±8
Text #50 deepseek-v4-pro-high-20260813 DeepSeek 1459±8
WebDev #1 qwen3.8-max-0902 Alibaba 1691+19/-19Preliminary
WebDev #2 claude-opus-5-max Anthropic 1687+8/-8
WebDev #3 kimi-k3-max Moonshot AI 1674+11/-11
WebDev #4 SpaceXAIgrok-4.6-high Other 1629+17/-17Preliminary
WebDev #5 Tencenthy4-preview Tencent 1627+17/-17Preliminary
WebDev #6 gpt-5.6-sol-xhigh (codex-harness) OpenAI 1616+7/-7
WebDev #7 Z.aiglm-5.3-max Other 1609+13/-13
WebDev #8 gemini-3.7-flash-high Google 1587+12/-12Preliminary
WebDev #9 deepseek-v4-pro-high-20260813 DeepSeek 1583+11/-11
WebDev #10 MetaMetamuse-spark-1.1 Other 1540+8/-8
WebDev #11 BytedanceBytedanceseed-2.1-pro-preview Other 1522+8/-8
WebDev #12 MiniMaxminimax-m3 Other 1487+7/-7
WebDev #13 Xiaomimimo-v2.5-pro Other 1476+6/-6
WebDev #14 Thinking MachinesThinkyinkling Other 1408+8/-8
WebDev #15 Upstagesolar-pro4 Other 1371+17/-17
WebDev #16 Poolsidelaguna-m.1 Other 1347+10/-10
WebDev #17 Mistralmistral-medium-3.5 Other 1265+15/-15
WebDev #18 KAT-Coder-Pro-V1 Other 1254+20/-20
WebDev #19 granite-4.1-8b IBM 1192+19/-19
WebDev #20 mercury-2 Inception AI 1166+25/-25

WebDev leader Qwen 3.8 Max 0902 is still marked Preliminary, so its position may move as more votes arrive. Confidence intervals for the top Text models also overlap heavily; a difference of a few points is not a decisive win.

View the Text leaderboard · View the WebDev leaderboard

4. APEX-Agents: long professional workflows

APEX-Agents tests whether an AI can work like an investment banking analyst, management consultant, or corporate lawyer. Tasks require research, analysis, document creation, multi-app operation, and planning over long workflows.

The benchmark contains 33 work environments and 480 tasks. It is more useful than a standard question-answer benchmark when you need an agent to finish complex professional work.

Shown: all 28 configurations under the default Mean Score + Loop view.

Rank Model Provider Mean Score
#1 Fable 5.1 Max Anthropic 62.0% ± 3.5%
#2 Opus 5 Max Anthropic 60.6% ± 3.6%
#3 Fable 5.1 High Anthropic 60.0% ± 3.5%
#4 Fable 5 Max Anthropic 59.2% ± 3.7%
#5 Muse Spark 1.1 Xhigh Meta 58.1% ± 3.4%
#6 Grok 4.6 High xAI 57.5% ± 3.5%
#7 GPT-5.6 Sol Max OpenAI 56.7% ± 3.3%
#8 GPT-5.6 Sol Max OpenAI 56.4% ± 3.4%
#9 Opus 4.8 Max Anthropic 56.2% ± 3.5%
#10 GPT-5.5 Xhigh OpenAI 55.5% ± 3.6%
#11 Kimi K3 Max Kimi 55.4% ± 3.3%
#12 GPT-5.6 Sol Xhigh OpenAI 54.3% ± 3.3%
#13 GLM-5.2 Zhipu 52.2% ± 3.5%
#14 DeepSeek-V4-Flash Max DeepSeek 51.6% ± 3.4%
#15 GPT-5.4 Xhigh OpenAI 51.3% ± 3.6%
#16 Opus 4.7 Max Anthropic 50.1% ± 3.5%
#17 Opus 4.6 High Anthropic 48.7% ± 3.5%
#18 Gemini 3.1 Pro High Google 48.7% ± 3.3%
#19 Sonnet 5 High Anthropic 48.5% ± 3.2%
#20 GPT-5.2 Xhigh OpenAI 47.8% ± 3.7%
#21 Grok 4.5 High xAI 47.1% ± 3.4%
#22 Gemini 3 Pro High Google 46.7% ± 3.4%
#23 GPT-5.3-Codex High OpenAI 44.6% ± 3.4%
#24 GPT-5.2-Codex High OpenAI 42.5% ± 3.2%
#25 Kimi K2.7 Code Auto Kimi 41.9% ± 3.0%
#26 Applied Compute: Small Applied Compute 40.1% ± 3.2%
#27 Opus 4.5 High Anthropic 37.1% ± 3.2%
#28 GLM-4.7 Zhipu 21.2% ± 2.2%

View APEX-Agents

5. Vals Finance Agent: financial analyst tasks

This leaderboard focuses on financial analysis, including researching companies and SEC filings, forming conclusions, and producing evidence-based answers. It is useful for selecting a financial research assistant, not for judging general chat or coding ability.

This table uses the Finance Agent v1.1 Accuracy data currently published on the page, updated June 4, 2026. A leading score near 64% still shows that complex financial research is far from fully reliable.

Shown: Top 50 of 51 models.

Rank Model Provider Accuracy
#1 claude-opus-4-7 Anthropic 64.373%
#2 claude-sonnet-4-6 Anthropic 63.331%
#3 muse-spark Meta 60.595%
#4 deepseek-v4-pro DeepSeek 60.389%
#5 claude-opus-4-6-thinking Anthropic 60.046%
#6 gpt-5.5 OpenAI 59.963%
#7 gemini-3.1-pro-preview Google 59.717%
#8 claude-opus-4-5-20251101-thinking Anthropic 58.810%
#9 gpt-5.2-2025-12-11 OpenAI 58.535%
#10 glm-5.1 Z.ai 57.655%
#11 gpt-5.4-2026-03-05 OpenAI 57.152%
#12 kimi-k2.6 Moonshot AI 57.056%
#13 gpt-5.1-2025-11-13 OpenAI 55.309%
#14 gemini-3-pro-preview Google 55.154%
#15 qwen3.6-plus Alibaba 54.627%
#16 claude-sonnet-4-5-20250929-thinking Anthropic 54.500%
#17 qwen3.5-plus-thinking Alibaba 54.475%
#18 grok-4.3 Grok 53.812%
#19 grok-4-0709 Grok 53.506%
#20 gpt-5.4-mini-2026-03-17 OpenAI 53.405%
#21 glm-5-thinking Z.ai 53.182%
#22 qwen3.6-max-preview Alibaba 52.785%
#23 grok-4-1-fast-reasoning Grok 52.448%
#24 grok-4.20-0309-reasoning Grok 52.295%
#25 gpt-5-2025-08-07 OpenAI 52.151%
#26 gpt-5-mini-2025-08-07 OpenAI 51.928%
#27 gemma-4-31b-it Google 50.788%
#28 kimi-k2.5-thinking Moonshot AI 50.622%
#29 MiniMax-M2.7 MiniMax 48.402%
#30 gpt-5.4-nano-2026-03-17 OpenAI 47.801%
#31 gemini-3-flash-preview Google 47.598%
#32 claude-haiku-4-5-20251001-thinking Anthropic 46.931%
#33 gemini-3.1-flash-lite-preview Google 46.123%
#34 mistral-medium-3.5 Mistralai 46.113%
#35 grok-4-fast-reasoning Grok 46.084%
#36 glm-4.7 Z.ai 45.977%
#37 qwen3.5-flash Alibaba 45.639%
#38 grok-4-1-fast-non-reasoning Grok 44.362%
#39 qwen3-max Alibaba 44.295%
#40 gemini-2.5-pro Google 41.589%
#41 MiniMax-M2.5 MiniMax 38.579%
#42 kimi-k2-thinking Moonshot AI 36.647%
#43 glm-4.6 Z.ai 36.480%
#44 MiniMax-M2.1 MiniMax 33.350%
#45 gpt-oss-120b Fireworks 21.541%
#46 mistral-large-2512 Mistralai 18.049%
#47 gpt-4o-2024-08-06 OpenAI 8.064%
#48 command-a-03-2025 Cohere 4.226%
#49 deepseek-v3p2-thinking Fireworks 2.345%
#50 jamba-large-1.7 AI21 Labs 0.370%

View Vals Finance Agent

6. DeepSWE v1.1: fixing real repositories

DeepSWE gives agents original, long-running software engineering tasks. A model must understand a repository, edit the code, submit a patch, and pass tests inside a clean verification container.

Version 1.1 contains 113 tasks, and the current snapshot was updated on August 26, 2026. Models run through the same agent framework, making comparison easier than with scattered vendor-reported results.

Shown: all 19 models in the Best view.

Rank Model Provider Pass@1
#1 claude-opus-5 max Anthropic 74% ± 4%
#2 gpt-5.6-sol max OpenAI 73% ± 3%
#3 claude-fable-5 max Anthropic 70% ± 4%
#4 glm-5.3 max ZAI.png 69% ± 3%
#5 kimi-k3 max Moonshot.png 69% ± 5%
#6 gpt-5.6-luna max OpenAI 67% ± 4%
#7 gpt-5.5 xhigh OpenAI 67% ± 6%
#8 grok-4.6 xhigh xAI 67% ± 2%
#9 gemini-3.7-flash high Google 65% ± 2%
#10 glm-5.3-flash max ZAI.png 63% ± 4%
#11 deepseek-v4-pro max DeepSeek 63% ± 6%
#12 claude-opus-4.8 max Anthropic 59% ± 2%
#13 qwen3.8-max xhigh Qwen.png 57% ± 3%
#14 muse-spark-1.2 xhigh Meta 55% ± 2%
#15 claude-sonnet-5 max Anthropic 54% ± 4%
#16 deepseek-v4-flash max DeepSeek 53% ± 4%
#17 gemini-3.6-flash high Google 47% ± 4%
#18 glm-5.2 max ZAI.png 44% ± 2%
#19 gemini-3.5-flash high Google 36% ± 4%

View DeepSWE v1.1

7. TapTap Maker: making games that actually run

TapTap Maker asks AI agents to write Lua game code for the UrhoX engine. A real engine executes and scores the result; another language model does not act as the judge.

The current ranking uses snapshot 2026-08-27-v5. The top scores are close, and the confidence intervals for second and third place are wide, so the leaders are better viewed as one top tier.

Shown: all 29 configurations.

Rank Model Provider Source score
#1 claude-fable-5 high Anthropic 89.02%
#2 claude-opus-5 xhigh Anthropic 88.89%
#3 grok-4.6 xhigh xAI 87.30%
#4 gpt-5.6-sol xhigh OpenAI 86.51%
#5 claude-opus-5 high Anthropic 86.51%
#6 gemini-3.7-flash high Google 84.13%
#7 grok-4.6 high xAI 84.13%
#8 glm-5.3 max Z.ai 83.60%
#9 claude-opus-4-8 xhigh Anthropic 82.76%
#10 qwen3.8-max xhigh Alibaba 82.54%
#11 gpt-5.6-terra xhigh OpenAI 81.75%
#12 qwen3.8-flash xhigh Alibaba 81.48%
#13 glm-5.3-flash max Z.ai 81.48%
#14 kimi-k3 Moonshot AI 79.72%
#15 x-ai/grok-4.5 high xAI 78.31%
#16 gpt-5.6-luna xhigh OpenAI 77.78%
#17 deepseek-v4-pro DeepSeek 77.78%
#18 gemini-3.6-flash Google 76.19%
#19 claude-opus-4-6 xhigh Anthropic 76.19%
#20 deepseek-v4-flash max DeepSeek 74.60%
#21 glm-5.2 Z.ai 74.34%
#22 claude-sonnet-5 xhigh Anthropic 73.28%
#23 hy3 Tencent 73.02%
#24 doubao-seed-evolving high ByteDance 72.75%
#25 qwen3.8-27b xhigh Alibaba 68.25%
#26 gemini-3.1-pro-preview Google 64.70%
#27 gemini-3.5-flash Google 63.89%
#28 MiniMax/MiniMax-M3 MiniMax 60.05%
#29 claude-haiku-4-5 high Anthropic 43.65%

View TapTap Maker

Which leaderboard should you use?

  • For overall capability, use Artificial Analysis and LiveBench.
  • For real-user preference, use LMArena Text.
  • For web and front-end development, use LMArena WebDev.
  • For long professional workflows, use APEX-Agents.
  • For financial research, use Vals Finance Agent.
  • For software engineering, use DeepSWE.
  • For game development, use TapTap Maker.

The practical approach is not to chase one overall ranking. Start with your task, consult the relevant leaderboard, then run a small test using your own work. A leaderboard narrows the field; your real workflow makes the final choice.

Link copied