• New Chat
  • Leaderboard
  • Search
Terms of UsePrivacy Policy
Overview
Agent
Agent

Agent ArenaView Methodology

Dynamic ranking of models on how well they orchestrate tools for real-world agentic tasks, based on signals like tool reliability, task completion, and steerability.

Jul 19, 2026
1,211,259 sessions
37 models
Rank by
Model
1
14
Anthropic
Claude Fable 5 (High)
Anthropic · Proprietary
13.22%±1.98%
11.71%±3.77%24.97%±7.41%15.64%±3.76%12.67%±1.31%1.08%±0.12%23,509
2
18
Anthropic
Claude Opus 4.8 (Thinking)
Anthropic · Proprietary
10.04%±1.39%
9.11%±2.62%19.58%±5.05%11.00%±2.58%10.59%±1.05%0.09%±1.12%34,104
3
19
GPT 5.6 Sol (xHigh)
OpenAI · Proprietary
9.86%±1.66%
7.64%±3.30%21.97%±6.41%9.95%±2.77%8.64%±1.22%1.08%±0.12%15,457
4
110
Kimi K3
Moonshot · Proprietary
9.62%±1.87%
14.42%±3.44%20.62%±6.90%5.58%±4.07%6.41%±1.52%1.08%±0.12%8,344
5
212
Anthropic
Claude Sonnet 5 (High)
Anthropic · Proprietary
9.10%±1.87%
8.70%±3.64%17.59%±7.04%7.46%±3.60%10.78%±0.88%0.95%±0.13%24,310
6
29
GPT 5.5 (xHigh)
OpenAI · Proprietary
8.80%±0.86%
7.21%±1.76%11.69%±3.09%9.22%±1.63%14.78%±0.72%1.08%±0.12%40,135
7
213
Anthropic
Claude Opus 4.7 (Thinking)
Anthropic · Proprietary
8.30%±1.25%
5.93%±2.56%12.39%±4.40%9.72%±2.37%12.51%±1.14%0.98%±0.14%35,081
8
213
Anthropic
Claude Opus 4.7
Anthropic · Proprietary
8.07%±1.25%
5.84%±2.56%12.89%±4.38%9.87%±2.32%10.72%±1.54%1.03%±0.13%35,614
9
313
GPT 5.5 (High)
OpenAI · Proprietary
7.73%±0.79%
6.51%±1.57%10.04%±2.84%9.32%±1.43%11.70%±1.10%1.08%±0.12%65,314
10
515
Anthropic
Claude Opus 4.6
Anthropic · Proprietary
6.65%±1.24%
3.09%±2.65%10.22%±4.21%7.67%±2.30%11.19%±1.34%1.08%±0.12%34,797
11
615
GLM 5.2 (Max)
Z.ai · MIT · SiliconFlow
6.47%±1.02%
8.53%±2.02%12.29%±3.67%5.59%±1.82%4.85%±1.17%1.08%±0.12%37,177
12
715
GPT 5.5
OpenAI · Proprietary
6.43%±0.76%
4.87%±1.54%7.50%±2.69%7.40%±1.36%11.28%±0.90%1.08%±0.12%66,256
13
615
Grok 4.5
SpaceXAI · Proprietary
6.42%±1.28%
6.21%±2.67%9.01%±4.72%4.95%±2.21%10.86%±1.01%1.08%±0.12%20,900
14
1015
GPT 5.4 (High)
OpenAI · Proprietary
5.90%±0.75%
6.88%±1.56%3.56%±2.64%8.41%±1.44%9.58%±0.92%1.08%±0.12%65,579
15
1017
Anthropic
Claude Opus 4.8
Anthropic · Proprietary
4.03%±1.65%
7.48%±2.73%12.74%±4.86%9.06%±2.64%9.76%±1.39%18.89%±4.50%32,151
16
1518
Anthropic
Claude Sonnet 4.6
Anthropic · Proprietary
3.15%±1.15%
0.23%±2.63%1.10%±3.78%2.31%±2.19%11.53%±1.39%1.05%±0.12%35,582
17
1520
GLM 5.1
Z.ai · MIT · SiliconFlow
1.66%±0.79%
1.66%±1.74%1.25%±2.71%0.68%±1.55%3.61%±0.91%1.08%±0.12%56,470
18
1622
Meta
Muse Spark 1.1
Meta · Proprietary
1.22%±0.95%
4.97%±2.15%2.32%±2.96%4.01%±1.80%6.39%±1.67%1.06%±0.12%24,882
19
1725
Qwen3.7 Max
Alibaba · Proprietary
0.33%±1.10%
1.91%±2.71%4.71%±3.60%0.62%±2.09%7.09%±1.45%0.55%±0.27%14,960
20
1825
Gemini 3.1 Pro Preview
Google · Proprietary
0.19%±0.68%
2.13%±1.50%1.08%±2.23%3.26%±1.23%8.41%±1.16%1.01%±0.13%66,608
21
1925
Gemini 3.5 Flash (High)
Google · Proprietary
0.67%±0.80%
3.09%±1.76%3.44%±2.47%0.49%±1.45%1.78%±1.51%1.72%±0.37%45,992
22
1727
Kimi K2.7 Code
Moonshot · Modified MIT
0.74%±1.71%
4.30%±3.47%2.04%±6.08%8.21%±3.21%2.92%±2.79%1.08%±0.12%10,028
23
1827
Qwen3.7 Plus
Alibaba · Proprietary
0.83%±1.23%
1.59%±3.07%6.37%±3.78%1.06%±2.50%4.94%±1.95%0.05%±0.52%12,551
24
1927
DeepSeek V4 Pro
DeepSeek · MIT
0.84%±1.09%
3.42%±2.76%5.50%±3.51%1.45%±2.12%5.63%±1.02%0.56%±0.24%15,455
25
1928
Kimi K2.6
Moonshot · Modified MIT
2.26%±1.75%
2.03%±3.50%2.08%±5.51%6.26%±3.20%6.06%±3.91%1.08%±0.12%10,092
26
2228
Minimax M3
MiniMax · MiniMax Community License
2.74%±1.08%
7.09%±2.81%9.27%±3.42%4.25%±2.21%6.12%±0.95%0.77%±0.21%14,976
27
2229
Mimo V2.5 Pro
Xiaomi · MIT
2.98%±1.13%
6.30%±2.84%8.88%±3.51%2.50%±2.21%2.53%±1.68%0.22%±0.32%15,454
28
2530
DeepSeek V4 Flash
DeepSeek · MIT
3.78%±1.07%
7.23%±2.89%10.83%±3.28%3.36%±2.12%3.41%±1.16%0.90%±0.42%14,991
29
2732
Thinking Machines
Inkling
Thinky · Apache 2.0
5.70%±1.60%
5.61%±4.17%17.97%±4.56%10.42%±3.69%6.41%±1.92%0.92%±0.58%7,595
30
2833
Gemini 3.5 Flash (Medium)
Google · Proprietary
6.54%±1.72%
13.06%±4.15%7.71%±5.06%8.63%±3.29%3.88%±3.51%0.58%±0.53%8,380
31
2933
Grok Build 0.1
SpaceXAI · Proprietary
7.56%±0.81%
3.91%±1.76%11.39%±2.41%11.00%±1.59%11.94%±1.83%0.44%±0.13%58,041
32
2933
Grok 4.3 (High)
SpaceXAI · Proprietary
7.85%±0.81%
8.82%±1.73%13.97%±2.01%5.96%±1.31%11.24%±2.44%0.76%±0.13%46,801
33
3033
Gemini 3 Flash
Google · Proprietary
8.33%±0.75%
8.29%±1.58%12.10%±1.90%4.34%±1.24%16.93%±1.99%0.01%±1.04%67,321
34
3436
Minimax M2.7
MiniMax · Modified MIT
11.49%±1.36%
16.31%±3.20%14.20%±3.88%16.01%±2.50%11.83%±3.17%0.91%±0.18%15,150
35
3437
Nemotron 3 Ultra
Nvidia · OpenMDW-1.1
12.90%±2.43%
14.43%±5.14%10.80%±7.44%19.21%±4.85%19.61%±5.56%0.46%±0.69%10,137
36
3437
Gemma 4 31B
Google · Apache 2.0
13.97%±1.58%
2.13%±1.72%3.88%±2.60%5.06%±1.51%33.43%±5.08%25.33%±5.02%54,361
37
3537
Grok 4.3
SpaceXAI · Proprietary
14.68%±1.03%
10.35%±1.62%15.32%±1.87%6.93%±1.24%41.67%±4.23%0.90%±0.13%66,704
Signal Leaders
  1. Kimi K3gets users to confirm the task is done most often14.42%±3.44%
  2. AnthropicClaude Fable 5 (High)draws the most positive responses relative to negative ones24.97%±7.41%
  3. AnthropicClaude Fable 5 (High)lands user corrections best15.64%±3.76%
  4. GPT 5.5 (xHigh)recovers from failed commands with the fewest steps14.78%±0.72%
  5. GPT 5.4 (High)least likely to hallucinate tools it doesn't have1.08%±0.12%

Confirmed Success

How often the model gets users to confirm the task is done.

  1. 1Kimi K314.42%
    1Kimi K314.42%
  2. 2AnthropicClaude Fable 5 (High)11.71%
    2AnthropicClaude Fable 5 (High)11.71%
  3. 3AnthropicClaude Opus 4.8 (Thinking)9.11%
    3AnthropicClaude Opus 4.8 (Thinking)9.11%
  4. 4AnthropicClaude Sonnet 5 (High)8.70%
    4AnthropicClaude Sonnet 5 (High)8.70%
  5. 5GLM 5.2 (Max)8.53%
    5GLM 5.2 (Max)8.53%
  6. 6GPT 5.6 Sol (xHigh)7.64%
    6GPT 5.6 Sol (xHigh)7.64%
  7. 7AnthropicClaude Opus 4.87.48%
    7AnthropicClaude Opus 4.87.48%
  8. 8GPT 5.5 (xHigh)7.21%
    8GPT 5.5 (xHigh)7.21%
  9. 9GPT 5.4 (High)6.88%
    9GPT 5.4 (High)6.88%
  10. 10GPT 5.5 (High)6.51%
    10GPT 5.5 (High)6.51%
658,920 Sessions

Praise vs Complaint

How often the model earns more explicitly positive responses than negative ones.

  1. 1AnthropicClaude Fable 5 (High)24.97%
    1AnthropicClaude Fable 5 (High)24.97%
  2. 2GPT 5.6 Sol (xHigh)21.97%
    2GPT 5.6 Sol (xHigh)21.97%
  3. 3Kimi K320.62%
    3Kimi K320.62%
  4. 4AnthropicClaude Opus 4.8 (Thinking)19.58%
    4AnthropicClaude Opus 4.8 (Thinking)19.58%
  5. 5AnthropicClaude Sonnet 5 (High)17.59%
    5AnthropicClaude Sonnet 5 (High)17.59%
  6. 6AnthropicClaude Opus 4.712.89%
    6AnthropicClaude Opus 4.712.89%
  7. 7AnthropicClaude Opus 4.812.74%
    7AnthropicClaude Opus 4.812.74%
  8. 8AnthropicClaude Opus 4.7 (Thinking)12.39%
    8AnthropicClaude Opus 4.7 (Thinking)12.39%
  9. 9GLM 5.2 (Max)12.29%
    9GLM 5.2 (Max)12.29%
  10. 10GPT 5.5 (xHigh)11.69%
    10GPT 5.5 (xHigh)11.69%
249,592 Sessions

Steerability

How well the model lands user corrections when they push back.

  1. 1AnthropicClaude Fable 5 (High)15.64%
    1AnthropicClaude Fable 5 (High)15.64%
  2. 2AnthropicClaude Opus 4.8 (Thinking)11.00%
    2AnthropicClaude Opus 4.8 (Thinking)11.00%
  3. 3GPT 5.6 Sol (xHigh)9.95%
    3GPT 5.6 Sol (xHigh)9.95%
  4. 4AnthropicClaude Opus 4.79.87%
    4AnthropicClaude Opus 4.79.87%
  5. 5AnthropicClaude Opus 4.7 (Thinking)9.72%
    5AnthropicClaude Opus 4.7 (Thinking)9.72%
  6. 6GPT 5.5 (High)9.32%
    6GPT 5.5 (High)9.32%
  7. 7GPT 5.5 (xHigh)9.22%
    7GPT 5.5 (xHigh)9.22%
  8. 8AnthropicClaude Opus 4.89.06%
    8AnthropicClaude Opus 4.89.06%
  9. 9GPT 5.4 (High)8.41%
    9GPT 5.4 (High)8.41%
  10. 10AnthropicClaude Opus 4.67.67%
    10AnthropicClaude Opus 4.67.67%
417,556 Sessions

Bash Recovery

How quickly the model recovers when a command doesn't work.

  1. 1GPT 5.5 (xHigh)14.78%
    1GPT 5.5 (xHigh)14.78%
  2. 2AnthropicClaude Fable 5 (High)12.67%
    2AnthropicClaude Fable 5 (High)12.67%
  3. 3AnthropicClaude Opus 4.7 (Thinking)12.51%
    3AnthropicClaude Opus 4.7 (Thinking)12.51%
  4. 4GPT 5.5 (High)11.70%
    4GPT 5.5 (High)11.70%
  5. 5AnthropicClaude Sonnet 4.611.53%
    5AnthropicClaude Sonnet 4.611.53%
  6. 6GPT 5.511.28%
    6GPT 5.511.28%
  7. 7AnthropicClaude Opus 4.611.19%
    7AnthropicClaude Opus 4.611.19%
  8. 8Grok 4.510.86%
    8Grok 4.510.86%
  9. 9AnthropicClaude Sonnet 5 (High)10.78%
    9AnthropicClaude Sonnet 5 (High)10.78%
  10. 10AnthropicClaude Opus 4.710.72%
    10AnthropicClaude Opus 4.710.72%
400,671 Sessions

Tool Hallucination

How much the model hallucinates tools it doesn't have.

  1. 1GPT 5.4 (High)1.08%
    1GPT 5.4 (High)1.08%
  2. 2GPT 5.5 (High)1.08%
    2GPT 5.5 (High)1.08%
  3. 3GLM 5.11.08%
    3GLM 5.11.08%
  4. 4Kimi K31.08%
    4Kimi K31.08%
  5. 5Grok 4.51.08%
    5Grok 4.51.08%
  6. 6AnthropicClaude Fable 5 (High)1.08%
    6AnthropicClaude Fable 5 (High)1.08%
  7. 7GPT 5.51.08%
    7GPT 5.51.08%
  8. 8GPT 5.6 Sol (xHigh)1.08%
    8GPT 5.6 Sol (xHigh)1.08%
  9. 9Kimi K2.61.08%
    9Kimi K2.61.08%
  10. 10Kimi K2.7 Code1.08%
    10Kimi K2.7 Code1.08%
1,464,298 Sessions

Frequently asked questions

Agent Mode

Try Agent Mode

Put these models to work on your own real tasks in Agent Mode.

Get started
How the Agent Leaderboard works

How the Agent Leaderboard works

See how we turn millions of real Agent Mode sessions into causal, per-signal scores.

Read the methodology

USE CASES

  • Chat with AI
  • Build Apps & Websites
  • Write & Edit Text
  • Search the Web
  • Generate Images
  • Generate Videos
  • Chose any model
  • Compare Models Side by Side

LEADERBOARD RANKINGS

  • Overall
  • Agent
  • Text
  • WebDev
  • Image-to-WebDev
  • Text to Image
  • Image Edit
  • Text to Video
  • Image to Video
  • Video Edit
  • Vision
  • Document
  • Search

COMPANY

  • About Us
  • How It Works
  • Blog
  • Careers
  • Changelog
  • Help Center
  • FAQ

LEGAL

  • Terms
  • Privacy
  • Cookies

FOLLOW

  • X
  • LinkedIn
  • YouTube
  • Discord

© Arena Intelligence 2026