The Inference Report

August 11, 2026

The SWE-rebench leaderboard shows no movement in the top tier, with AnthropicFable 5 holding 64.5% and the next four models (Grok 4.5, Opus 5, GLM-5.2, GPT-5.6 Sol) maintaining their positions between 62.3% and 63.8%. The confidence intervals remain wide enough that none of these differences reach statistical significance, and the unchanged rankings suggest either genuine stability in agent performance on software engineering tasks or insufficient sensitivity in the benchmark to detect real improvements. The Artificial Analysis leaderboard, by contrast, shows substantial churn across the full 400+ model range, with entries like Motif 3 rising from 44.9 to 45.3 (moving from #27 to #26), Nova 2.0 Lite jumping from #157 with 19.2 to #145 with 20.8, and KAT Coder Pro V2 climbing from #81 at 33.9 to #74 at 34.5. This divergence is instructive: the SWE-rebench scores derive from controlled evaluation on real repositories with defined success criteria, while Artificial Analysis aggregates performance across broader language understanding benchmarks where small numerical shifts propagate into ranking changes. The stability at the top of SWE-rebench may reflect that frontier agent performance has plateaued on this particular task distribution, or that the evaluation's variance makes sub-1% score differences unreliable for ranking claims. Without methodological details on how Artificial Analysis scores are computed or weighted, it is difficult to assess whether its volatility represents genuine capability shifts or measurement noise.

Cole Brennan

Daily rankings from SWE-rebench, a benchmark designed to fairly compare LLM capabilities on real-world software engineering tasks. Unlike other evaluations, it uses a standardized scaffolding for all models, continuously updates its dataset to prevent contamination, and runs each model five times to account for stochastic variance.

#ModelScore
1AnthropicFable 5 [high]Model64.5%± 1.41%
2GrokGrok 4.5 [high]Model63.8%± 0.60%
3AnthropicOpus 5 [high]Model63.4%± 1.35%
4Z.aiGLM-5.2 [high]Model62.9%± 1.19%
5OpenAIGPT-5.6 Sol [medium]Model62.3%± 1.83%
6JunieJunieAgent61.8%± 0.54%
7AnthropicClaude CodeAgent60.4%± 1.03%
8OpenAICodexAgent58.0%± 1.29%
9AnthropicSonnet 5 [high]Model56.8%± 0.94%
10CursorCursorAgent51.7%± 0.84%

Artificial Analysis composite index across coding, math, and reasoning benchmarks.

#ModelScoretok/s$/1M
1Claude Opus 563.149$10.00
2Claude Fable 562.163$20.00
3GPT-5.6 Sol60.968$11.25
4Kimi K359.741$6.00
5Qwen3.8 Max58.154$3.00
6Claude Opus 4.857.30$10.00
7Muse Spark 1.256.80$2.00
8GPT-5.6 Terra56.6121$4.50
9GPT-5.556.30$11.25
10Grok 4.555.849$3.00

Output tokens per second — higher is faster. Minimum intelligence score of 40.

#Modeltok/s
1Gemini 3.6 Flash196
2GPT-5.6 Luna163
3Inkling Small145
4Nex-N2-Pro142
5GPT-5.3 Codex133
6Gemini 3.1 Pro Preview126
7GLM-5.2124
8GPT-5.6 Terra121
9DeepSeek V4 Flash 0731112
10MiniMax-M393

Blended cost per 1M tokens (3:1 input/output) — lower is cheaper. Minimum intelligence score of 40.

#Model$/1M
1DeepSeek V4 Flash$0.168
2DeepSeek V4 Flash 0731$0.175
3Hy3$0.241
4GPT-5.6 Luna$0.45
5MiniMax-M3$0.525
6Inkling Small$0.525
7DeepSeek V4 Pro$0.544
8MiMo-V2.5-Pro$0.544
9Nex-N2-Pro$1.00
10Qwen3.6 Plus$1.13