The Inference Report

July 31, 2026

The SWE-rebench rankings remain unchanged from the previous cycle, with Anthropic's Fable 5 holding the top position at 64.5 percent, followed by Grok 4.5 at 63.8 percent and Opus 5 at 63.4 percent; the stability across the top tier reflects the difficulty of moving beyond these performance plateaus, though the confidence intervals suggest room for meaningful separation if the variance decreases. The Artificial Analysis leaderboard, by contrast, shows considerable churn in the middle ranks, with DeepSeek V4 Flash 0731 entering at position 17 with a score of 49.9, displacing the previous entry and pushing Claude Sonnet 4.6 from 17th to 18th, yet this movement involves models scoring within a compressed band between 47 and 50 where differentiation becomes fragile. At the top of Artificial Analysis, Claude Opus 5 leads at 60.7, a position it does not hold on SWE-rebench where it ranks third; this discrepancy points to benchmark sensitivity in methodology or task composition, suggesting that neither leaderboard alone captures the full picture of coding capability. The SWE-rebench data carries reported confidence intervals, which provide some guard against false precision, whereas Artificial Analysis reports point estimates without uncertainty quantification, making it harder to assess whether observed shifts represent genuine performance changes or measurement noise. What stands out is not volatility at the extremes but rather the emergence of a dense cluster of models in the 40 to 50 range on both benchmarks, indicating that the frontier is broadening even as the absolute ceiling remains contested.

Cole Brennan

Daily rankings from SWE-rebench, a benchmark designed to fairly compare LLM capabilities on real-world software engineering tasks. Unlike other evaluations, it uses a standardized scaffolding for all models, continuously updates its dataset to prevent contamination, and runs each model five times to account for stochastic variance.

#ModelScore
1AnthropicFable 5 [high]Model64.5%± 1.41%
2GrokGrok 4.5 [high]Model63.8%± 0.60%
3AnthropicOpus 5 [high]Model63.4%± 1.35%
4Z.aiGLM-5.2 [high]Model62.9%± 1.19%
5OpenAIGPT-5.6 Sol [medium]Model62.3%± 1.83%
6JunieJunieAgent61.8%± 0.54%
7AnthropicClaude CodeAgent60.4%± 1.03%
8OpenAICodexAgent58.0%± 1.29%
9AnthropicSonnet 5 [high]Model56.8%± 0.94%
10CursorCursorAgent51.7%± 0.84%

Artificial Analysis composite index across coding, math, and reasoning benchmarks.

#ModelScoretok/s$/1M
1Claude Opus 560.755$10.00
2Claude Fable 559.959$20.00
3GPT-5.6 Sol58.967$11.25
4Kimi K357.133$6.00
5Claude Opus 4.855.752$10.00
6GPT-5.6 Terra55128$4.50
7GPT-5.554.80$11.25
8Grok 4.553.855$3.00
9Claude Opus 4.753.50$10.00
10Claude Sonnet 553.485$4.00

Output tokens per second — higher is faster. Minimum intelligence score of 40.

#Modeltok/s
1Gemini 3.5 Flash220
2Gemini 3.6 Flash207
3Qwen3.7 Max200
4GPT-5.6 Luna187
5Muse Spark 1.1140
6Gemini 3.1 Pro Preview136
7Nex-N2-Pro133
8GPT-5.6 Terra128
9GPT-5.3 Codex108
10GLM-5.2107

Blended cost per 1M tokens (3:1 input/output) — lower is cheaper. Minimum intelligence score of 40.

#Model$/1M
1DeepSeek V4 Flash 0731$0.175
2DeepSeek V4 Flash$0.175
3Hy3$0.241
4GPT-5.6 Luna$0.45
5MiniMax-M3$0.525
6Inkling Small$0.525
7DeepSeek V4 Pro$0.544
8MiMo-V2.5-Pro$0.544
9Nex-N2-Pro$1.00
10GPT-5.4 mini$1.69