The Inference Report

July 30, 2026

The SWE-rebench standings remain frozen at their previous positions, with AnthropicFable 5 holding 64.5% ± 1.41%, GrokGrok 4.5 at 63.8% ± 0.60%, and AnthropicOpus 5 at 63.4% ± 1.35% across the top three slots, suggesting either a pause in model iteration or an evaluation cycle that hasn't refreshed since the last report. The Artificial Analysis benchmark, by contrast, shows substantial churn throughout its 415-entry roster, with Claude Opus 5 now leading at 60.7 (up from prior position), Claude Fable 5 at 59.9, and GPT-5.6 Sol at 58.9, though the methodology underlying these scores remains opaque and the entries lack confidence intervals that would allow assessment of whether observed ranking shifts exceed noise. A notable addition appears at position 36: Inkling Small enters the Artificial Analysis rankings at 40.2, bumping prior entries down, while Mistral Medium 3.1 enters at 188 with a score of 14.7, suggesting either new model releases or expanded evaluation coverage. The two benchmarks diverge sharply in their top performers, SWE-rebench privileges Anthropic and Grok models at the frontier, while Artificial Analysis splits leadership across Anthropic and OpenAI, a divergence that likely reflects different problem distributions and evaluation rigor rather than a meaningful signal about absolute capability. Without documentation of Artificial Analysis's evaluation protocol, sample sizes, or inter-rater agreement, the ranking movements cannot be distinguished from reordering noise, making the SWE-rebench's stable top tier the more reliable reference point despite its slower refresh cycle.

Cole Brennan

Daily rankings from SWE-rebench, a benchmark designed to fairly compare LLM capabilities on real-world software engineering tasks. Unlike other evaluations, it uses a standardized scaffolding for all models, continuously updates its dataset to prevent contamination, and runs each model five times to account for stochastic variance.

#ModelScore
1AnthropicFable 5 [high]Model64.5%± 1.41%
2GrokGrok 4.5 [high]Model63.8%± 0.60%
3AnthropicOpus 5 [high]Model63.4%± 1.35%
4Z.aiGLM-5.2 [high]Model62.9%± 1.19%
5OpenAIGPT-5.6 Sol [medium]Model62.3%± 1.83%
6JunieJunieAgent61.8%± 0.54%
7AnthropicClaude CodeAgent60.4%± 1.03%
8OpenAICodexAgent58.0%± 1.29%
9AnthropicSonnet 5 [high]Model56.8%± 0.94%
10CursorCursorAgent51.7%± 0.84%

Artificial Analysis composite index across coding, math, and reasoning benchmarks.

#ModelScoretok/s$/1M
1Claude Opus 560.755$10.00
2Claude Fable 559.958$20.00
3GPT-5.6 Sol58.967$11.25
4Kimi K357.132$6.00
5Claude Opus 4.855.760$10.00
6GPT-5.6 Terra55135$4.50
7GPT-5.554.80$11.25
8Grok 4.553.855$3.00
9Claude Opus 4.753.50$10.00
10Claude Sonnet 553.485$4.00

Output tokens per second — higher is faster. Minimum intelligence score of 40.

#Modeltok/s
1Gemini 3.5 Flash220
2Gemini 3.6 Flash214
3Qwen3.7 Max202
4GPT-5.6 Luna187
5Gemini 3.1 Pro Preview136
6GPT-5.6 Terra135
7Muse Spark 1.1135
8Nex-N2-Pro133
9GLM-5.2113
10DeepSeek V4 Flash112

Blended cost per 1M tokens (3:1 input/output) — lower is cheaper. Minimum intelligence score of 40.

#Model$/1M
1DeepSeek V4 Flash$0.175
2Hy3$0.241
3GPT-5.6 Luna$0.45
4MiniMax-M3$0.525
5Inkling Small$0.525
6DeepSeek V4 Pro$0.544
7MiMo-V2.5-Pro$0.544
8Nex-N2-Pro$1.00
9GPT-5.4 mini$1.69
10Kimi K2.6$1.71