The Inference Report

August 5, 2026

The SWE-rebench rankings show no movement from the previous cycle, with AnthropicFable 5 holding the top position at 64.5% (±1.41%), followed by GrokGrok 4.5 at 63.8% (±0.60%) and AnthropicOpus 5 at 63.4% (±1.35%). The confidence intervals are wide enough that several adjacent pairs could plausibly swap on different test runs, yet the top tier remains stable across both measurement periods. By contrast, the Artificial Analysis benchmark presents a different ordering entirely, with Claude Opus 5 ranking first at 60.7 rather than second, and Claude Fable 5 at 59.9 rather than first, suggesting the two benchmarks are measuring distinct performance dimensions or using fundamentally different evaluation protocols. The SWE-rebench methodology appears to isolate specific coding task behaviors that differ from the broader capability assessment in Artificial Analysis, but without documentation of task construction, test set composition, or inter-rater reliability for either benchmark, it remains unclear whether the divergence reflects genuine performance variation across problem types or systematic differences in how each benchmark weights model capabilities. The lack of movement in SWE-rebench rankings over this period could indicate either stable model performance on this specific task distribution or insufficient statistical power to detect real changes given the uncertainty bands, particularly for models in the 40-60% range where error margins approach or exceed the gaps between adjacent entries.

Cole Brennan

Daily rankings from SWE-rebench, a benchmark designed to fairly compare LLM capabilities on real-world software engineering tasks. Unlike other evaluations, it uses a standardized scaffolding for all models, continuously updates its dataset to prevent contamination, and runs each model five times to account for stochastic variance.

#ModelScore
1AnthropicFable 5 [high]Model64.5%± 1.41%
2GrokGrok 4.5 [high]Model63.8%± 0.60%
3AnthropicOpus 5 [high]Model63.4%± 1.35%
4Z.aiGLM-5.2 [high]Model62.9%± 1.19%
5OpenAIGPT-5.6 Sol [medium]Model62.3%± 1.83%
6JunieJunieAgent61.8%± 0.54%
7AnthropicClaude CodeAgent60.4%± 1.03%
8OpenAICodexAgent58.0%± 1.29%
9AnthropicSonnet 5 [high]Model56.8%± 0.94%
10CursorCursorAgent51.7%± 0.84%

Artificial Analysis composite index across coding, math, and reasoning benchmarks.

#ModelScoretok/s$/1M
1Claude Opus 560.762$10.00
2Claude Fable 559.971$20.00
3GPT-5.6 Sol58.976$11.25
4Kimi K357.136$6.00
5Claude Opus 4.855.70$10.00
6GPT-5.6 Terra55149$4.50
7GPT-5.554.80$11.25
8Grok 4.553.867$3.00
9Claude Opus 4.753.50$10.00
10Claude Sonnet 553.497$4.00

Output tokens per second — higher is faster. Minimum intelligence score of 40.

#Modeltok/s
1Gemini 3.5 Flash288
2Gemini 3.6 Flash236
3Muse Spark 1.1217
4Qwen3.7 Max213
5GLM-5.2182
6GPT-5.6 Luna178
7GPT-5.6 Terra149
8Gemini 3.1 Pro Preview141
9Nex-N2-Pro138
10GPT-5.3 Codex127

Blended cost per 1M tokens (3:1 input/output) — lower is cheaper. Minimum intelligence score of 40.

#Model$/1M
1DeepSeek V4 Flash$0.171
2DeepSeek V4 Flash 0731$0.175
3Hy3$0.241
4GPT-5.6 Luna$0.45
5MiniMax-M3$0.525
6Inkling Small$0.525
7DeepSeek V4 Pro$0.544
8MiMo-V2.5-Pro$0.544
9Nex-N2-Pro$1.00
10GPT-5.4 mini$1.69