The Inference Report

August 1, 2026

The SWE-rebench leaderboard shows no movement in the top tier, with AnthropicFable 5 maintaining 64.5% ± 1.41%, GrokGrok 4.5 holding 63.8% ± 0.60%, and AnthropicOpus 5 steady at 63.4% ± 1.35%. The Artificial Analysis benchmark, however, reveals more churn across its 417-entry roster, though the top positions remain occupied by Claude Opus 5 (60.7) and Claude Fable 5 (59.9). One new entry appears at rank 227 in Artificial Analysis: Celeris-1 at 11.8, which triggered a cascade of single-position shifts downward through the lower-middle tiers. The confidence intervals on SWE-rebench range from 0.54% to 1.83%, suggesting adequate precision for detecting real differences between adjacent models, yet the frozen top rankings across both benchmarks indicate either genuine convergence at the frontier or evaluation ceiling effects. The Artificial Analysis list, substantially longer and more volatile, shows that differentiation persists below the top 30, where scores drop from the 50s into the 40s and below. Without prior Artificial Analysis snapshots in the historical data, it is unclear whether the Celeris-1 insertion represents a new capability or a recalibration of the ranking system itself. SWE-rebench's stability at the top is consistent with the narrow margin between first and third place (1.1 percentage points), where confidence intervals overlap, making further separation difficult to resolve empirically.

Cole Brennan

Daily rankings from SWE-rebench, a benchmark designed to fairly compare LLM capabilities on real-world software engineering tasks. Unlike other evaluations, it uses a standardized scaffolding for all models, continuously updates its dataset to prevent contamination, and runs each model five times to account for stochastic variance.

#ModelScore
1AnthropicFable 5 [high]Model64.5%± 1.41%
2GrokGrok 4.5 [high]Model63.8%± 0.60%
3AnthropicOpus 5 [high]Model63.4%± 1.35%
4Z.aiGLM-5.2 [high]Model62.9%± 1.19%
5OpenAIGPT-5.6 Sol [medium]Model62.3%± 1.83%
6JunieJunieAgent61.8%± 0.54%
7AnthropicClaude CodeAgent60.4%± 1.03%
8OpenAICodexAgent58.0%± 1.29%
9AnthropicSonnet 5 [high]Model56.8%± 0.94%
10CursorCursorAgent51.7%± 0.84%

Artificial Analysis composite index across coding, math, and reasoning benchmarks.

#ModelScoretok/s$/1M
1Claude Opus 560.756$10.00
2Claude Fable 559.960$20.00
3GPT-5.6 Sol58.967$11.25
4Kimi K357.134$6.00
5Claude Opus 4.855.70$10.00
6GPT-5.6 Terra55128$4.50
7GPT-5.554.80$11.25
8Grok 4.553.852$3.00
9Claude Opus 4.753.50$10.00
10Claude Sonnet 553.479$4.00

Output tokens per second — higher is faster. Minimum intelligence score of 40.

#Modeltok/s
1Gemini 3.5 Flash220
2Gemini 3.6 Flash207
3Qwen3.7 Max200
4GPT-5.6 Luna172
5Nex-N2-Pro132
6Muse Spark 1.1131
7Gemini 3.1 Pro Preview129
8GPT-5.6 Terra128
9GPT-5.3 Codex109
10DeepSeek V4 Flash102

Blended cost per 1M tokens (3:1 input/output) — lower is cheaper. Minimum intelligence score of 40.

#Model$/1M
1DeepSeek V4 Flash 0731$0.175
2DeepSeek V4 Flash$0.175
3Hy3$0.241
4GPT-5.6 Luna$0.45
5MiniMax-M3$0.525
6Inkling Small$0.525
7DeepSeek V4 Pro$0.544
8MiMo-V2.5-Pro$0.544
9Nex-N2-Pro$1.00
10GPT-5.4 mini$1.69