The Inference Report

August 8, 2026

The SWE-rebench leaderboard holds stable at the top, with AnthropicFable 5 maintaining 64.5% and the top five models clustered within 2.2 percentage points, confidence intervals overlapping substantially enough that ranking shifts between them would not represent genuine performance differences. The Artificial Analysis benchmark, by contrast, shows considerable churn across its 422-entry list, with models shuffling positions frequently even when scores differ by 0.1 to 0.3 points, a pattern that raises questions about whether such fine-grained distinctions reflect real capability gaps or measurement noise. Notable movements include Ling 3.0 Flash dropping from rank 59 (37.8) to unranked status, NVIDIA Nemotron 3 Super 120B falling from 119 (25.7) to unranked, Ling 3.0 Tiny climbing from 125 (23.9) to 121 (24.5), and Phi-4 Mini Instruct rising from 326 (5.7) to 316 (6.2). The SWE-rebench methodology, which tests agents on actual software engineering tasks with controlled evaluation conditions, produces the more interpretable signal: the top tier of models genuinely solves 56 to 65 percent of problems, with uncertainty bands that reflect real variability. The Artificial Analysis scores, spanning from 64.5 down to 1.0 across hundreds of entries, compress evaluation into a single dimension that may conflate different failure modes or reflect task-specific brittleness rather than general capability. Without visibility into Artificial Analysis's evaluation protocol, whether it uses held-out test sets, how it handles edge cases, or whether scoring is normalized, the frequent micro-movements and aggressive differentiation at the tail end suggest the ranking may be sensitive to evaluation artifacts rather than tracking reproducible differences in model behavior.

Cole Brennan

Daily rankings from SWE-rebench, a benchmark designed to fairly compare LLM capabilities on real-world software engineering tasks. Unlike other evaluations, it uses a standardized scaffolding for all models, continuously updates its dataset to prevent contamination, and runs each model five times to account for stochastic variance.

#ModelScore
1AnthropicFable 5 [high]Model64.5%± 1.41%
2GrokGrok 4.5 [high]Model63.8%± 0.60%
3AnthropicOpus 5 [high]Model63.4%± 1.35%
4Z.aiGLM-5.2 [high]Model62.9%± 1.19%
5OpenAIGPT-5.6 Sol [medium]Model62.3%± 1.83%
6JunieJunieAgent61.8%± 0.54%
7AnthropicClaude CodeAgent60.4%± 1.03%
8OpenAICodexAgent58.0%± 1.29%
9AnthropicSonnet 5 [high]Model56.8%± 0.94%
10CursorCursorAgent51.7%± 0.84%

Artificial Analysis composite index across coding, math, and reasoning benchmarks.

#ModelScoretok/s$/1M
1Claude Opus 563.148$10.00
2Claude Fable 562.158$20.00
3GPT-5.6 Sol60.963$11.25
4Kimi K359.737$6.00
5Qwen3.8 Max58.169$3.00
6Claude Opus 4.857.30$10.00
7Muse Spark 1.256.80$2.00
8GPT-5.6 Terra56.6115$4.50
9GPT-5.556.30$11.25
10Grok 4.555.851$3.00

Output tokens per second — higher is faster. Minimum intelligence score of 40.

#Modeltok/s
1Gemini 3.5 Flash219
2Qwen3.7 Max195
3Muse Spark 1.1191
4Gemini 3.6 Flash189
5GPT-5.6 Luna176
6Inkling Small130
7Gemini 3.1 Pro Preview127
8GLM-5.2119
9Nex-N2-Pro118
10GPT-5.6 Terra115

Blended cost per 1M tokens (3:1 input/output) — lower is cheaper. Minimum intelligence score of 40.

#Model$/1M
1DeepSeek V4 Flash$0.168
2DeepSeek V4 Flash 0731$0.175
3Hy3$0.241
4GPT-5.6 Luna$0.45
5MiniMax-M3$0.525
6Inkling Small$0.525
7DeepSeek V4 Pro$0.544
8MiMo-V2.5-Pro$0.544
9Nex-N2-Pro$1.00
10Qwen3.6 Plus$1.13