The Inference Report

August 6, 2026

The SWE-rebench rankings show no movement from the previous cycle, with AnthropicFable 5 holding 64.5% ± 1.41% at the top, followed by GrokGrok 4.5 at 63.8% ± 0.60% and AnthropicOpus 5 at 63.4% ± 1.35%, but the Artificial Analysis benchmark reveals substantial churn across its 421-model roster, with two new entries disrupting the upper tier: Qwen3.8 Max enters at number 5 with 56.2, pushing Claude Opus 4.8 down one slot to 6, and Muse Spark 1.2 debuts at number 9 with 54.1, while Ling-3.0-flash appears at number 57 with 37.4. The stability in SWE-rebench contrasts sharply with the Artificial Analysis churn and raises a methodological question: SWE-rebench uses confidence intervals (ranging from ± 0.54% to ± 1.83%) suggesting repeated trials or statistical sampling, whereas Artificial Analysis reports single point scores without uncertainty estimates, making it unclear whether the apparent volatility reflects genuine model improvements, evaluation methodology shifts, or simply different statistical rigor between benchmarks. The top-tier SWE-rebench models cluster tightly between 62.3% and 64.5%, overlapping within their confidence bands, which means the ranking order itself carries limited discriminative power at that level. Without historical Artificial Analysis rankings to confirm whether these entries are genuinely new models or data collection artifacts, the apparent movement cannot be cleanly separated from benchmark drift.

Cole Brennan

Daily rankings from SWE-rebench, a benchmark designed to fairly compare LLM capabilities on real-world software engineering tasks. Unlike other evaluations, it uses a standardized scaffolding for all models, continuously updates its dataset to prevent contamination, and runs each model five times to account for stochastic variance.

#ModelScore
1AnthropicFable 5 [high]Model64.5%± 1.41%
2GrokGrok 4.5 [high]Model63.8%± 0.60%
3AnthropicOpus 5 [high]Model63.4%± 1.35%
4Z.aiGLM-5.2 [high]Model62.9%± 1.19%
5OpenAIGPT-5.6 Sol [medium]Model62.3%± 1.83%
6JunieJunieAgent61.8%± 0.54%
7AnthropicClaude CodeAgent60.4%± 1.03%
8OpenAICodexAgent58.0%± 1.29%
9AnthropicSonnet 5 [high]Model56.8%± 0.94%
10CursorCursorAgent51.7%± 0.84%

Artificial Analysis composite index across coding, math, and reasoning benchmarks.

#ModelScoretok/s$/1M
1Claude Opus 560.758$10.00
2Claude Fable 559.971$20.00
3GPT-5.6 Sol58.970$11.25
4Kimi K357.137$6.00
5Qwen3.8 Max56.253$3.00
6Claude Opus 4.855.70$10.00
7GPT-5.6 Terra55127$4.50
8GPT-5.554.80$11.25
9Muse Spark 1.254.10$2.00
10Grok 4.553.859$3.00

Output tokens per second — higher is faster. Minimum intelligence score of 40.

#Modeltok/s
1Gemini 3.5 Flash255
2Gemini 3.6 Flash236
3Qwen3.7 Max201
4Muse Spark 1.1192
5GPT-5.6 Luna166
6GLM-5.2160
7Nex-N2-Pro138
8Gemini 3.1 Pro Preview134
9GPT-5.6 Terra127
10Inkling Small118

Blended cost per 1M tokens (3:1 input/output) — lower is cheaper. Minimum intelligence score of 40.

#Model$/1M
1DeepSeek V4 Flash$0.171
2DeepSeek V4 Flash 0731$0.175
3Hy3$0.241
4GPT-5.6 Luna$0.45
5MiniMax-M3$0.525
6Inkling Small$0.525
7DeepSeek V4 Pro$0.544
8MiMo-V2.5-Pro$0.544
9Nex-N2-Pro$1.00
10GPT-5.4 mini$1.69