The Inference Report

August 15, 2026

The SWE-rebench leaderboard shows stability at the top tier, with no movement in the top ten positions. AnthropicFable 5 maintains 64.5% ± 1.41%, followed by GrokGrok 4.5 at 63.8% ± 0.60% and AnthropicOpus 5 at 63.4% ± 1.35%, margins well within their confidence intervals. The Artificial Analysis benchmark, by contrast, reveals substantial reshuffling across the middle and lower ranks, with Claude Opus 5 rising to #1 at 63.1 (from #2 at 63.1), Claude Fable 5 at #2 with 62.1, and notable drops for models like DeepSeek V4 Pro, which fell from #16 at 53.2 to #30 at 45.3. The discrepancy between these two benchmarks warrants scrutiny: SWE-rebench uses controlled problem-solving conditions with uncertainty quantification, while Artificial Analysis provides point scores without error bounds, making direct comparison difficult. Within SWE-rebench's narrower scope, the consistency suggests the benchmark has sufficient resolution to distinguish performance in the 40-65% range but may be reaching saturation at the frontier, where confidence intervals begin to overlap meaningfully. The Artificial Analysis shifts hint either at different evaluation methodologies, data drift, or sensitivity to model versioning that SWE-rebench does not capture. Until both benchmarks publish their evaluation protocols in detail, the divergence remains interpretable but not fully explainable.

Cole Brennan

Daily rankings from SWE-rebench, a benchmark designed to fairly compare LLM capabilities on real-world software engineering tasks. Unlike other evaluations, it uses a standardized scaffolding for all models, continuously updates its dataset to prevent contamination, and runs each model five times to account for stochastic variance.

#ModelScore
1AnthropicFable 5 [high]Model64.5%± 1.41%
2GrokGrok 4.5 [high]Model63.8%± 0.60%
3AnthropicOpus 5 [high]Model63.4%± 1.35%
4Z.aiGLM-5.2 [high]Model62.9%± 1.19%
5OpenAIGPT-5.6 Sol [medium]Model62.3%± 1.83%
6JunieJunieAgent61.8%± 0.54%
7AnthropicClaude CodeAgent60.4%± 1.03%
8OpenAICodexAgent58.0%± 1.29%
9AnthropicSonnet 5 [high]Model56.8%± 0.94%
10CursorCursorAgent51.7%± 0.84%

Artificial Analysis composite index across coding, math, and reasoning benchmarks.

#ModelScoretok/s$/1M
1Claude Opus 563.146$10.00
2Claude Fable 562.158$20.00
3GPT-5.6 Sol60.960$11.25
4Grok 4.660.955$3.00
5Kimi K359.736$6.00
6Qwen3.8 Max58.143$3.00
7Qwen3.8 2.4T A95B57.748$3.00
8Claude Opus 4.857.30$10.00
9Muse Spark 1.256.80$2.00
10GPT-5.6 Terra56.6102$4.50

Output tokens per second — higher is faster. Minimum intelligence score of 40.

#Modeltok/s
1Gemini 3.7 Flash515
2Gemini 3.6 Flash205
3GPT-5.6 Luna148
4Nex-N2-Pro142
5Gemini 3.1 Pro Preview128
6Inkling Small113
7GPT-5.3 Codex112
8GLM-5.2104
9DeepSeek V4 Flash 0731103
10GPT-5.6 Terra102

Blended cost per 1M tokens (3:1 input/output) — lower is cheaper. Minimum intelligence score of 40.

#Model$/1M
1DeepSeek V4 Flash$0.168
2Hy3$0.241
3GPT-5.6 Luna$0.45
4MiniMax-M3$0.525
5Solar Pro 4$0.525
6Inkling Small$0.525
7DeepSeek V4 Pro$0.544
8MiMo-V2.5-Pro$0.544
9DeepSeek V4 Flash 0731$0.66
10Nex-N2-Pro$1.00