The Inference Report

August 24, 2026

The SWE-rebench leaderboard shows no movement in the top positions; AnthropicFable 5 remains at 64.5% plus or minus 1.41%, with GrokGrok 4.5 and AnthropicOpus 5 holding their second and third positions at 63.8% and 63.4% respectively. The confidence intervals are wide enough that these three models occupy a statistical plateau where ranking shifts would require changes well outside current uncertainty bounds. Below the top tier, the spread widens noticeably: Z.aiGLM-5.2 sits at 62.9% plus or minus 1.19%, OpenAIGPT-5.6 Sol at 62.3% plus or minus 1.83%, and JunieJunieAgent at 61.8% plus or minus 0.54%, indicating that agents with tighter confidence intervals (like Junie) may have undergone more stable evaluation conditions than models with larger error margins. The Artificial Analysis benchmark presents a different picture, with Claude Opus 5 leading at 63.1 rather than Fable 5, suggesting the two benchmarks measure or weight code-solving performance differently or that the models respond differently to their evaluation protocols. The SWE-rebench methodology appears to favor Anthropic's Fable variant and Grok 4.5 specifically, while Artificial Analysis ranks Opus higher, a divergence worth noting when interpreting which benchmark better predicts real coding task performance. Neither benchmark shows dramatic velocity in the upper ranks, which either reflects genuine saturation on the problem set or indicates that incremental improvements at this performance ceiling require substantially larger modeling or data investments than the recent release cycle has provided.

Cole Brennan

Daily rankings from SWE-rebench, a benchmark designed to fairly compare LLM capabilities on real-world software engineering tasks. Unlike other evaluations, it uses a standardized scaffolding for all models, continuously updates its dataset to prevent contamination, and runs each model five times to account for stochastic variance.

#ModelScore
1AnthropicFable 5 [high]Model64.5%± 1.41%
2GrokGrok 4.5 [high]Model63.8%± 0.60%
3AnthropicOpus 5 [high]Model63.4%± 1.35%
4Z.aiGLM-5.2 [high]Model62.9%± 1.19%
5OpenAIGPT-5.6 Sol [medium]Model62.3%± 1.83%
6JunieJunieAgent61.8%± 0.54%
7AnthropicClaude CodeAgent60.4%± 1.03%
8OpenAICodexAgent58.0%± 1.29%
9AnthropicSonnet 5 [high]Model56.8%± 0.94%
10CursorCursorAgent51.7%± 0.84%

Artificial Analysis composite index across coding, math, and reasoning benchmarks.

#ModelScoretok/s$/1M
1Claude Opus 563.152$10.00
2Claude Fable 562.163$20.00
3GPT-5.6 Sol60.964$11.25
4Grok 4.660.956$3.00
5Kimi K359.734$6.00
6GLM-5.359.50$2.15
7Qwen3.8 Max58.143$3.00
8Qwen3.8 2.4T A95B57.745$3.00
9Claude Opus 4.857.30$10.00
10Muse Spark 1.256.80$2.00

Output tokens per second — higher is faster. Minimum intelligence score of 40.

#Modeltok/s
1Gemini 3.7 Flash327
2Gemini 3.6 Flash196
3GPT-5.6 Luna139
4Nex-N2-Pro136
5Inkling Small135
6MiniMax-M3133
7GPT-5.3 Codex128
8Gemini 3.1 Pro Preview123
9GPT-5.6 Terra111
10DeepSeek V4 Flash 0731109

Blended cost per 1M tokens (3:1 input/output) — lower is cheaper. Minimum intelligence score of 40.

#Model$/1M
1DeepSeek V4 Flash$0.168
2Hy3$0.241
3GPT-5.6 Luna$0.45
4MiniMax-M3$0.525
5Solar Pro 4$0.525
6Inkling Small$0.525
7DeepSeek V4 Pro$0.544
8MiMo-V2.5-Pro$0.544
9DeepSeek V4 Flash 0731$0.66
10Nex-N2-Pro$1.00