The Inference Report

August 12, 2026

The SWE-rebench rankings show no movement in the top tier, with Anthropic's Fable 5 holding at 64.5%, Grok 4.5 at 63.8%, and Opus 5 at 63.4%, all within their confidence intervals and unchanged from the previous cycle. The Artificial Analysis benchmark, by contrast, shows substantial reshuffling: Nemotron 3 Ultra 550B A55B enters at position 128, pushing Gemini 2.5 Pro Preview down one slot, a cascade that ripples through 297 positions below. Claude Opus 5 leads Artificial Analysis at 63.1, a 0.8-point gap from Fable 5 at 62.1, reversing the SWE-rebench ordering and suggesting the benchmarks measure different problem spaces or that the evaluation protocols diverge in their sensitivity to model capability. The SWE-rebench data carries tighter confidence bounds (Grok 4.5 at 0.60%, Junie Agent at 0.54%) compared to the broader spread in Artificial Analysis, indicating more controlled experimental conditions for the coding task, though the lack of movement across 17 entries raises questions about whether the test set has reached saturation or whether the evaluation is insensitive to recent model improvements. Artificial Analysis's dense 424-entry ranking with models clustered in the 1.0 to 3.7 range at the bottom suggests either a different difficulty calibration or inclusion of older or smaller-parameter models that SWE-rebench does not test, making direct comparison between the two benchmarks unreliable for assessing true performance shifts.

Cole Brennan

Daily rankings from SWE-rebench, a benchmark designed to fairly compare LLM capabilities on real-world software engineering tasks. Unlike other evaluations, it uses a standardized scaffolding for all models, continuously updates its dataset to prevent contamination, and runs each model five times to account for stochastic variance.

#ModelScore
1AnthropicFable 5 [high]Model64.5%± 1.41%
2GrokGrok 4.5 [high]Model63.8%± 0.60%
3AnthropicOpus 5 [high]Model63.4%± 1.35%
4Z.aiGLM-5.2 [high]Model62.9%± 1.19%
5OpenAIGPT-5.6 Sol [medium]Model62.3%± 1.83%
6JunieJunieAgent61.8%± 0.54%
7AnthropicClaude CodeAgent60.4%± 1.03%
8OpenAICodexAgent58.0%± 1.29%
9AnthropicSonnet 5 [high]Model56.8%± 0.94%
10CursorCursorAgent51.7%± 0.84%

Artificial Analysis composite index across coding, math, and reasoning benchmarks.

#ModelScoretok/s$/1M
1Claude Opus 563.153$10.00
2Claude Fable 562.166$20.00
3GPT-5.6 Sol60.965$11.25
4Kimi K359.740$6.00
5Qwen3.8 Max58.151$3.00
6Claude Opus 4.857.30$10.00
7Muse Spark 1.256.80$2.00
8GPT-5.6 Terra56.6121$4.50
9GPT-5.556.30$11.25
10Grok 4.555.850$3.00

Output tokens per second — higher is faster. Minimum intelligence score of 40.

#Modeltok/s
1Gemini 3.6 Flash206
2GPT-5.6 Luna164
3Nex-N2-Pro144
4Gemini 3.1 Pro Preview126
5GLM-5.2125
6Inkling Small125
7GPT-5.3 Codex124
8GPT-5.6 Terra121
9DeepSeek V4 Flash 0731118
10MiniMax-M396

Blended cost per 1M tokens (3:1 input/output) — lower is cheaper. Minimum intelligence score of 40.

#Model$/1M
1DeepSeek V4 Flash$0.168
2DeepSeek V4 Flash 0731$0.175
3Hy3$0.241
4GPT-5.6 Luna$0.45
5MiniMax-M3$0.525
6Inkling Small$0.525
7DeepSeek V4 Pro$0.544
8MiMo-V2.5-Pro$0.544
9Nex-N2-Pro$1.00
10Qwen3.6 Plus$1.13