The Inference Report

August 29, 2026

The SWE-rebench results show no movement in the top tier, with Anthropic Fable 5 maintaining 64.5% ± 1.41% at rank one and the full top ten remaining locked in place. Grok 4.5 stays at 63.8% ± 0.60%, Opus 5 at 63.4% ± 1.35%, and the spread between positions one and ten spans 12.8 percentage points. The Artificial Analysis benchmark, by contrast, exhibits substantial churn across its 442 entries, with Claude Opus 5 leading at 63.1 and K-EXAONE 2.0 newly entering at rank 113 with 31.0, while K-EXAONE 2.0 0803 simultaneously dropped from that same position. This divergence signals a methodological gap worth noting: SWE-rebench isolates a narrow, high-signal task (software engineering problem resolution) where the top performers have reached a plateau tight enough that uncertainty intervals overlap, whereas Artificial Analysis aggregates across a broader evaluation surface where models shuffle constantly. The SWE-rebench stability is not stagnation but rather statistical saturation among frontier systems; further discrimination at this level would require either larger sample sizes to tighten confidence bands or harder problem instances. The Artificial Analysis flux, meanwhile, reflects sensitivity to model versioning and evaluation scope, making it less reliable for tracking real capability shifts in code-specific reasoning.

Cole Brennan

Daily rankings from SWE-rebench, a benchmark designed to fairly compare LLM capabilities on real-world software engineering tasks. Unlike other evaluations, it uses a standardized scaffolding for all models, continuously updates its dataset to prevent contamination, and runs each model five times to account for stochastic variance.

#ModelScore
1AnthropicFable 5 [high]Model64.5%± 1.41%
2GrokGrok 4.5 [high]Model63.8%± 0.60%
3AnthropicOpus 5 [high]Model63.4%± 1.35%
4Z.aiGLM-5.2 [high]Model62.9%± 1.19%
5OpenAIGPT-5.6 Sol [medium]Model62.3%± 1.83%
6JunieJunieAgent61.8%± 0.54%
7AnthropicClaude CodeAgent60.4%± 1.03%
8OpenAICodexAgent58.0%± 1.29%
9AnthropicSonnet 5 [high]Model56.8%± 0.94%
10CursorCursorAgent51.7%± 0.84%

Artificial Analysis composite index across coding, math, and reasoning benchmarks.

#ModelScoretok/s$/1M
1Claude Opus 563.153$10.00
2Claude Fable 562.165$20.00
3GPT-5.6 Sol60.981$8.00
4Grok 4.660.960$3.00
5Kimi K359.736$6.00
6GLM-5.359.575$2.15
7Qwen3.8 Max58.127$3.00
8Qwen3.8 2.4T A95B57.724$3.00
9GLM-5.3-Flash57.549$0.237
10Claude Opus 4.857.30$10.00

Output tokens per second — higher is faster. Minimum intelligence score of 40.

#Modeltok/s
1Gemini 3.7 Flash376
2Gemini 3.6 Flash177
3Agnes 2.5 Pro Beta151
4Nex-N2-Pro137
5DeepSeek V4 Flash 0731135
6GPT-5.3 Codex126
7GPT-5.6 Luna124
8Gemini 3.1 Pro Preview122
9DeepSeek V4 Flash Vision117
10GPT-5.6 Terra107

Blended cost per 1M tokens (3:1 input/output) — lower is cheaper. Minimum intelligence score of 40.

#Model$/1M
1Agnes 2.5 Pro Beta$0.15
2DeepSeek V4 Flash$0.168
3Qwen3.8-Flash-Next$0.23
4GLM-5.3-Flash$0.237
5Hy3$0.241
6GPT-5.6 Luna$0.45
7MiniMax-M3$0.525
8Solar Pro 4$0.525
9Inkling Small$0.525
10DeepSeek V4 Pro$0.544