The Inference Report

August 28, 2026

The SWE-rebench results show no movement since the previous cycle: AnthropicFable 5 holds position one at 64.5% with a confidence interval of 1.41%, followed by GrokGrok 4.5 at 63.8% and AnthropicOpus 5 at 63.4%, identical to prior rankings. The Artificial Analysis leaderboard exhibits modest churn across its 442-entry roster, with two new entrants, Granite 4.2 30B at position 142 (23.7) and Agnes 2.5 Pro Beta at position 29 (49.1), displacing models that previously held those slots. The coding benchmark's stability across top performers suggests either convergence toward a ceiling or insufficient temporal resolution to detect meaningful gains; the tight error bars (0.54 to 1.83 percentage points) indicate the measurements themselves are reliable, but the lack of differentiation between cycles raises a question about whether SWE-rebench remains sensitive to model improvements at this performance tier. Conversely, the Artificial Analysis benchmark's broader movement pattern, particularly the entry of newer variants like Granite 4.2 and Agnes 2.5 Pro Beta, implies that general-capability benchmarks continue to register iterative progress across the field, though the magnitude of individual shifts remains modest and the methodology underlying Artificial Analysis scores is opaque, making it difficult to assess whether the observed reordering reflects genuine capability changes or variance in evaluation conditions.

Cole Brennan

Daily rankings from SWE-rebench, a benchmark designed to fairly compare LLM capabilities on real-world software engineering tasks. Unlike other evaluations, it uses a standardized scaffolding for all models, continuously updates its dataset to prevent contamination, and runs each model five times to account for stochastic variance.

#ModelScore
1AnthropicFable 5 [high]Model64.5%± 1.41%
2GrokGrok 4.5 [high]Model63.8%± 0.60%
3AnthropicOpus 5 [high]Model63.4%± 1.35%
4Z.aiGLM-5.2 [high]Model62.9%± 1.19%
5OpenAIGPT-5.6 Sol [medium]Model62.3%± 1.83%
6JunieJunieAgent61.8%± 0.54%
7AnthropicClaude CodeAgent60.4%± 1.03%
8OpenAICodexAgent58.0%± 1.29%
9AnthropicSonnet 5 [high]Model56.8%± 0.94%
10CursorCursorAgent51.7%± 0.84%

Artificial Analysis composite index across coding, math, and reasoning benchmarks.

#ModelScoretok/s$/1M
1Claude Opus 563.157$10.00
2Claude Fable 562.171$20.00
3GPT-5.6 Sol60.978$8.00
4Grok 4.660.960$3.00
5Kimi K359.738$6.00
6GLM-5.359.566$2.15
7Qwen3.8 Max58.126$3.00
8Qwen3.8 2.4T A95B57.724$3.00
9GLM-5.3-Flash57.549$0.237
10Claude Opus 4.857.30$10.00

Output tokens per second — higher is faster. Minimum intelligence score of 40.

#Modeltok/s
1Gemini 3.7 Flash399
2Gemini 3.6 Flash186
3Agnes 2.5 Pro Beta141
4Nex-N2-Pro139
5DeepSeek V4 Flash 0731137
6GPT-5.3 Codex126
7GPT-5.6 Luna123
8Gemini 3.1 Pro Preview123
9DeepSeek V4 Flash Vision117
10GPT-5.6 Terra107

Blended cost per 1M tokens (3:1 input/output) — lower is cheaper. Minimum intelligence score of 40.

#Model$/1M
1Agnes 2.5 Pro Beta$0.15
2DeepSeek V4 Flash$0.168
3Qwen3.8-Flash-Next$0.23
4GLM-5.3-Flash$0.237
5Hy3$0.241
6GPT-5.6 Luna$0.45
7MiniMax-M3$0.525
8Solar Pro 4$0.525
9Inkling Small$0.525
10DeepSeek V4 Pro$0.544