The Inference Report

September 4, 2026

The SWE-rebench rankings show no movement from the previous cycle: the top twelve entries remain in identical positions with matching scores and confidence intervals. AnthropicFable 5 holds first at 64.5% ± 1.41%, followed by GrokGrok 4.5 at 63.8% ± 0.60% and AnthropicOpus 5 at 63.4% ± 1.35%, with all downstream positions locked through rank seventeen. The stability across a full benchmark cycle suggests either that the test set has reached saturation or that the underlying model capabilities have plateaued relative to the task difficulty. On Artificial Analysis, by contrast, substantial reshuffling occurred in the top tier: Muse Spark 1.3 and GPT-6 Astra entered the ranking as new entries at positions three and five respectively, pushing prior occupants down, while Claude Fable 5.1 moved to first place at 65.7 from a prior position lower in the list. The broader Artificial Analysis leaderboard shows consistent but minor position shifts throughout the middle ranks, with no entries dropping off entirely and only two confirmed new additions in the visible window. The divergence between the two benchmarks, perfect stability on SWE-rebench versus active reordering on Artificial Analysis, raises a methodological question: SWE-rebench's narrow confidence intervals and lack of movement across its top entries suggest either tighter experimental controls or a ceiling effect that Artificial Analysis does not exhibit. Without access to sample sizes or retest protocols for either benchmark, it remains unclear whether the SWE-rebench plateau reflects genuine convergence in model performance or methodological constraints that limit discriminative power at the frontier.

Cole Brennan

Daily rankings from SWE-rebench, a benchmark designed to fairly compare LLM capabilities on real-world software engineering tasks. Unlike other evaluations, it uses a standardized scaffolding for all models, continuously updates its dataset to prevent contamination, and runs each model five times to account for stochastic variance.

#ModelScore
1AnthropicFable 5 [high]Model64.5%± 1.41%
2GrokGrok 4.5 [high]Model63.8%± 0.60%
3AnthropicOpus 5 [high]Model63.4%± 1.35%
4Z.aiGLM-5.2 [high]Model62.9%± 1.19%
5OpenAIGPT-5.6 Sol [medium]Model62.3%± 1.83%
6JunieJunieAgent61.8%± 0.54%
7AnthropicClaude CodeAgent60.4%± 1.03%
8OpenAICodexAgent58.0%± 1.29%
9AnthropicSonnet 5 [high]Model56.8%± 0.94%
10CursorCursorAgent51.7%± 0.84%

Artificial Analysis composite index across coding, math, and reasoning benchmarks.

#ModelScoretok/s$/1M
1Claude Fable 5.165.770$20.00
2Claude Opus 563.149$10.00
3Muse Spark 1.362.10$0.00
4Claude Fable 562.167$20.00
5GPT-6 Astra61.20$20.00
6GPT-5.6 Sol60.980$8.00
7Grok 4.660.960$3.00
8Kimi K359.738$6.00
9GLM-5.359.574$2.15
10Gemini 3.8 Flash58.7312$1.50

Output tokens per second — higher is faster. Minimum intelligence score of 40.

#Modeltok/s
1Gemini 3.8 Flash312
2Gemini 3.7 Flash292
3Muse Spark 1.2201
4Gemini 3.6 Flash188
5Quasar 438B187
6Agnes 2.5 Pro Beta157
7Nex-N2-Pro134
8DeepSeek V4 Flash 0731130
9GPT-5.3 Codex126
10Gemini 3.1 Pro Preview117

Blended cost per 1M tokens (3:1 input/output) — lower is cheaper. Minimum intelligence score of 40.

#Model$/1M
1Agnes 2.5 Pro Beta$0.15
2DeepSeek V4 Flash$0.168
3Qwen3.8-Flash-Next$0.23
4GLM-5.3-Flash$0.237
5Hy3$0.241
6GPT-5.6 Luna$0.45
7MiniMax-M3$0.525
8Solar Pro 4$0.525
9Inkling Small$0.525
10DeepSeek V4 Pro$0.544