The Inference Report

September 19, 2026

The SWE-rebench rankings remain static across both snapshots, with AnthropicFable 5 holding 64.5% and the top seventeen models unchanged in position and score. This stability itself is noteworthy: the coding benchmark shows no movement whatsoever, suggesting either that evaluation cycles are infrequent or that the models assessed have not undergone meaningful revision. The confidence intervals are tight enough to rule out measurement noise as an explanation for the lack of change (GrokGrok 4.5's ±0.60% is especially precise), indicating the benchmark has sufficient resolution to detect real differences. By contrast, the Artificial Analysis rankings show extensive reshuffling throughout the list, with Step 5 Preview entering at #11, GPT-5.6 Terra dropping from #11 to #12, and cascading shifts down the entire leaderboard. This divergence between benchmarks raises a methodological question: SWE-rebench appears to measure something more stable or resistant to model iteration than the Artificial Analysis suite, or the two are sampling fundamentally different capability dimensions. Without knowing the composition or evaluation protocol of either benchmark, it is difficult to assess whether the SWE-rebench stasis reflects genuine plateau in coding agent performance or simply less frequent or less sensitive evaluation. The Artificial Analysis churn, conversely, could indicate either rapid iteration in the broader model ecosystem or greater sensitivity to minor capability shifts, but the absence of clear winners and losers across the full 459-model ranking makes it hard to extract signal about what has actually improved.

Cole Brennan

Daily rankings from SWE-rebench, a benchmark designed to fairly compare LLM capabilities on real-world software engineering tasks. Unlike other evaluations, it uses a standardized scaffolding for all models, continuously updates its dataset to prevent contamination, and runs each model five times to account for stochastic variance.

#ModelScore
1AnthropicFable 5 [high]Model64.5%± 1.41%
2GrokGrok 4.5 [high]Model63.8%± 0.60%
3AnthropicOpus 5 [high]Model63.4%± 1.35%
4Z.aiGLM-5.2 [high]Model62.9%± 1.19%
5OpenAIGPT-5.6 Sol [medium]Model62.3%± 1.83%
6JunieJunieAgent61.8%± 0.54%
7AnthropicClaude CodeAgent60.4%± 1.03%
8OpenAICodexAgent58.0%± 1.29%
9AnthropicSonnet 5 [high]Model56.8%± 0.94%
10CursorCursorAgent51.7%± 0.84%

Artificial Analysis composite index across coding, math, and reasoning benchmarks.

#ModelScoretok/s$/1M
1Claude Fable 5.153.471$20.00
2GPT-6 Astra52.856$20.00
3Claude Opus 550.753$10.00
4Claude Fable 549.766$20.00
5Muse Spark 1.348.2228$2.00
6GPT-5.6 Sol47.160$8.00
7Qwen3.8 Max45.441$3.00
8GLM-5.344.963$2.15
9Grok 4.644.455$3.00
10Kimi K343.838$6.00

Output tokens per second — higher is faster. Minimum intelligence score of 40.

#Modeltok/s
1Gemini 3.8 Flash306
2Muse Spark 1.3228
3GLM 5.3 Flash97
4Step 5 Preview93
5GPT-5.6 Terra85
6Claude Fable 5.171
7Claude Fable 566
8GLM-5.363
9GPT-5.6 Sol60
10GPT-6 Astra56

Blended cost per 1M tokens (3:1 input/output) — lower is cheaper. Minimum intelligence score of 40.

#Model$/1M
1GLM 5.3 Flash$0.237
2Step 5 Preview$1.43
3Gemini 3.8 Flash$1.50
4Muse Spark 1.3$2.00
5GLM-5.3$2.15
6Qwen3.8 Max$3.00
7Grok 4.6$3.00
8Qwen3.8 2.4T A95B$3.00
9GPT-5.6 Terra$4.50
10Kimi K3$6.00