The Inference Report

October 6, 2026

The SWE-rebench rankings show no movement from the previous cycle, with AnthropicFable 5 maintaining its lead at 64.5 percent and the top seven models holding identical positions and scores. Conversely, the Artificial Analysis benchmark exhibits substantial churn throughout its 468-entry list, with models reordering across the full range: A.X-K2 drops from position 104 to 114, Ling-3.0-flash-Fin climbs from 105 to 104, and dozens of mid-tier entries shuffle positions by single digits. The stability on SWE-rebench suggests either that the test captures a genuine performance plateau among leading models, that confidence intervals overlap sufficiently to mask real differences, or that the benchmark's methodology constrains the signal available for discrimination. The volatility in Artificial Analysis, by contrast, points toward either greater sensitivity in its evaluation protocol or systematic variance across evaluation runs. Without visibility into the specific test conditions, sample sizes, or whether these benchmarks measure overlapping capabilities, the divergence between frozen and fluid rankings resists clean interpretation: SWE-rebench's consistency could reflect robustness or insensitivity, while Artificial Analysis's movement could reflect responsiveness or noise. The two benchmarks appear to be capturing different aspects of model behavior, or operating under different levels of measurement precision, rather than converging on a shared picture of the field.

Cole Brennan

Daily rankings from SWE-rebench, a benchmark designed to fairly compare LLM capabilities on real-world software engineering tasks. Unlike other evaluations, it uses a standardized scaffolding for all models, continuously updates its dataset to prevent contamination, and runs each model five times to account for stochastic variance.

#ModelScore
1AnthropicFable 5 [high]Model64.5%± 1.41%
2GrokGrok 4.5 [high]Model63.8%± 0.60%
3AnthropicOpus 5 [high]Model63.4%± 1.35%
4Z.aiGLM-5.2 [high]Model62.9%± 1.19%
5OpenAIGPT-5.6 Sol [medium]Model62.3%± 1.83%
6JunieJunieAgent61.8%± 0.54%
7AnthropicClaude CodeAgent60.4%± 1.03%
8OpenAICodexAgent58.0%± 1.29%
9AnthropicSonnet 5 [high]Model56.8%± 0.94%
10CursorCursorAgent51.7%± 0.84%

Artificial Analysis composite index across coding, math, and reasoning benchmarks.

#ModelScoretok/s$/1M
1Claude Opus 5.557.697$8.00
2Claude Sonnet 5.556137$4.00
3Claude Fable 5.153.467$20.00
4GPT-6 Astra52.764$20.00
5Gemini 4 Argon52.60$4.00
6GPT-6.1 Sol51.859$4.00
7Claude Opus 550.80$10.00
8Claude Fable 549.60$20.00
9Muse Spark 1.348.1148$2.00
10GPT-6 Sol47.60$4.00

Output tokens per second — higher is faster. Minimum intelligence score of 40.

#Modeltok/s
1Ling 3.1 Flash216
2Gemini 3.8 Flash209
3Muse Spark 1.3148
4Claude Sonnet 5.5137
5GPT-5.6 Terra111
6Claude Opus 5.597
7Step 5 Preview90
8GLM-5.373
9Grok 4.771
10Claude Fable 5.167

Blended cost per 1M tokens (3:1 input/output) — lower is cheaper. Minimum intelligence score of 40.

#Model$/1M
1GLM 5.3 Flash$0.237
2Ling 3.1 Flash$0.45
3MiMo-V2.6-Pro$0.544
4Step 5 Preview$1.43
5Gemini 3.8 Flash$1.50
6Muse Spark 1.3$2.00
7GLM-5.3$2.15
8Grok 4.7$3.00
9Qwen3.8 Max$3.00
10Grok 4.6$3.00