The Inference Report

October 3, 2026

The SWE-rebench rankings show no movement in the top tier: AnthropicFable 5 holds first at 64.5% ± 1.41%, GrokGrok 4.5 remains second at 63.8% ± 0.60%, and the next five positions remain unchanged through OpenAIGPT-5.6 Sol at 62.3% ± 1.83%. The Artificial Analysis benchmark, by contrast, underwent substantial reorganization across its full 467-model roster, with NVIDIA Nemotron 3 Nano 4B dropping from position 321 to 465, and multiple models in the 320s range shifting position by single or double digits. The SWE-rebench evaluation tests code generation agents on real GitHub issues with controlled execution environments, while Artificial Analysis appears to assess a broader range of model capabilities across a much larger set of entries, making direct comparison difficult. The stability of the SWE-rebench top tier suggests either that these models have reached a performance plateau on that particular task or that the evaluation's variance bands (ranging from ±0.54% to ±1.83%) are narrow enough to prevent ranking churn. The extensive reshuffling in Artificial Analysis, which uses point scores without reported confidence intervals, may reflect either more sensitive discrimination between models or differences in evaluation methodology. Without access to the specific problem distributions, success criteria, or statistical methodology behind each benchmark, it remains unclear whether the SWE-rebench stability reflects genuine performance plateaus or whether the benchmark's design constrains differentiation among top performers.

Cole Brennan

Daily rankings from SWE-rebench, a benchmark designed to fairly compare LLM capabilities on real-world software engineering tasks. Unlike other evaluations, it uses a standardized scaffolding for all models, continuously updates its dataset to prevent contamination, and runs each model five times to account for stochastic variance.

#ModelScore
1AnthropicFable 5 [high]Model64.5%± 1.41%
2GrokGrok 4.5 [high]Model63.8%± 0.60%
3AnthropicOpus 5 [high]Model63.4%± 1.35%
4Z.aiGLM-5.2 [high]Model62.9%± 1.19%
5OpenAIGPT-5.6 Sol [medium]Model62.3%± 1.83%
6JunieJunieAgent61.8%± 0.54%
7AnthropicClaude CodeAgent60.4%± 1.03%
8OpenAICodexAgent58.0%± 1.29%
9AnthropicSonnet 5 [high]Model56.8%± 0.94%
10CursorCursorAgent51.7%± 0.84%

Artificial Analysis composite index across coding, math, and reasoning benchmarks.

#ModelScoretok/s$/1M
1Claude Opus 5.557.698$8.00
2Claude Sonnet 5.556138$4.00
3Claude Fable 5.153.471$20.00
4GPT-6 Astra52.756$20.00
5Gemini 4 Argon52.60$4.00
6GPT-6.1 Sol51.859$4.00
7Claude Opus 550.80$10.00
8Claude Fable 549.60$20.00
9Muse Spark 1.348.1157$2.00
10GPT-6 Sol47.6111$4.00

Output tokens per second — higher is faster. Minimum intelligence score of 40.

#Modeltok/s
1Gemini 3.8 Flash256
2Muse Spark 1.3157
3Claude Sonnet 5.5138
4GPT-6 Sol111
5GPT-5.6 Terra110
6Claude Opus 5.598
7Step 5 Preview84
8Grok 4.780
9GLM-5.375
10Claude Fable 5.171

Blended cost per 1M tokens (3:1 input/output) — lower is cheaper. Minimum intelligence score of 40.

#Model$/1M
1GLM 5.3 Flash$0.237
2MiMo-V2.6-Pro$0.544
3Step 5 Preview$1.43
4Gemini 3.8 Flash$1.50
5Muse Spark 1.3$2.00
6GLM-5.3$2.15
7Grok 4.7$3.00
8Qwen3.8 Max$3.00
9Grok 4.6$3.00
10Claude Sonnet 5.5$4.00