The Inference Report

September 23, 2026

On the SWE-rebench coding benchmark, the rankings remain stable from the previous cycle: AnthropicFable 5 holds position one at 64.5% (±1.41%), followed by GrokGrok 4.5 at 63.8% (±0.60%), and AnthropicOpus 5 at 63.4% (±1.35%). The top seven positions cluster tightly between 60.4% and 64.5%, with confidence intervals that overlap substantially, making any claim of meaningful separation within this tier premature given the measurement uncertainty. The Artificial Analysis benchmark tells a different story: Claude Opus 5.5 enters at position one with a score of 57.6, displacing Claude Fable 5.1 to second place at 53.4, while GPT-6 Astra moves to third at 52.7. The new entries Claude Opus 5.5 and GPT-6 Sol (at 47.5, position seven) suggest recent model releases, though the Artificial Analysis methodology differs fundamentally from SWE-rebench and measures different problem classes, making direct cross-benchmark comparison invalid. Within each benchmark independently, the movements are modest: on SWE-rebench, the top performers show no ranking changes from the previous cycle, indicating evaluation consistency or saturation at the frontier; on Artificial Analysis, the displacement of Fable 5.1 by Opus 5.5 represents a 4.2-point gain that, while directional, lacks context about whether this reflects genuine capability advancement or variation in task difficulty. Without details on the Artificial Analysis evaluation methodology, confidence intervals, or the specific problems tested, whether these shifts constitute meaningful progress or normal measurement variance remains unclear.

Cole Brennan

Daily rankings from SWE-rebench, a benchmark designed to fairly compare LLM capabilities on real-world software engineering tasks. Unlike other evaluations, it uses a standardized scaffolding for all models, continuously updates its dataset to prevent contamination, and runs each model five times to account for stochastic variance.

#ModelScore
1AnthropicFable 5 [high]Model64.5%± 1.41%
2GrokGrok 4.5 [high]Model63.8%± 0.60%
3AnthropicOpus 5 [high]Model63.4%± 1.35%
4Z.aiGLM-5.2 [high]Model62.9%± 1.19%
5OpenAIGPT-5.6 Sol [medium]Model62.3%± 1.83%
6JunieJunieAgent61.8%± 0.54%
7AnthropicClaude CodeAgent60.4%± 1.03%
8OpenAICodexAgent58.0%± 1.29%
9AnthropicSonnet 5 [high]Model56.8%± 0.94%
10CursorCursorAgent51.7%± 0.84%

Artificial Analysis composite index across coding, math, and reasoning benchmarks.

#ModelScoretok/s$/1M
1Claude Opus 5.557.60$8.00
2Claude Fable 5.153.464$20.00
3GPT-6 Astra52.754$20.00
4Claude Opus 550.852$10.00
5Claude Fable 549.60$20.00
6Muse Spark 1.348.1220$2.00
7GPT-6 Sol47.5116$4.00
8GPT-5.6 Sol4762$8.00
9Grok 4.746.450$3.00
10MiMo-V2.6-Pro46.354$0.544

Output tokens per second — higher is faster. Minimum intelligence score of 40.

#Modeltok/s
1Gemini 3.8 Flash301
2Muse Spark 1.3220
3GPT-6 Sol116
4GPT-5.6 Terra89
5Step 5 Preview75
6Grok 4.670
7Claude Fable 5.164
8GLM-5.363
9GPT-5.6 Sol62
10GPT-6 Astra54

Blended cost per 1M tokens (3:1 input/output) — lower is cheaper. Minimum intelligence score of 40.

#Model$/1M
1GLM 5.3 Flash$0.237
2MiMo-V2.6-Pro$0.544
3Step 5 Preview$1.43
4Gemini 3.8 Flash$1.50
5Muse Spark 1.3$2.00
6GLM-5.3$2.15
7Grok 4.7$3.00
8Qwen3.8 Max$3.00
9Grok 4.6$3.00
10GPT-6 Sol$4.00