The Inference Report

September 22, 2026

The SWE-rebench rankings remain static across both snapshots, with AnthropicFable 5 holding first place at 64.5% and the top seventeen models unchanged in position and score. The Artificial Analysis benchmark, by contrast, shows wholesale reshuffling: Grok 4.7 and MiMo-V2.6-Pro enter the top ten as new entries at positions 7 and 8, displacing Qwen3.8 Max and GLM-5.3 downward by two spots each, while every other model in the 300-entry list shifts position despite most scores remaining identical to the previous cycle. This pattern suggests the Artificial Analysis ranking operates on a different methodology than SWE-rebench, likely incorporating recency weighting, model release date, or other temporal factors that cause constant reordering without corresponding score changes, whereas SWE-rebench appears to measure against a fixed problem set with genuine performance stability at the top. The stability of SWE-rebench scores across the full range, paired with tight confidence intervals (most under 1.5%), indicates controlled experimental conditions and reproducible results; Artificial Analysis's perpetual repositioning of identical-scoring models raises questions about whether the ranking reflects meaningful differentiation or administrative sorting unmoored from the underlying evaluation data. Without visibility into how Artificial Analysis recalculates rankings when scores don't change, treating its movement as a signal of capability shift would be premature.

Cole Brennan

Daily rankings from SWE-rebench, a benchmark designed to fairly compare LLM capabilities on real-world software engineering tasks. Unlike other evaluations, it uses a standardized scaffolding for all models, continuously updates its dataset to prevent contamination, and runs each model five times to account for stochastic variance.

#ModelScore
1AnthropicFable 5 [high]Model64.5%± 1.41%
2GrokGrok 4.5 [high]Model63.8%± 0.60%
3AnthropicOpus 5 [high]Model63.4%± 1.35%
4Z.aiGLM-5.2 [high]Model62.9%± 1.19%
5OpenAIGPT-5.6 Sol [medium]Model62.3%± 1.83%
6JunieJunieAgent61.8%± 0.54%
7AnthropicClaude CodeAgent60.4%± 1.03%
8OpenAICodexAgent58.0%± 1.29%
9AnthropicSonnet 5 [high]Model56.8%± 0.94%
10CursorCursorAgent51.7%± 0.84%

Artificial Analysis composite index across coding, math, and reasoning benchmarks.

#ModelScoretok/s$/1M
1Claude Fable 5.153.471$20.00
2GPT-6 Astra52.762$20.00
3Claude Opus 550.857$10.00
4Claude Fable 549.60$20.00
5Muse Spark 1.348.1281$2.00
6GPT-5.6 Sol4773$8.00
7Grok 4.746.444$3.00
8MiMo-V2.6-Pro46.3116$0.544
9Qwen3.8 Max45.442$3.00
10GLM-5.344.865$2.15

Output tokens per second — higher is faster. Minimum intelligence score of 40.

#Modeltok/s
1Gemini 3.8 Flash365
2Muse Spark 1.3281
3MiMo-V2.6-Pro116
4GPT-5.6 Terra109
5Step 5 Preview93
6GLM 5.3 Flash79
7GPT-5.6 Sol73
8Claude Fable 5.171
9Grok 4.670
10GLM-5.365

Blended cost per 1M tokens (3:1 input/output) — lower is cheaper. Minimum intelligence score of 40.

#Model$/1M
1GLM 5.3 Flash$0.237
2MiMo-V2.6-Pro$0.544
3Step 5 Preview$1.43
4Gemini 3.8 Flash$1.50
5Muse Spark 1.3$2.00
6GLM-5.3$2.15
7Grok 4.7$3.00
8Qwen3.8 Max$3.00
9Grok 4.6$3.00
10GPT-5.6 Terra$4.50