The Inference Report

September 29, 2026

The SWE-rebench rankings remain static across the top tier, with AnthropicFable 5 holding 64.5% ± 1.41% at first place and the next six positions unchanged through OpenAICodexAgent at 58.0% ± 1.29%. This stability in the coding benchmark suggests either that the test set has reached saturation at the frontier or that recent model updates have not substantially altered performance on this particular evaluation. The Artificial Analysis benchmark, by contrast, shows considerable churn: Claude Sonnet 5.5 enters at rank 2 with 56.0, displacing Claude Fable 5.1 down one position to 3, while the rest of the top 20 shifts accordingly with no score changes, only ranking adjustments. The gap between the two benchmarks is worth noting. SWE-rebench scores cluster tightly in the 60s and 50s for leading models, with error bars ranging from ±0.54% to ±1.83%, reflecting controlled experimental conditions. Artificial Analysis scores begin at 57.6 for Claude Opus 5.5 and decline more steeply down the list, reaching single digits by rank 250. This divergence hints at different task difficulty profiles: SWE-rebench may be measuring a narrower, more saturated capability space (software engineering problem solving with defined test cases), while Artificial Analysis appears to cover a broader skill distribution. Neither benchmark reveals methodological details in the provided data, making it difficult to assess whether the stability in SWE-rebench reflects genuine parity or simply coarse-grained measurement. The entry of Claude Sonnet 5.5 into Artificial Analysis's top two is the only concrete movement worth tracking, but without prior scores for this model on that benchmark, its significance cannot be determined.

Cole Brennan

Daily rankings from SWE-rebench, a benchmark designed to fairly compare LLM capabilities on real-world software engineering tasks. Unlike other evaluations, it uses a standardized scaffolding for all models, continuously updates its dataset to prevent contamination, and runs each model five times to account for stochastic variance.

#ModelScore
1AnthropicFable 5 [high]Model64.5%± 1.41%
2GrokGrok 4.5 [high]Model63.8%± 0.60%
3AnthropicOpus 5 [high]Model63.4%± 1.35%
4Z.aiGLM-5.2 [high]Model62.9%± 1.19%
5OpenAIGPT-5.6 Sol [medium]Model62.3%± 1.83%
6JunieJunieAgent61.8%± 0.54%
7AnthropicClaude CodeAgent60.4%± 1.03%
8OpenAICodexAgent58.0%± 1.29%
9AnthropicSonnet 5 [high]Model56.8%± 0.94%
10CursorCursorAgent51.7%± 0.84%

Artificial Analysis composite index across coding, math, and reasoning benchmarks.

#ModelScoretok/s$/1M
1Claude Opus 5.557.696$8.00
2Claude Sonnet 5.556145$4.00
3Claude Fable 5.153.469$20.00
4GPT-6 Astra52.762$20.00
5Claude Opus 550.80$10.00
6Claude Fable 549.60$20.00
7Muse Spark 1.348.1189$2.00
8GPT-6 Sol47.584$4.00
9GPT-5.6 Sol470$8.00
10Grok 4.746.472$3.00

Output tokens per second — higher is faster. Minimum intelligence score of 40.

#Modeltok/s
1Gemini 3.8 Flash238
2Muse Spark 1.3189
3Claude Sonnet 5.5145
4GPT-5.6 Terra109
5Claude Opus 5.596
6GPT-6 Sol84
7Step 5 Preview84
8GLM-5.375
9Grok 4.675
10Grok 4.772

Blended cost per 1M tokens (3:1 input/output) — lower is cheaper. Minimum intelligence score of 40.

#Model$/1M
1GLM 5.3 Flash$0.237
2MiMo-V2.6-Pro$0.544
3Step 5 Preview$1.43
4Gemini 3.8 Flash$1.50
5Muse Spark 1.3$2.00
6GLM-5.3$2.15
7Grok 4.7$3.00
8Qwen3.8 Max$3.00
9Grok 4.6$3.00
10Claude Sonnet 5.5$4.00