The Inference Report

September 8, 2026

The SWE-rebench results remain static across the top tier, with Anthropic Fable 5 holding 64.5% and the next four positions unchanged through Grok 4.5, Opus 5, GLM-5.2, and GPT-5.6 Sol, all within their confidence intervals from the previous cycle. The Artificial Analysis benchmark shows more churn in the middle ranks but reveals a pattern worth noting: Claude Opus 4.8 gained 0.4 points to reach 47.8, Claude Opus 4.7 held steady at 44.3, and Claude Sonnet 5 remained at 45.1, suggesting Anthropic's line is consolidating performance across different model sizes. DeepSeek V3 jumped 2.1 points from 8.3 to 10.4, placing it at #212 and marking the largest single-step gain in the visible range. Qwen3 32B moved from 5.7 to 8.0 (entry #232), a 2.3-point improvement. Lower down the list, gpt-oss-20b gained 1.7 points to 10.8, and Trinity Large Thinking moved from #190 to #181 with a 1.1-point gain to 13.1. The methodology here differs between benchmarks: SWE-rebench measures code generation on real software engineering tasks with confidence intervals, while Artificial Analysis reports single-point scores without error bounds, making the two incomparable directly. Neither benchmark shows the kind of concentrated improvement at the top that would suggest a fundamental breakthrough. The gains are distributed, incremental, and within the noise of typical model iteration.

Cole Brennan

Daily rankings from SWE-rebench, a benchmark designed to fairly compare LLM capabilities on real-world software engineering tasks. Unlike other evaluations, it uses a standardized scaffolding for all models, continuously updates its dataset to prevent contamination, and runs each model five times to account for stochastic variance.

#ModelScore
1AnthropicFable 5 [high]Model64.5%± 1.41%
2GrokGrok 4.5 [high]Model63.8%± 0.60%
3AnthropicOpus 5 [high]Model63.4%± 1.35%
4Z.aiGLM-5.2 [high]Model62.9%± 1.19%
5OpenAIGPT-5.6 Sol [medium]Model62.3%± 1.83%
6JunieJunieAgent61.8%± 0.54%
7AnthropicClaude CodeAgent60.4%± 1.03%
8OpenAICodexAgent58.0%± 1.29%
9AnthropicSonnet 5 [high]Model56.8%± 0.94%
10CursorCursorAgent51.7%± 0.84%

Artificial Analysis composite index across coding, math, and reasoning benchmarks.

#ModelScoretok/s$/1M
1Claude Fable 5.156.868$20.00
2GPT-6 Astra54.763$20.00
3Claude Opus 554.152$10.00
4Claude Fable 553.262$20.00
5Muse Spark 1.353221$2.00
6GPT-5.6 Sol51.373$8.00
7Grok 4.650.657$3.00
8Kimi K350.242$6.00
9GLM-5.348.675$2.15
10Claude Opus 4.847.80$10.00

Output tokens per second — higher is faster. Minimum intelligence score of 40.

#Modeltok/s
1Gemini 3.8 Flash356
2Gemini 3.7 Flash325
3Muse Spark 1.2262
4Muse Spark 1.3221
5Gemini 3.6 Flash188
6GPT-5.6 Terra121
7DeepSeek V4 Flash Vision120
8DeepSeek V4 Flash 0731118
9GPT-5.6 Luna117
10Claude Sonnet 580

Blended cost per 1M tokens (3:1 input/output) — lower is cheaper. Minimum intelligence score of 40.

#Model$/1M
1Qwen3.8-Flash-Next$0.23
2GLM-5.3-Flash$0.237
3GPT-5.6 Luna$0.45
4DeepSeek V4 Flash 0731$0.66
5DeepSeek V4 Flash Vision$0.66
6Qwen3.8 27B$1.13
7Gemini 3.8 Flash$1.50
8Gemini 3.7 Flash$1.50
9Gemini 3.6 Flash$1.50
10DeepSeek V4 Pro 0813$1.98