The Inference Report

September 20, 2026

The SWE-rebench rankings hold steady at the top, with AnthropicFable 5 maintaining 64.5% and the next four models unchanged in position and score, suggesting the coding agent benchmark has stabilized around these performers. On the Artificial Analysis evaluation, the movement is modest but directional: GPT-6 Astra slipped from 52.8 to 52.7, Claude Fable 5 dropped from 49.7 to 49.6, and most models in the upper tiers lost 0.1 to 0.3 points, indicating either marginal measurement variance or a tightening of performance gaps as the field matures. Lower-ranked models show more volatility, with Inkling Small jumping from 26.1 to 27.8 (rank 69 to 57), Ling 3.0 Tiny climbing from 11.9 to 15.3 (rank 194 to 151), and Gemma 4 31B surging from 15.4 to 19.0 (rank 149 to 125), suggesting either improved model versions or recalibration of the evaluation methodology for smaller models. The Artificial Analysis benchmark's broader coverage (459 models versus 17 on SWE-rebench) means its shifts reflect both genuine capability changes and the inherent noise of a more distributed leaderboard, where single-point movements can swing rankings significantly. Without access to confidence intervals on the Artificial Analysis scores, it is difficult to distinguish signal from noise in the middle and lower tiers, and the modest fractional changes at the top suggest the benchmark lacks sufficient granularity to detect meaningful progress in the current generation of frontier models.

Cole Brennan

Daily rankings from SWE-rebench, a benchmark designed to fairly compare LLM capabilities on real-world software engineering tasks. Unlike other evaluations, it uses a standardized scaffolding for all models, continuously updates its dataset to prevent contamination, and runs each model five times to account for stochastic variance.

#ModelScore
1AnthropicFable 5 [high]Model64.5%± 1.41%
2GrokGrok 4.5 [high]Model63.8%± 0.60%
3AnthropicOpus 5 [high]Model63.4%± 1.35%
4Z.aiGLM-5.2 [high]Model62.9%± 1.19%
5OpenAIGPT-5.6 Sol [medium]Model62.3%± 1.83%
6JunieJunieAgent61.8%± 0.54%
7AnthropicClaude CodeAgent60.4%± 1.03%
8OpenAICodexAgent58.0%± 1.29%
9AnthropicSonnet 5 [high]Model56.8%± 0.94%
10CursorCursorAgent51.7%± 0.84%

Artificial Analysis composite index across coding, math, and reasoning benchmarks.

#ModelScoretok/s$/1M
1Claude Fable 5.153.473$20.00
2GPT-6 Astra52.760$20.00
3Claude Opus 550.854$10.00
4Claude Fable 549.60$20.00
5Muse Spark 1.348.1251$2.00
6GPT-5.6 Sol4767$8.00
7Qwen3.8 Max45.439$3.00
8GLM-5.344.877$2.15
9Grok 4.644.358$3.00
10Step 5 Preview43.793$1.43

Output tokens per second — higher is faster. Minimum intelligence score of 40.

#Modeltok/s
1Gemini 3.8 Flash314
2Muse Spark 1.3251
3Step 5 Preview93
4GPT-5.6 Terra93
5GLM 5.3 Flash92
6GLM-5.377
7Claude Fable 5.173
8GPT-5.6 Sol67
9GPT-6 Astra60
10Grok 4.658

Blended cost per 1M tokens (3:1 input/output) — lower is cheaper. Minimum intelligence score of 40.

#Model$/1M
1GLM 5.3 Flash$0.237
2Step 5 Preview$1.43
3Gemini 3.8 Flash$1.50
4Muse Spark 1.3$2.00
5GLM-5.3$2.15
6Qwen3.8 Max$3.00
7Grok 4.6$3.00
8GPT-5.6 Terra$4.50
9Kimi K3$6.00
10GPT-5.6 Sol$8.00