The Inference Report

September 6, 2026

The SWE-rebench rankings remain static across both surveys, with AnthropicFable 5 holding the top position at 64.5 percent and the same seventeen models occupying positions one through seventeen in identical order. The Artificial Analysis benchmark shows marginal movement in its broader field: Muse Spark 1.3 gained 0.3 points to reach 53.0, climbing from position five to maintain that slot; Mistral Large 3 moved from 9.7 to 11.1 points and rose from position 211 to 202; Gemini 3.5 Flash fell from 41.9 to 39.7 points, dropping from position 28 to 32; and Gemini 3.6 Flash shifted from position 32 to 31 while holding steady at 40.3 points. The SWE-rebench evaluation uses controlled conditions with reported confidence intervals, though the methodology for the Artificial Analysis benchmark remains opaque regarding test set composition, evaluation protocol, and statistical rigor, making it difficult to assess whether these fractional gains represent meaningful capability differences or measurement noise. The coding agent benchmark's stability at the top tier suggests the frontier models have reached a performance plateau on this task, while the Artificial Analysis results show typical variance patterns consistent with evaluation noise rather than systematic improvement across the broader model landscape.

Cole Brennan

Daily rankings from SWE-rebench, a benchmark designed to fairly compare LLM capabilities on real-world software engineering tasks. Unlike other evaluations, it uses a standardized scaffolding for all models, continuously updates its dataset to prevent contamination, and runs each model five times to account for stochastic variance.

#ModelScore
1AnthropicFable 5 [high]Model64.5%± 1.41%
2GrokGrok 4.5 [high]Model63.8%± 0.60%
3AnthropicOpus 5 [high]Model63.4%± 1.35%
4Z.aiGLM-5.2 [high]Model62.9%± 1.19%
5OpenAIGPT-5.6 Sol [medium]Model62.3%± 1.83%
6JunieJunieAgent61.8%± 0.54%
7AnthropicClaude CodeAgent60.4%± 1.03%
8OpenAICodexAgent58.0%± 1.29%
9AnthropicSonnet 5 [high]Model56.8%± 0.94%
10CursorCursorAgent51.7%± 0.84%

Artificial Analysis composite index across coding, math, and reasoning benchmarks.

#ModelScoretok/s$/1M
1Claude Fable 5.156.870$20.00
2GPT-6 Astra54.761$20.00
3Claude Opus 554.149$10.00
4Claude Fable 553.259$20.00
5Muse Spark 1.353177$2.00
6GPT-5.6 Sol51.379$8.00
7Grok 4.650.657$3.00
8Kimi K350.239$6.00
9GLM-5.348.677$2.15
10Gemini 3.8 Flash47.1340$1.50

Output tokens per second — higher is faster. Minimum intelligence score of 40.

#Modeltok/s
1Gemini 3.8 Flash340
2Gemini 3.7 Flash284
3Muse Spark 1.2224
4Gemini 3.6 Flash186
5Muse Spark 1.3177
6DeepSeek V4 Flash Vision121
7DeepSeek V4 Flash 0731118
8GPT-5.6 Luna110
9GPT-5.6 Terra104
10GPT-5.6 Sol79

Blended cost per 1M tokens (3:1 input/output) — lower is cheaper. Minimum intelligence score of 40.

#Model$/1M
1Qwen3.8-Flash-Next$0.23
2GLM-5.3-Flash$0.237
3GPT-5.6 Luna$0.45
4DeepSeek V4 Flash Vision$0.66
5DeepSeek V4 Flash 0731$0.66
6Qwen3.8 27B$1.13
7Gemini 3.8 Flash$1.50
8Gemini 3.7 Flash$1.50
9Gemini 3.6 Flash$1.50
10DeepSeek V4 Pro 0813$1.98