The Inference Report

September 5, 2026

The SWE-rebench standings show minimal movement at the top, with AnthropicFable 5 holding the lead at 64.5% (±1.41%) and the top six models clustered within 2.5 percentage points, but the Artificial Analysis benchmark reveals a sharp, systematic decline across the entire leaderboard that defies the stability suggested by SWE-rebench alone. Claude Fable 5.1 dropped from 65.7 to 56.8, a 13% relative decline, while GPT-6 Astra fell from 61.2 to 54.7, and Claude Opus 5 from 63.1 to 54.1, suggesting either a methodological shift in Artificial Analysis's evaluation protocol or a recalibration of its test set that hit all models uniformly rather than differentially. The pattern is not random: every model in the Artificial Analysis top 100 lost between 8 and 15 percentage points, with the median drop around 11 points, while lower-ranked models (those below rank 200) showed smaller absolute losses, implying the benchmark either tightened its criteria, introduced harder test cases, or corrected for prior inflation. The SWE-rebench data, by contrast, appears internally consistent with its previous iteration, suggesting these are genuinely separate evaluation regimes measuring different aspects of code generation capability. Without access to Artificial Analysis's methodology change documentation, the most parsimonious explanation is that one benchmark recalibrated while the other did not, making cross-benchmark comparisons unreliable and highlighting the risk of relying on a single leaderboard source for model assessment.

Cole Brennan

Daily rankings from SWE-rebench, a benchmark designed to fairly compare LLM capabilities on real-world software engineering tasks. Unlike other evaluations, it uses a standardized scaffolding for all models, continuously updates its dataset to prevent contamination, and runs each model five times to account for stochastic variance.

#ModelScore
1AnthropicFable 5 [high]Model64.5%± 1.41%
2GrokGrok 4.5 [high]Model63.8%± 0.60%
3AnthropicOpus 5 [high]Model63.4%± 1.35%
4Z.aiGLM-5.2 [high]Model62.9%± 1.19%
5OpenAIGPT-5.6 Sol [medium]Model62.3%± 1.83%
6JunieJunieAgent61.8%± 0.54%
7AnthropicClaude CodeAgent60.4%± 1.03%
8OpenAICodexAgent58.0%± 1.29%
9AnthropicSonnet 5 [high]Model56.8%± 0.94%
10CursorCursorAgent51.7%± 0.84%

Artificial Analysis composite index across coding, math, and reasoning benchmarks.

#ModelScoretok/s$/1M
1Claude Fable 5.156.871$20.00
2GPT-6 Astra54.70$20.00
3Claude Opus 554.152$10.00
4Claude Fable 553.266$20.00
5Muse Spark 1.352.7177$2.00
6GPT-5.6 Sol51.379$8.00
7Grok 4.650.660$3.00
8Kimi K350.238$6.00
9GLM-5.348.678$2.15
10Gemini 3.8 Flash47.1429$1.50

Output tokens per second — higher is faster. Minimum intelligence score of 40.

#Modeltok/s
1Gemini 3.8 Flash429
2Gemini 3.7 Flash295
3Muse Spark 1.2226
4Gemini 3.6 Flash207
5Muse Spark 1.3177
6DeepSeek V4 Flash 0731130
7DeepSeek V4 Flash Vision118
8GPT-5.6 Luna110
9GPT-5.6 Terra104
10Claude Sonnet 581

Blended cost per 1M tokens (3:1 input/output) — lower is cheaper. Minimum intelligence score of 40.

#Model$/1M
1Qwen3.8-Flash-Next$0.23
2GLM-5.3-Flash$0.237
3GPT-5.6 Luna$0.45
4DeepSeek V4 Flash Vision$0.66
5DeepSeek V4 Flash 0731$0.66
6Qwen3.8 27B$1.13
7Gemini 3.8 Flash$1.50
8Gemini 3.7 Flash$1.50
9Gemini 3.6 Flash$1.50
10DeepSeek V4 Pro 0813$1.98