The Inference Report

August 27, 2026

The SWE-rebench leaderboard shows no movement since the previous cycle, with AnthropicFable 5 holding the top position at 64.5 percent and the full ranking unchanged across all seventeen entries. On Artificial Analysis, two new entrants appear: GLM-5.3-Flash enters at position nine with a score of 57.5, pushing Claude Opus 4.8 from ninth to tenth, and Granite 4.2 8B debuts at position 170 with 19.6, followed by Granite 4.2 3B at position 217 with 14.3. The stability in SWE-rebench contrasts sharply with the Artificial Analysis benchmark, where the expanded leaderboard now includes 439 models compared to the previous 437. Within the top tier, scores remain tightly clustered and unchanged: Claude Opus 5 leads at 63.1, followed by Claude Fable 5 at 62.1 and GPT-5.6 Sol and Grok 4.6 both at 60.9. The SWE-rebench methodology, which evaluates coding agents on real software engineering tasks with controlled conditions and reproducible metrics, appears sufficiently mature that week-to-week variation falls within confidence intervals marked by standard errors ranging from 0.54 to 1.83 percentage points. The absence of movement in the top positions suggests either that the tested models have reached a performance plateau on this benchmark or that new model releases are not substantially outpacing existing leaders. Artificial Analysis, by contrast, continues to absorb new model submissions, though the criteria for inclusion and the evaluation methodology are less transparent than SWE-rebench's publicly documented task construction.

Cole Brennan

Daily rankings from SWE-rebench, a benchmark designed to fairly compare LLM capabilities on real-world software engineering tasks. Unlike other evaluations, it uses a standardized scaffolding for all models, continuously updates its dataset to prevent contamination, and runs each model five times to account for stochastic variance.

#ModelScore
1AnthropicFable 5 [high]Model64.5%± 1.41%
2GrokGrok 4.5 [high]Model63.8%± 0.60%
3AnthropicOpus 5 [high]Model63.4%± 1.35%
4Z.aiGLM-5.2 [high]Model62.9%± 1.19%
5OpenAIGPT-5.6 Sol [medium]Model62.3%± 1.83%
6JunieJunieAgent61.8%± 0.54%
7AnthropicClaude CodeAgent60.4%± 1.03%
8OpenAICodexAgent58.0%± 1.29%
9AnthropicSonnet 5 [high]Model56.8%± 0.94%
10CursorCursorAgent51.7%± 0.84%

Artificial Analysis composite index across coding, math, and reasoning benchmarks.

#ModelScoretok/s$/1M
1Claude Opus 563.155$10.00
2Claude Fable 562.168$20.00
3GPT-5.6 Sol60.970$8.00
4Grok 4.660.961$3.00
5Kimi K359.738$6.00
6GLM-5.359.581$2.15
7Qwen3.8 Max58.124$3.00
8Qwen3.8 2.4T A95B57.724$3.00
9GLM-5.3-Flash57.542$0.237
10Claude Opus 4.857.30$10.00

Output tokens per second — higher is faster. Minimum intelligence score of 40.

#Modeltok/s
1Gemini 3.7 Flash366
2Gemini 3.6 Flash197
3Nex-N2-Pro142
4DeepSeek V4 Flash 0731129
5Gemini 3.1 Pro Preview125
6GPT-5.6 Luna124
7GPT-5.3 Codex121
8DeepSeek V4 Flash Vision120
9Inkling Small115
10GPT-5.6 Terra113

Blended cost per 1M tokens (3:1 input/output) — lower is cheaper. Minimum intelligence score of 40.

#Model$/1M
1DeepSeek V4 Flash$0.168
2GLM-5.3-Flash$0.237
3Hy3$0.241
4GPT-5.6 Luna$0.45
5MiniMax-M3$0.525
6Solar Pro 4$0.525
7Inkling Small$0.525
8DeepSeek V4 Pro$0.544
9MiMo-V2.5-Pro$0.544
10DeepSeek V4 Flash 0731$0.66