The Inference Report

August 26, 2026

The SWE-rebench rankings remain frozen at their previous positions, with AnthropicFable 5 holding first place at 64.5% ± 1.41%, followed by GrokGrok 4.5 at 63.8% ± 0.60%, and AnthropicOpus 5 at 63.4% ± 1.35%, suggesting either a measurement plateau or a pause in the release cadence for coding agents. On the Artificial Analysis benchmark, the roster has shifted substantially: Claude Opus 5 now leads at 63.1 (up from second), while Claude Fable 5 dropped to 62.1 (down from first), and a new entry, DeepSeek V4 Flash Vision at 51.5, has inserted itself at position 26, pushing all subsequent models down one slot. The divergence between these two benchmarks is notable, SWE-rebench evaluates coding task completion in a controlled sandbox environment with reproducible test suites, while Artificial Analysis appears to measure general capability across a broader evaluation matrix, which explains why the same models rank differently across the two tests. The stability in SWE-rebench contrasts with the churn in Artificial Analysis, where the top tier remains competitive but entries below position 25 experience consistent downward pressure as newer variants appear. Neither benchmark methodology is disclosed in detail here, but the consistency of SWE-rebench's top performers and their tight clustering (all within 8.3 percentage points) suggests the task distribution may be saturating for frontier models, while Artificial Analysis's wider spread indicates greater differentiation across capability profiles.

Cole Brennan

Daily rankings from SWE-rebench, a benchmark designed to fairly compare LLM capabilities on real-world software engineering tasks. Unlike other evaluations, it uses a standardized scaffolding for all models, continuously updates its dataset to prevent contamination, and runs each model five times to account for stochastic variance.

#ModelScore
1AnthropicFable 5 [high]Model64.5%± 1.41%
2GrokGrok 4.5 [high]Model63.8%± 0.60%
3AnthropicOpus 5 [high]Model63.4%± 1.35%
4Z.aiGLM-5.2 [high]Model62.9%± 1.19%
5OpenAIGPT-5.6 Sol [medium]Model62.3%± 1.83%
6JunieJunieAgent61.8%± 0.54%
7AnthropicClaude CodeAgent60.4%± 1.03%
8OpenAICodexAgent58.0%± 1.29%
9AnthropicSonnet 5 [high]Model56.8%± 0.94%
10CursorCursorAgent51.7%± 0.84%

Artificial Analysis composite index across coding, math, and reasoning benchmarks.

#ModelScoretok/s$/1M
1Claude Opus 563.155$10.00
2Claude Fable 562.164$20.00
3GPT-5.6 Sol60.970$8.00
4Grok 4.660.957$3.00
5Kimi K359.738$6.00
6GLM-5.359.580$2.15
7Qwen3.8 Max58.124$3.00
8Qwen3.8 2.4T A95B57.724$3.00
9Claude Opus 4.857.30$10.00
10Muse Spark 1.256.80$2.00

Output tokens per second — higher is faster. Minimum intelligence score of 40.

#Modeltok/s
1Gemini 3.7 Flash345
2Gemini 3.6 Flash195
3MiniMax-M3139
4Nex-N2-Pro133
5Inkling Small131
6GPT-5.6 Luna130
7Gemini 3.1 Pro Preview123
8DeepSeek V4 Flash 0731122
9GPT-5.3 Codex121
10DeepSeek V4 Flash Vision120

Blended cost per 1M tokens (3:1 input/output) — lower is cheaper. Minimum intelligence score of 40.

#Model$/1M
1DeepSeek V4 Flash$0.168
2Hy3$0.241
3GPT-5.6 Luna$0.45
4MiniMax-M3$0.525
5Solar Pro 4$0.525
6Inkling Small$0.525
7DeepSeek V4 Pro$0.544
8MiMo-V2.5-Pro$0.544
9DeepSeek V4 Flash 0731$0.66
10DeepSeek V4 Flash Vision$0.66