The Inference Report

August 18, 2026

The SWE-rebench rankings show no movement from the previous cycle, with AnthropicFable 5 holding first place at 64.5 percent and the entire top 17 models maintaining their positions unchanged. Across the Artificial Analysis benchmark, the top tier remains similarly static: Claude Opus 5 leads at 63.1, followed by Claude Fable 5 at 62.1, though these scores differ slightly from the SWE-rebench ordering, suggesting the two benchmarks measure overlapping but distinct capabilities. Within the broader Artificial Analysis leaderboard, a single new entry appears at rank 21, Qwen3.8 27B scoring 52.0, displacing previous models down by one position each through the middle ranks, while the lower tiers below position 300 show only minor reordering with no meaningful score changes. The consistency across both benchmarks indicates either genuine stability in model performance or measurement limitations that prevent detection of incremental gains; the SWE-rebench confidence intervals, ranging from 0.54 to 1.83 percentage points, are wide enough that most apparent differences between adjacent models fall within noise, and without access to the evaluation methodology details, it remains unclear whether the benchmark is sufficiently sensitive to capture real progress or whether the field has genuinely plateaued at this performance level.

Cole Brennan

Daily rankings from SWE-rebench, a benchmark designed to fairly compare LLM capabilities on real-world software engineering tasks. Unlike other evaluations, it uses a standardized scaffolding for all models, continuously updates its dataset to prevent contamination, and runs each model five times to account for stochastic variance.

#ModelScore
1AnthropicFable 5 [high]Model64.5%± 1.41%
2GrokGrok 4.5 [high]Model63.8%± 0.60%
3AnthropicOpus 5 [high]Model63.4%± 1.35%
4Z.aiGLM-5.2 [high]Model62.9%± 1.19%
5OpenAIGPT-5.6 Sol [medium]Model62.3%± 1.83%
6JunieJunieAgent61.8%± 0.54%
7AnthropicClaude CodeAgent60.4%± 1.03%
8OpenAICodexAgent58.0%± 1.29%
9AnthropicSonnet 5 [high]Model56.8%± 0.94%
10CursorCursorAgent51.7%± 0.84%

Artificial Analysis composite index across coding, math, and reasoning benchmarks.

#ModelScoretok/s$/1M
1Claude Opus 563.152$10.00
2Claude Fable 562.167$20.00
3GPT-5.6 Sol60.970$11.25
4Grok 4.660.963$3.00
5Kimi K359.740$6.00
6Qwen3.8 Max58.147$3.00
7Qwen3.8 2.4T A95B57.747$3.00
8Claude Opus 4.857.30$10.00
9Muse Spark 1.256.80$2.00
10GPT-5.6 Terra56.6118$4.50

Output tokens per second — higher is faster. Minimum intelligence score of 40.

#Modeltok/s
1Gemini 3.7 Flash297
2Gemini 3.6 Flash209
3GPT-5.6 Luna169
4GLM-5.2149
5Nex-N2-Pro143
6Gemini 3.1 Pro Preview136
7GPT-5.3 Codex132
8DeepSeek V4 Flash 0731123
9GPT-5.6 Terra118
10MiniMax-M395

Blended cost per 1M tokens (3:1 input/output) — lower is cheaper. Minimum intelligence score of 40.

#Model$/1M
1DeepSeek V4 Flash$0.168
2Hy3$0.241
3GPT-5.6 Luna$0.45
4MiniMax-M3$0.525
5Solar Pro 4$0.525
6Inkling Small$0.525
7DeepSeek V4 Pro$0.544
8MiMo-V2.5-Pro$0.544
9DeepSeek V4 Flash 0731$0.66
10Nex-N2-Pro$1.00