The Inference Report

August 14, 2026

On SWE-rebench, the coding agent benchmark, the top rankings show no movement: AnthropicFable 5 holds first place at 64.5% ± 1.41%, followed by GrokGrok 4.5 at 63.8% ± 0.60% and AnthropicOpus 5 at 63.4% ± 1.35%, with confidence intervals that do not overlap meaningfully across the top five positions. The Artificial Analysis benchmark tells a different story, where two new entries appear in the top tier: Qwen3.8 2.4T A95B enters at #7 with 57.7, pushing Claude Opus 4.8 down one position, and Gemini 3.7 Flash debuts at #12 with 56.0. The broader rankings remain largely stable across 432 positions, with most models retaining their previous slots or shifting by single positions. The consistency across SWE-rebench contrasts with modest churn in Artificial Analysis, where new models occasionally displace prior entrants but without dramatic reshuffling. Neither benchmark shows the kind of compression at the top that would indicate a breakthrough in reasoning or coding capability; instead, the data reflects incremental refinement and occasional new model releases entering established hierarchies. The stability of the SWE-rebench top ten, despite its larger confidence intervals, suggests the evaluation has reached a plateau where further gains require architectural innovation rather than parameter tuning.

Cole Brennan

Daily rankings from SWE-rebench, a benchmark designed to fairly compare LLM capabilities on real-world software engineering tasks. Unlike other evaluations, it uses a standardized scaffolding for all models, continuously updates its dataset to prevent contamination, and runs each model five times to account for stochastic variance.

#ModelScore
1AnthropicFable 5 [high]Model64.5%± 1.41%
2GrokGrok 4.5 [high]Model63.8%± 0.60%
3AnthropicOpus 5 [high]Model63.4%± 1.35%
4Z.aiGLM-5.2 [high]Model62.9%± 1.19%
5OpenAIGPT-5.6 Sol [medium]Model62.3%± 1.83%
6JunieJunieAgent61.8%± 0.54%
7AnthropicClaude CodeAgent60.4%± 1.03%
8OpenAICodexAgent58.0%± 1.29%
9AnthropicSonnet 5 [high]Model56.8%± 0.94%
10CursorCursorAgent51.7%± 0.84%

Artificial Analysis composite index across coding, math, and reasoning benchmarks.

#ModelScoretok/s$/1M
1Claude Opus 563.150$10.00
2Claude Fable 562.160$20.00
3GPT-5.6 Sol60.961$11.25
4Grok 4.660.954$3.00
5Kimi K359.737$6.00
6Qwen3.8 Max58.144$3.00
7Qwen3.8 2.4T A95B57.750$3.00
8Claude Opus 4.857.30$10.00
9Muse Spark 1.256.80$2.00
10GPT-5.6 Terra56.6106$4.50

Output tokens per second — higher is faster. Minimum intelligence score of 40.

#Modeltok/s
1Gemini 3.7 Flash515
2Gemini 3.6 Flash207
3GPT-5.6 Luna156
4Nex-N2-Pro139
5Gemini 3.1 Pro Preview127
6Inkling Small125
7GLM-5.2110
8DeepSeek V4 Flash 0731109
9GPT-5.3 Codex109
10GPT-5.6 Terra106

Blended cost per 1M tokens (3:1 input/output) — lower is cheaper. Minimum intelligence score of 40.

#Model$/1M
1DeepSeek V4 Flash$0.168
2DeepSeek V4 Flash 0731$0.175
3Hy3$0.241
4GPT-5.6 Luna$0.45
5MiniMax-M3$0.525
6Solar Pro 4$0.525
7Inkling Small$0.525
8DeepSeek V4 Pro$0.544
9MiMo-V2.5-Pro$0.544
10Nex-N2-Pro$1.00