The Inference Report

August 25, 2026

The SWE-rebench top tier shows no movement from the previous cycle, with AnthropicFable 5 holding at 64.5 percent, GrokGrok 4.5 at 63.8 percent, and AnthropicOpus 5 at 63.4 percent, all within their confidence intervals. The gap between first and fifth place spans just 2.2 percentage points, indicating a plateau in frontier performance where the leading models have converged. Notably, the confidence intervals remain wide for high-confidence models like AnthropicFable 5 (±1.41%) and OpenAIGPT-5.6 Sol (±1.83%), suggesting that SWE-rebench, despite its focus on concrete software engineering tasks, still exhibits considerable variance in its measurements. The Artificial Analysis benchmark tells a different story: Claude Opus 5 has moved to first at 63.1, displacing Claude Fable 5 to second at 62.1, a swap that contradicts the SWE-rebench ordering where Fable 5 ranks higher. This discordance between benchmarks reflects fundamental differences in what each measures. SWE-rebench targets repository-level bug fixes and feature implementations with deterministic evaluation, whereas Artificial Analysis appears to capture broader reasoning or integration capabilities where Opus 5 edges ahead. The lower-tier Artificial Analysis rankings show substantial churn, with models like QwQ 32B and Qwen3 VL 30B A3B swapping positions at ranks 224-225, and several entries like Granite 4.0 variants clustered at 1.0, suggesting floor effects or evaluation instability at scale. Neither benchmark demonstrates dramatic movement that would indicate a methodological breakthrough; instead, both suggest incremental consolidation among established leaders and noise among the field.

Cole Brennan

Daily rankings from SWE-rebench, a benchmark designed to fairly compare LLM capabilities on real-world software engineering tasks. Unlike other evaluations, it uses a standardized scaffolding for all models, continuously updates its dataset to prevent contamination, and runs each model five times to account for stochastic variance.

#ModelScore
1AnthropicFable 5 [high]Model64.5%± 1.41%
2GrokGrok 4.5 [high]Model63.8%± 0.60%
3AnthropicOpus 5 [high]Model63.4%± 1.35%
4Z.aiGLM-5.2 [high]Model62.9%± 1.19%
5OpenAIGPT-5.6 Sol [medium]Model62.3%± 1.83%
6JunieJunieAgent61.8%± 0.54%
7AnthropicClaude CodeAgent60.4%± 1.03%
8OpenAICodexAgent58.0%± 1.29%
9AnthropicSonnet 5 [high]Model56.8%± 0.94%
10CursorCursorAgent51.7%± 0.84%

Artificial Analysis composite index across coding, math, and reasoning benchmarks.

#ModelScoretok/s$/1M
1Claude Opus 563.154$10.00
2Claude Fable 562.166$20.00
3GPT-5.6 Sol60.962$8.00
4Grok 4.660.954$3.00
5Kimi K359.735$6.00
6GLM-5.359.50$2.15
7Qwen3.8 Max58.123$3.00
8Qwen3.8 2.4T A95B57.727$3.00
9Claude Opus 4.857.30$10.00
10Muse Spark 1.256.80$2.00

Output tokens per second — higher is faster. Minimum intelligence score of 40.

#Modeltok/s
1Gemini 3.7 Flash358
2Gemini 3.6 Flash196
3GPT-5.6 Luna139
4MiniMax-M3139
5Nex-N2-Pro139
6Inkling Small137
7GPT-5.3 Codex136
8Gemini 3.1 Pro Preview122
9GPT-5.6 Terra114
10DeepSeek V4 Flash 0731109

Blended cost per 1M tokens (3:1 input/output) — lower is cheaper. Minimum intelligence score of 40.

#Model$/1M
1DeepSeek V4 Flash$0.168
2Hy3$0.241
3GPT-5.6 Luna$0.45
4MiniMax-M3$0.525
5Solar Pro 4$0.525
6Inkling Small$0.525
7DeepSeek V4 Pro$0.544
8MiMo-V2.5-Pro$0.544
9DeepSeek V4 Flash 0731$0.66
10Nex-N2-Pro$1.00