The Inference Report

August 2, 2026

AnthropicFable 5 maintains its lead on SWE-rebench at 64.5 percent with a confidence interval of plus or minus 1.41 percent, unchanged from the previous measurement, while GrokGrok 4.5 holds second place at 63.8 percent and AnthropicOpus 5 remains third at 63.4 percent, all three separated by margins within their measurement uncertainty. The top tier shows stability: the spread between first and fifth place spans just 2.2 percentage points, with Z.aiGLM-5.2 at 62.9 percent and OpenAIGPT-5.6 Sol at 62.3 percent, both marked as high-capability models. Below this tier, the rankings diverge sharply between the two evaluations, particularly for agent-based systems like JunieJunieAgent (61.8 percent on SWE-rebench, absent from Artificial Analysis) and AnthropicClaude CodeAgent (60.4 percent on SWE-rebench, unranked in Artificial Analysis), suggesting the two benchmarks measure distinct aspects of coding capability or that agent implementations perform better on repository-level tasks than on the broader evaluation set. The Artificial Analysis ranking places Claude Opus 5 first at 60.7 points and Claude Fable 5 second at 59.9 points, reversing their order on SWE-rebench and indicating either different test distributions or that general instruction-following performance does not correlate cleanly with software engineering problem-solving on controlled tasks. The SWE-rebench methodology appears more sensitive to agentic scaffolding and iteration, given agent systems' elevated positions there, while Artificial Analysis weights raw model capability more heavily, as evidenced by reasoning models like o1 and o3 appearing far lower (ranks 121 and 93 respectively) despite their problem-solving reputation. Neither benchmark shows the dramatic shifts that would indicate methodological drift, but the consistent reordering of Anthropic and OpenAI variants across the two systems warrants attention to how each constructs its evaluation tasks, particularly whether SWE-rebench's focus on repository context and tool use versus Artificial Analysis's emphasis on single-turn reasoning creates genuinely different capability profiles or simply reflects different measurement error.

Cole Brennan

Daily rankings from SWE-rebench, a benchmark designed to fairly compare LLM capabilities on real-world software engineering tasks. Unlike other evaluations, it uses a standardized scaffolding for all models, continuously updates its dataset to prevent contamination, and runs each model five times to account for stochastic variance.

#ModelScore
1AnthropicFable 5 [high]Model64.5%± 1.41%
2GrokGrok 4.5 [high]Model63.8%± 0.60%
3AnthropicOpus 5 [high]Model63.4%± 1.35%
4Z.aiGLM-5.2 [high]Model62.9%± 1.19%
5OpenAIGPT-5.6 Sol [medium]Model62.3%± 1.83%
6JunieJunieAgent61.8%± 0.54%
7AnthropicClaude CodeAgent60.4%± 1.03%
8OpenAICodexAgent58.0%± 1.29%
9AnthropicSonnet 5 [high]Model56.8%± 0.94%
10CursorCursorAgent51.7%± 0.84%

Artificial Analysis composite index across coding, math, and reasoning benchmarks.

#ModelScoretok/s$/1M
1Claude Opus 560.754$10.00
2Claude Fable 559.961$20.00
3GPT-5.6 Sol58.973$11.25
4Kimi K357.133$6.00
5Claude Opus 4.855.70$10.00
6GPT-5.6 Terra55123$4.50
7GPT-5.554.80$11.25
8Grok 4.553.853$3.00
9Claude Opus 4.753.50$10.00
10Claude Sonnet 553.475$4.00

Output tokens per second — higher is faster. Minimum intelligence score of 40.

#Modeltok/s
1Gemini 3.5 Flash221
2Qwen3.7 Max198
3Gemini 3.6 Flash195
4GPT-5.6 Luna162
5GLM-5.2142
6Muse Spark 1.1133
7Nex-N2-Pro131
8GPT-5.6 Terra123
9Gemini 3.1 Pro Preview120
10GPT-5.3 Codex120

Blended cost per 1M tokens (3:1 input/output) — lower is cheaper. Minimum intelligence score of 40.

#Model$/1M
1DeepSeek V4 Flash 0731$0.175
2DeepSeek V4 Flash$0.175
3Hy3$0.241
4GPT-5.6 Luna$0.45
5MiniMax-M3$0.525
6Inkling Small$0.525
7DeepSeek V4 Pro$0.544
8MiMo-V2.5-Pro$0.544
9Nex-N2-Pro$1.00
10GPT-5.4 mini$1.69