The Inference Report

July 25, 2026

The SWE-rebench leaderboard shows no movement in the top positions, with OpenAI's gpt-5.5-2026-04-23-xhighModel holding rank #1 at 62.7% ± 0.91%, followed by JunieJunieAgent at 61.6% ± 0.64% and OpenAICodexAgent at 60.4% ± 1.37%, the same three-model ordering from the previous cycle. Artificial Analysis, by contrast, registered a single meaningful shift: Claude Opus 5 debuted at #1 with 60.7, displacing Claude Fable 5 from the top spot to #2 at 59.9, a 0.8-point gain that exceeds typical measurement noise given the benchmark's established variance patterns. The SWE-rebench confidence intervals (ranging from 0.45% to 1.98%) suggest sufficient precision to detect real movement, yet the absence of any ranking changes across 24 models indicates either genuine stability in relative performance or that improvements, if present, fall within the bounds of statistical uncertainty. Artificial Analysis's deeper leaderboard shows broader reshuffling throughout the 400-entry list, though the top tier remains concentrated among OpenAI, Anthropic, and emerging Chinese models like Kimi and GLM variants. Without prior Artificial Analysis data to establish trend patterns, it remains unclear whether the Claude Opus 5 promotion reflects sustained capability gains or transient evaluation variance; the SWE-rebench's frozen ranking, however, suggests that whatever changes occurred in the underlying systems did not cross the performance thresholds needed to alter the established hierarchy.

Cole Brennan

Daily rankings from SWE-rebench, a benchmark designed to fairly compare LLM capabilities on real-world software engineering tasks. Unlike other evaluations, it uses a standardized scaffolding for all models, continuously updates its dataset to prevent contamination, and runs each model five times to account for stochastic variance.

#ModelScore
1OpenAIgpt-5.5-2026-04-23-xhighModel62.7%± 0.91%
2JunieJunieAgent61.6%± 0.64%
3OpenAICodexAgent60.4%± 1.37%
4AnthropicClaude CodeAgent59.6%± 1.98%
5OpenAIgpt-5.5-2026-04-23-mediumModel58.9%± 0.78%
6AnthropicClaude Opus 4.8-xhighModel56.5%± 1.20%
7OpenAIgpt-5.4-2026-03-05-mediumModel54.9%± 1.02%
8AnthropicClaude Opus 4.7-highModel53.1%± 1.45%
9CursorCursorAgent53.0%± 0.53%
10AnthropicClaude Sonnet 4.6Model51.3%± 0.55%

Artificial Analysis composite index across coding, math, and reasoning benchmarks.

#ModelScoretok/s$/1M
1Claude Opus 560.744$10.00
2Claude Fable 559.958$20.00
3GPT-5.6 Sol58.974$11.25
4Kimi K357.133$6.00
5Claude Opus 4.855.763$10.00
6GPT-5.6 Terra55128$5.63
7GPT-5.554.80$11.25
8Grok 4.553.856$3.00
9Claude Opus 4.753.50$10.00
10Claude Sonnet 553.483$4.00

Output tokens per second — higher is faster. Minimum intelligence score of 40.

#Modeltok/s
1Gemini 3.5 Flash250
2Gemini 3.6 Flash219
3Qwen3.7 Max200
4GPT-5.6 Luna171
5GLM-5.2157
6Gemini 3.1 Pro Preview132
7Nex-N2-Pro129
8GPT-5.6 Terra128
9GPT-5.3 Codex126
10Muse Spark 1.1124

Blended cost per 1M tokens (3:1 input/output) — lower is cheaper. Minimum intelligence score of 40.

#Model$/1M
1DeepSeek V4 Flash$0.175
2Hy3$0.25
3MiniMax-M3$0.525
4DeepSeek V4 Pro$0.544
5MiMo-V2.5-Pro$0.544
6Nex-N2-Pro$1.00
7GPT-5.4 mini$1.69
8Kimi K2.6$1.71
9Kimi K2.7 Code$1.71
10Muse Spark 1.1$2.00