The Inference Report

July 23, 2026

The SWE-rebench rankings show no movement from the previous cycle, with OpenAI's gpt-5.5-2026-04-23-xhighModel holding the top position at 62.7% ± 0.91%, followed by JunieJunieAgent at 61.6% ± 0.64% and OpenAICodexAgent at 60.4% ± 1.37%. The confidence intervals across the top performers remain wide enough that several models within the top ten overlap in their true performance ranges, particularly Claude CodeAgent (59.6% ± 1.98%) and the medium variant of gpt-5.5 (58.9% ± 0.78%), suggesting that ranking precision at these levels may exceed what the benchmark's variance actually supports. The Artificial Analysis benchmark, by contrast, shows substantial churn in the lower ranks: Nemotron Cascade 2 30B A3B dropped from position 131 to 164, while models in the 130-165 range shuffled significantly, yet the top tier remains stable with Claude Fable 5 at 59.9 and GPT-5.6 Sol at 58.9. The absence of movement in SWE-rebench's top positions across evaluation cycles, combined with overlapping error bars, suggests either genuine plateau in agent performance on this benchmark or that the test set may be approaching saturation for the leading approaches. The Artificial Analysis churn in lower positions reflects typical ranking volatility among closely-scored models rather than meaningful capability shifts.

Cole Brennan

Daily rankings from SWE-rebench, a benchmark designed to fairly compare LLM capabilities on real-world software engineering tasks. Unlike other evaluations, it uses a standardized scaffolding for all models, continuously updates its dataset to prevent contamination, and runs each model five times to account for stochastic variance.

#ModelScore
1OpenAIgpt-5.5-2026-04-23-xhighModel62.7%± 0.91%
2JunieJunieAgent61.6%± 0.64%
3OpenAICodexAgent60.4%± 1.37%
4AnthropicClaude CodeAgent59.6%± 1.98%
5OpenAIgpt-5.5-2026-04-23-mediumModel58.9%± 0.78%
6AnthropicClaude Opus 4.8-xhighModel56.5%± 1.20%
7OpenAIgpt-5.4-2026-03-05-mediumModel54.9%± 1.02%
8AnthropicClaude Opus 4.7-highModel53.1%± 1.45%
9CursorCursorAgent53.0%± 0.53%
10AnthropicClaude Sonnet 4.6Model51.3%± 0.55%

Artificial Analysis composite index across coding, math, and reasoning benchmarks.

#ModelScoretok/s$/1M
1Claude Fable 559.969$20.00
2GPT-5.6 Sol58.964$11.25
3Kimi K357.137$6.00
4Claude Opus 4.855.764$10.00
5GPT-5.6 Terra55134$5.63
6GPT-5.554.893$11.25
7Grok 4.553.864$3.00
8Claude Opus 4.753.561$10.00
9Claude Sonnet 553.483$4.00
10GPT-5.451.4149$5.63

Output tokens per second — higher is faster. Minimum intelligence score of 40.

#Modeltok/s
1Gemini 3.5 Flash276
2Gemini 3.6 Flash256
3Qwen3.7 Max207
4GPT-5.6 Luna200
5GPT-5.4 mini181
6GLM-5.2167
7GPT-5.2 Codex163
8GPT-5.4149
9GPT-5.6 Terra134
10Gemini 3.1 Pro Preview131

Blended cost per 1M tokens (3:1 input/output) — lower is cheaper. Minimum intelligence score of 40.

#Model$/1M
1DeepSeek V4 Flash$0.175
2Hy3$0.25
3MiniMax-M3$0.525
4DeepSeek V4 Pro$0.544
5MiMo-V2.5-Pro$0.544
6Nex-N2-Pro$1.00
7GPT-5.4 mini$1.69
8Kimi K2.6$1.71
9Kimi K2.7 Code$1.71
10Muse Spark 1.1$2.00