The Inference Report

July 20, 2026

The SWE-rebench rankings remain unchanged from the previous snapshot, with OpenAI's gpt-5.5-2026-04-23-xhighModel holding first place at 62.7% ± 0.91%, followed by JunieJunieAgent at 61.6% ± 0.64% and OpenAICodexAgent at 60.4% ± 1.37%. The confidence intervals overlap substantially across the top tier, suggesting the performance gap between positions one through five is within measurement noise. Artificial Analysis data shows broader dispersion across 408 models, with Claude Fable 5 leading at 59.9 and models dropping below 2% in the long tail, but this benchmark uses different evaluation methodology and does not track software engineering task completion in the same way as SWE-rebench. The absence of movement in the SWE-rebench leaderboard across two reporting cycles could indicate either genuine stability in the coding agent landscape or insufficient statistical power to detect meaningful shifts given the modest confidence intervals. Without prior historical snapshots beyond these two identical readings, it remains unclear whether the SWE-rebench results represent a genuine plateau in agent performance or simply reflect the inherent variance of the evaluation itself.

Cole Brennan

Daily rankings from SWE-rebench, a benchmark designed to fairly compare LLM capabilities on real-world software engineering tasks. Unlike other evaluations, it uses a standardized scaffolding for all models, continuously updates its dataset to prevent contamination, and runs each model five times to account for stochastic variance.

#ModelScore
1OpenAIgpt-5.5-2026-04-23-xhighModel62.7%± 0.91%
2JunieJunieAgent61.6%± 0.64%
3OpenAICodexAgent60.4%± 1.37%
4AnthropicClaude CodeAgent59.6%± 1.98%
5OpenAIgpt-5.5-2026-04-23-mediumModel58.9%± 0.78%
6AnthropicClaude Opus 4.8-xhighModel56.5%± 1.20%
7OpenAIgpt-5.4-2026-03-05-mediumModel54.9%± 1.02%
8AnthropicClaude Opus 4.7-highModel53.1%± 1.45%
9CursorCursorAgent53.0%± 0.53%
10AnthropicClaude Sonnet 4.6Model51.3%± 0.55%

Artificial Analysis composite index across coding, math, and reasoning benchmarks.

#ModelScoretok/s$/1M
1Claude Fable 559.966$20.00
2GPT-5.6 Sol58.971$11.25
3Kimi K357.10$6.00
4Claude Opus 4.855.761$10.00
5GPT-5.6 Terra55140$5.63
6GPT-5.554.886$11.25
7Grok 4.553.873$3.00
8Claude Opus 4.753.553$10.00
9Claude Sonnet 553.481$4.00
10GPT-5.451.4165$5.63

Output tokens per second — higher is faster. Minimum intelligence score of 40.

#Modeltok/s
1Gemini 3.5 Flash291
2GPT-5.6 Luna209
3Qwen3.7 Max205
4GLM-5.2185
5GPT-5.4 mini178
6GPT-5.4165
7GPT-5.2 Codex154
8Nex-N2-Pro142
9GPT-5.6 Terra140
10Gemini 3.1 Pro Preview137

Blended cost per 1M tokens (3:1 input/output) — lower is cheaper. Minimum intelligence score of 40.

#Model$/1M
1DeepSeek V4 Flash$0.175
2MiniMax-M3$0.525
3DeepSeek V4 Pro$0.544
4MiMo-V2.5-Pro$0.544
5Nex-N2-Pro$1.00
6GPT-5.4 mini$1.69
7Kimi K2.6$1.71
8Kimi K2.7 Code$1.71
9Muse Spark 1.1$2.00
10GLM-5.2$2.15