The Inference Report

July 24, 2026

The SWE-rebench and Artificial Analysis rankings show stability at the top with minor consolidation below. On SWE-rebench, the top four remain unchanged: OpenAI gpt-5.5-2026-04-23-xhighModel at 62.7%, JunieJunieAgent at 61.6%, OpenAICodexAgent at 60.4%, and AnthropicClaude CodeAgent at 59.6%, with confidence intervals suggesting these gaps are meaningful given the ±0.64% to ±1.98% uncertainty bands. The Artificial Analysis benchmark shows more churn in the 40-180 range, where Agnes 2.5 Pro Alpha enters at position 45 with 38.8 points, pushing prior entries down one slot, and K-EXAONE drops from 113 to 125 (24.7 to 22.1), while G9v3-3B enters at 176 with 16.1 points. Neither benchmark exhibits dramatic shifts at the frontier; the SWE-rebench top performers remain separated by gaps of 0.7 to 1.1 percentage points that exceed most reported error margins, indicating the ranking reflects genuine capability differences rather than noise. Artificial Analysis shows more granular differentiation across its extended tail, but the lack of confidence intervals there limits interpretation of whether movements represent real changes or sampling variation. The stability of the top positions across both benchmarks suggests the leading agent architectures have reached a performance plateau on these tasks, at least relative to the measurement precision available.

Cole Brennan

Daily rankings from SWE-rebench, a benchmark designed to fairly compare LLM capabilities on real-world software engineering tasks. Unlike other evaluations, it uses a standardized scaffolding for all models, continuously updates its dataset to prevent contamination, and runs each model five times to account for stochastic variance.

#ModelScore
1OpenAIgpt-5.5-2026-04-23-xhighModel62.7%± 0.91%
2JunieJunieAgent61.6%± 0.64%
3OpenAICodexAgent60.4%± 1.37%
4AnthropicClaude CodeAgent59.6%± 1.98%
5OpenAIgpt-5.5-2026-04-23-mediumModel58.9%± 0.78%
6AnthropicClaude Opus 4.8-xhighModel56.5%± 1.20%
7OpenAIgpt-5.4-2026-03-05-mediumModel54.9%± 1.02%
8AnthropicClaude Opus 4.7-highModel53.1%± 1.45%
9CursorCursorAgent53.0%± 0.53%
10AnthropicClaude Sonnet 4.6Model51.3%± 0.55%

Artificial Analysis composite index across coding, math, and reasoning benchmarks.

#ModelScoretok/s$/1M
1Claude Fable 559.959$20.00
2GPT-5.6 Sol58.962$11.25
3Kimi K357.133$6.00
4Claude Opus 4.855.761$10.00
5GPT-5.6 Terra55122$5.63
6GPT-5.554.80$11.25
7Grok 4.553.856$3.00
8Claude Opus 4.753.50$10.00
9Claude Sonnet 553.481$4.00
10GPT-5.451.40$5.63

Output tokens per second — higher is faster. Minimum intelligence score of 40.

#Modeltok/s
1Gemini 3.6 Flash255
2Gemini 3.5 Flash252
3Qwen3.7 Max200
4GPT-5.6 Luna170
5GLM-5.2164
6Nex-N2-Pro130
7Gemini 3.1 Pro Preview128
8Muse Spark 1.1124
9GPT-5.6 Terra122
10DeepSeek V4 Flash120

Blended cost per 1M tokens (3:1 input/output) — lower is cheaper. Minimum intelligence score of 40.

#Model$/1M
1DeepSeek V4 Flash$0.175
2Hy3$0.25
3MiniMax-M3$0.525
4DeepSeek V4 Pro$0.544
5MiMo-V2.5-Pro$0.544
6Nex-N2-Pro$1.00
7GPT-5.4 mini$1.69
8Kimi K2.6$1.71
9Kimi K2.7 Code$1.71
10Muse Spark 1.1$2.00