The Inference Report

July 28, 2026

The two benchmark systems show stability at the top tier but divergence in methodology that warrants scrutiny. On SWE-rebench, the leader remains OpenAI gpt-5.5-2026-04-23-xhighModel at 62.7% ± 0.91%, with positions two through ten unchanged from the prior snapshot: JunieJunieAgent (61.6%), OpenAICodexAgent (60.4%), AnthropicClaude CodeAgent (59.6%), and five others holding their ranks. The confidence intervals are tight enough that these orderings reflect real performance separation, though the 1.1-point gap between first and second suggests the frontier remains contested. Artificial Analysis data tells a different story about which models matter. Claude Opus 5 leads there at 60.7, ahead of Claude Fable 5 (59.9) and GPT-5.6 Sol (58.9), a ranking that clusters Anthropic and OpenAI variants rather than privileging agent-based configurations. The two datasets agree on little: SWE-rebench elevates specialized agent frameworks (JunieAgent, Cursor) to positions two and nine, while Artificial Analysis buries agent variants in the mid-tier or lower. This gap reflects fundamental differences in evaluation scope. SWE-rebench appears narrowly calibrated to software engineering tasks with reproducible metrics and error bars; Artificial Analysis spans broader capability assessment and lacks reported confidence bounds, making score movements harder to interpret. Neither system has shifted meaningfully since the previous snapshot, suggesting either genuine stability or evaluation lag. The practical signal is clearest on SWE-rebench's top twenty, where the spread from 62.7% to 38.4% separates systems with meaningfully different engineering performance. Below that, the Artificial Analysis tail extends to models scoring 1.0, a floor that lacks discriminative power and suggests saturation in the evaluation methodology rather than genuine parity.

Cole Brennan

Daily rankings from SWE-rebench, a benchmark designed to fairly compare LLM capabilities on real-world software engineering tasks. Unlike other evaluations, it uses a standardized scaffolding for all models, continuously updates its dataset to prevent contamination, and runs each model five times to account for stochastic variance.

#ModelScore
1OpenAIgpt-5.5-2026-04-23-xhighModel62.7%± 0.91%
2JunieJunieAgent61.6%± 0.64%
3OpenAICodexAgent60.4%± 1.37%
4AnthropicClaude CodeAgent59.6%± 1.98%
5OpenAIgpt-5.5-2026-04-23-mediumModel58.9%± 0.78%
6AnthropicClaude Opus 4.8-xhighModel56.5%± 1.20%
7OpenAIgpt-5.4-2026-03-05-mediumModel54.9%± 1.02%
8AnthropicClaude Opus 4.7-highModel53.1%± 1.45%
9CursorCursorAgent53.0%± 0.53%
10AnthropicClaude Sonnet 4.6Model51.3%± 0.55%

Artificial Analysis composite index across coding, math, and reasoning benchmarks.

#ModelScoretok/s$/1M
1Claude Opus 560.763$10.00
2Claude Fable 559.971$20.00
3GPT-5.6 Sol58.990$11.25
4Kimi K357.133$6.00
5Claude Opus 4.855.765$10.00
6GPT-5.6 Terra55165$5.63
7GPT-5.554.80$11.25
8Grok 4.553.861$3.00
9Claude Opus 4.753.50$10.00
10Claude Sonnet 553.486$4.00

Output tokens per second — higher is faster. Minimum intelligence score of 40.

#Modeltok/s
1Gemini 3.5 Flash266
2Gemini 3.6 Flash255
3GPT-5.6 Luna220
4GLM-5.2219
5Qwen3.7 Max203
6GPT-5.6 Terra165
7GPT-5.3 Codex148
8Gemini 3.1 Pro Preview147
9Nex-N2-Pro142
10Muse Spark 1.1137

Blended cost per 1M tokens (3:1 input/output) — lower is cheaper. Minimum intelligence score of 40.

#Model$/1M
1DeepSeek V4 Flash$0.175
2Hy3$0.25
3MiniMax-M3$0.525
4DeepSeek V4 Pro$0.544
5MiMo-V2.5-Pro$0.544
6Nex-N2-Pro$1.00
7GPT-5.4 mini$1.69
8Kimi K2.6$1.71
9Kimi K2.7 Code$1.71
10Muse Spark 1.1$2.00