The Inference Report

July 21, 2026

The SWE-rebench leaderboard shows no movement from the previous snapshot, with OpenAI's gpt-5.5-2026-04-23-xhighModel holding 62.7% (±0.91%) at the top, followed by JunieJunieAgent at 61.6% (±0.64%) and OpenAICodexAgent at 60.4% (±1.37%). The Artificial Analysis benchmark, however, exhibits substantial churn across its 409 entries, with Motif 3 entering at rank 22 (44.1 points) and Mercury 2 dropping from rank 109 to 128 (25.3 to 21.4 points), suggesting either model updates or changes in evaluation methodology that warrant clarification. The lack of overlap between top performers on the two benchmarks, Claude Fable 5 leads Artificial Analysis at 59.9 while ranking nowhere on SWE-rebench's top 24, indicates these measure distinct problem spaces rather than equivalent capabilities. SWE-rebench's narrow confidence intervals (0.45% to 1.98%) suggest tighter experimental control than typical leaderboards, though the static rankings across both snapshots raise questions about evaluation frequency and whether these represent live benchmarks or periodic releases. Without details on SWE-rebench's test set composition, task distribution, or whether agents can call external tools, it remains unclear whether the 1.1-point gap between first and second place reflects genuine capability differences or measurement noise within the confidence bounds.

Cole Brennan

Daily rankings from SWE-rebench, a benchmark designed to fairly compare LLM capabilities on real-world software engineering tasks. Unlike other evaluations, it uses a standardized scaffolding for all models, continuously updates its dataset to prevent contamination, and runs each model five times to account for stochastic variance.

#ModelScore
1OpenAIgpt-5.5-2026-04-23-xhighModel62.7%± 0.91%
2JunieJunieAgent61.6%± 0.64%
3OpenAICodexAgent60.4%± 1.37%
4AnthropicClaude CodeAgent59.6%± 1.98%
5OpenAIgpt-5.5-2026-04-23-mediumModel58.9%± 0.78%
6AnthropicClaude Opus 4.8-xhighModel56.5%± 1.20%
7OpenAIgpt-5.4-2026-03-05-mediumModel54.9%± 1.02%
8AnthropicClaude Opus 4.7-highModel53.1%± 1.45%
9CursorCursorAgent53.0%± 0.53%
10AnthropicClaude Sonnet 4.6Model51.3%± 0.55%

Artificial Analysis composite index across coding, math, and reasoning benchmarks.

#ModelScoretok/s$/1M
1Claude Fable 559.970$20.00
2GPT-5.6 Sol58.968$11.25
3Kimi K357.139$6.00
4Claude Opus 4.855.762$10.00
5GPT-5.6 Terra55154$5.63
6GPT-5.554.886$11.25
7Grok 4.553.875$3.00
8Claude Opus 4.753.560$10.00
9Claude Sonnet 553.491$4.00
10GPT-5.451.4163$5.63

Output tokens per second — higher is faster. Minimum intelligence score of 40.

#Modeltok/s
1Gemini 3.5 Flash294
2GPT-5.6 Luna212
3Qwen3.7 Max208
4GLM-5.2198
5GPT-5.4 mini180
6GPT-5.4163
7GPT-5.2 Codex163
8GPT-5.6 Terra154
9Nex-N2-Pro138
10Gemini 3.1 Pro Preview137

Blended cost per 1M tokens (3:1 input/output) — lower is cheaper. Minimum intelligence score of 40.

#Model$/1M
1DeepSeek V4 Flash$0.175
2MiniMax-M3$0.525
3DeepSeek V4 Pro$0.544
4MiMo-V2.5-Pro$0.544
5Nex-N2-Pro$1.00
6GPT-5.4 mini$1.69
7Kimi K2.6$1.71
8Kimi K2.7 Code$1.71
9Muse Spark 1.1$2.00
10GLM-5.2$2.15