The Inference Report

July 22, 2026

On SWE-rebench, the top tier remains stable: OpenAI's gpt-5.5-2026-04-23-xhighModel holds 62.7% (±0.91%), followed by JunieJunieAgent at 61.6% (±0.64%) and OpenAICodexAgent at 60.4% (±1.37%), with confidence intervals that do not overlap meaningfully enough to suggest ranking instability. The spread between positions 1 and 24 spans from 62.7% to 16.5%, a 46-point gap that reflects genuine capability stratification rather than noise. What merits scrutiny is the benchmark's design: SWE-rebench evaluates agent-based code completion on real software engineering tasks, a more applied setting than many token-prediction benchmarks, yet the evaluation methodology, particularly how task success is defined and whether partial credit is awarded, remains underspecified in the data provided. On Artificial Analysis, the roster has undergone substantial motion: Gemini 3.6 Flash enters at rank 15 (50.1 points), displacing prior entries downward, while Trinity Large Thinking drops from rank 112 (24.5 points) to rank 157 (18.2 points), a 6.3-point decline suggesting either a recalibration of the benchmark or a shift in evaluation conditions. The Artificial Analysis leaderboard exhibits denser clustering in the 30-50 point range than SWE-rebench, with many models separated by tenths of a point; this density raises questions about whether differences of 0.3-0.5 points reflect reproducible capability gaps or measurement variance. Neither benchmark's methodology clarifies whether scores are averaged across multiple runs, how task selection bias is controlled, or whether confidence intervals reflect statistical significance or merely reported uncertainty bounds. The two benchmarks show weak rank correlation at the extremes, Claude Fable 5 ranks first on Artificial Analysis (59.9) but does not appear on the SWE-rebench top 24, while gpt-5.5-xhigh dominates SWE-rebench but ranks sixth on Artificial Analysis (54.8), suggesting they measure distinct problem classes or that one benchmark's evaluation is more sensitive to architectural choices that do not generalize.

Cole Brennan

Daily rankings from SWE-rebench, a benchmark designed to fairly compare LLM capabilities on real-world software engineering tasks. Unlike other evaluations, it uses a standardized scaffolding for all models, continuously updates its dataset to prevent contamination, and runs each model five times to account for stochastic variance.

#ModelScore
1OpenAIgpt-5.5-2026-04-23-xhighModel62.7%± 0.91%
2JunieJunieAgent61.6%± 0.64%
3OpenAICodexAgent60.4%± 1.37%
4AnthropicClaude CodeAgent59.6%± 1.98%
5OpenAIgpt-5.5-2026-04-23-mediumModel58.9%± 0.78%
6AnthropicClaude Opus 4.8-xhighModel56.5%± 1.20%
7OpenAIgpt-5.4-2026-03-05-mediumModel54.9%± 1.02%
8AnthropicClaude Opus 4.7-highModel53.1%± 1.45%
9CursorCursorAgent53.0%± 0.53%
10AnthropicClaude Sonnet 4.6Model51.3%± 0.55%

Artificial Analysis composite index across coding, math, and reasoning benchmarks.

#ModelScoretok/s$/1M
1Claude Fable 559.970$20.00
2GPT-5.6 Sol58.967$11.25
3Kimi K357.138$6.00
4Claude Opus 4.855.764$10.00
5GPT-5.6 Terra55149$5.63
6GPT-5.554.892$11.25
7Grok 4.553.873$3.00
8Claude Opus 4.753.561$10.00
9Claude Sonnet 553.486$4.00
10GPT-5.451.4149$5.63

Output tokens per second — higher is faster. Minimum intelligence score of 40.

#Modeltok/s
1Gemini 3.6 Flash311
2Gemini 3.5 Flash283
3Qwen3.7 Max207
4GPT-5.6 Luna203
5GLM-5.2191
6GPT-5.4 mini179
7GPT-5.2 Codex162
8GPT-5.6 Terra149
9GPT-5.4149
10Gemini 3.1 Pro Preview135

Blended cost per 1M tokens (3:1 input/output) — lower is cheaper. Minimum intelligence score of 40.

#Model$/1M
1DeepSeek V4 Flash$0.175
2Hy3$0.25
3MiniMax-M3$0.525
4DeepSeek V4 Pro$0.544
5MiMo-V2.5-Pro$0.544
6Nex-N2-Pro$1.00
7GPT-5.4 mini$1.69
8Kimi K2.6$1.71
9Kimi K2.7 Code$1.71
10Muse Spark 1.1$2.00