The Inference Report

July 19, 2026

Daily rankings from SWE-rebench, a benchmark designed to fairly compare LLM capabilities on real-world software engineering tasks. Unlike other evaluations, it uses a standardized scaffolding for all models, continuously updates its dataset to prevent contamination, and runs each model five times to account for stochastic variance.

#ModelScore
1OpenAIgpt-5.5-2026-04-23-xhighModel62.7%± 0.91%
2JunieJunieAgent61.6%± 0.64%
3OpenAICodexAgent60.4%± 1.37%
4AnthropicClaude CodeAgent59.6%± 1.98%
5OpenAIgpt-5.5-2026-04-23-mediumModel58.9%± 0.78%
6AnthropicClaude Opus 4.8-xhighModel56.5%± 1.20%
7OpenAIgpt-5.4-2026-03-05-mediumModel54.9%± 1.02%
8AnthropicClaude Opus 4.7-highModel53.1%± 1.45%
9CursorCursorAgent53.0%± 0.53%
10AnthropicClaude Sonnet 4.6Model51.3%± 0.55%

Artificial Analysis composite index across coding, math, and reasoning benchmarks.

#ModelScoretok/s$/1M
1Claude Fable 559.957$20.00
2GPT-5.6 Sol58.965$11.25
3Kimi K357.159$6.00
4Claude Opus 4.855.752$10.00
5GPT-5.6 Terra55137$5.63
6GPT-5.554.880$11.25
7Grok 4.553.874$3.00
8Claude Opus 4.753.550$10.00
9Claude Sonnet 553.478$4.00
10GPT-5.451.4159$5.63

Output tokens per second — higher is faster. Minimum intelligence score of 40.

#Modeltok/s
1Gemini 3.5 Flash282
2Qwen3.7 Max204
3GPT-5.6 Luna200
4GLM-5.2179
5GPT-5.4 mini177
6GPT-5.4159
7GPT-5.2 Codex149
8Nex-N2-Pro144
9GPT-5.6 Terra137
10Gemini 3.1 Pro Preview127

Blended cost per 1M tokens (3:1 input/output) — lower is cheaper. Minimum intelligence score of 40.

#Model$/1M
1DeepSeek V4 Flash$0.175
2MiniMax-M3$0.525
3DeepSeek V4 Pro$0.544
4MiMo-V2.5-Pro$0.544
5Nex-N2-Pro$1.00
6GPT-5.4 mini$1.69
7Kimi K2.6$1.71
8Kimi K2.7 Code$1.71
9Muse Spark 1.1$2.00
10GLM-5.2$2.15