The Inference Report

August 30, 2026

Daily rankings from SWE-rebench, a benchmark designed to fairly compare LLM capabilities on real-world software engineering tasks. Unlike other evaluations, it uses a standardized scaffolding for all models, continuously updates its dataset to prevent contamination, and runs each model five times to account for stochastic variance.

#ModelScore
1AnthropicFable 5 [high]Model64.5%± 1.41%
2GrokGrok 4.5 [high]Model63.8%± 0.60%
3AnthropicOpus 5 [high]Model63.4%± 1.35%
4Z.aiGLM-5.2 [high]Model62.9%± 1.19%
5OpenAIGPT-5.6 Sol [medium]Model62.3%± 1.83%
6JunieJunieAgent61.8%± 0.54%
7AnthropicClaude CodeAgent60.4%± 1.03%
8OpenAICodexAgent58.0%± 1.29%
9AnthropicSonnet 5 [high]Model56.8%± 0.94%
10CursorCursorAgent51.7%± 0.84%

Artificial Analysis composite index across coding, math, and reasoning benchmarks.

#ModelScoretok/s$/1M
1Claude Opus 563.154$10.00
2Claude Fable 562.162$20.00
3GPT-5.6 Sol60.984$8.00
4Grok 4.660.958$3.00
5Kimi K359.736$6.00
6GLM-5.359.575$2.15
7Qwen3.8 Max58.130$3.00
8Qwen3.8 2.4T A95B57.731$3.00
9GLM-5.3-Flash57.550$0.237
10Claude Opus 4.857.30$10.00

Output tokens per second — higher is faster. Minimum intelligence score of 40.

#Modeltok/s
1Gemini 3.7 Flash345
2Gemini 3.6 Flash175
3Agnes 2.5 Pro Beta151
4DeepSeek V4 Flash 0731135
5Nex-N2-Pro134
6GPT-5.6 Luna129
7GPT-5.3 Codex126
8MiniMax-M3126
9Gemini 3.1 Pro Preview120
10DeepSeek V4 Flash Vision115

Blended cost per 1M tokens (3:1 input/output) — lower is cheaper. Minimum intelligence score of 40.

#Model$/1M
1Agnes 2.5 Pro Beta$0.15
2DeepSeek V4 Flash$0.168
3Qwen3.8-Flash-Next$0.23
4GLM-5.3-Flash$0.237
5Hy3$0.241
6GPT-5.6 Luna$0.45
7MiniMax-M3$0.525
8Solar Pro 4$0.525
9Inkling Small$0.525
10DeepSeek V4 Pro$0.544