The Inference Report

August 3, 2026

The SWE-rebench rankings hold steady at the top, with AnthropicFable 5 retaining first place at 64.5% and the next five positions unchanged through GPT-5.6 Sol. The confidence intervals remain wide enough that the top tier shows genuine separation: Fable 5's 1.41% margin and Grok 4.5's 0.60% margin indicate different levels of measurement precision, yet both models occupy their positions with confidence bounds that don't overlap meaningfully with adjacent ranks. Below the top six, however, the data becomes noisier. JunieAgent holds sixth at 61.8% with a tight 0.54% margin, but Claude CodeAgent and OpenAICodexAgent follow at 60.4% and 58.0%, suggesting a drop-off in performance that the error bars don't fully resolve. The Artificial Analysis rankings tell a different story. Claude Opus 5 leads there at 60.7, while AnthropicFable 5 ranks second at 59.9, a reversal of the SWE-rebench hierarchy that points to benchmark-specific strengths rather than universal capability differences. GPT-5.6 Sol appears at position three in Artificial Analysis (58.9) but holds fifth on SWE-rebench (62.3%), indicating the two evaluation frameworks weight code-generation and general reasoning tasks differently. The divergence widens further down: DeepSeek-V4 Pro scores 40.2% on SWE-rebench but only 44.3 on Artificial Analysis, placing it 23 positions lower in the latter ranking. These inversions suggest neither benchmark captures the full picture of model capability, and that agentic code performance (measured by SWE-rebench) does not correlate linearly with broader reasoning benchmarks.

Cole Brennan

Daily rankings from SWE-rebench, a benchmark designed to fairly compare LLM capabilities on real-world software engineering tasks. Unlike other evaluations, it uses a standardized scaffolding for all models, continuously updates its dataset to prevent contamination, and runs each model five times to account for stochastic variance.

#ModelScore
1AnthropicFable 5 [high]Model64.5%± 1.41%
2GrokGrok 4.5 [high]Model63.8%± 0.60%
3AnthropicOpus 5 [high]Model63.4%± 1.35%
4Z.aiGLM-5.2 [high]Model62.9%± 1.19%
5OpenAIGPT-5.6 Sol [medium]Model62.3%± 1.83%
6JunieJunieAgent61.8%± 0.54%
7AnthropicClaude CodeAgent60.4%± 1.03%
8OpenAICodexAgent58.0%± 1.29%
9AnthropicSonnet 5 [high]Model56.8%± 0.94%
10CursorCursorAgent51.7%± 0.84%

Artificial Analysis composite index across coding, math, and reasoning benchmarks.

#ModelScoretok/s$/1M
1Claude Opus 560.759$10.00
2Claude Fable 559.971$20.00
3GPT-5.6 Sol58.977$11.25
4Kimi K357.134$6.00
5Claude Opus 4.855.70$10.00
6GPT-5.6 Terra55136$4.50
7GPT-5.554.80$11.25
8Grok 4.553.856$3.00
9Claude Opus 4.753.50$10.00
10Claude Sonnet 553.486$4.00

Output tokens per second — higher is faster. Minimum intelligence score of 40.

#Modeltok/s
1Gemini 3.5 Flash268
2Gemini 3.6 Flash230
3Qwen3.7 Max204
4GLM-5.2187
5GPT-5.6 Luna174
6Muse Spark 1.1172
7GPT-5.6 Terra136
8Gemini 3.1 Pro Preview129
9GPT-5.3 Codex129
10Nex-N2-Pro128

Blended cost per 1M tokens (3:1 input/output) — lower is cheaper. Minimum intelligence score of 40.

#Model$/1M
1DeepSeek V4 Flash$0.171
2DeepSeek V4 Flash 0731$0.175
3Hy3$0.241
4GPT-5.6 Luna$0.45
5MiniMax-M3$0.525
6Inkling Small$0.525
7DeepSeek V4 Pro$0.544
8MiMo-V2.5-Pro$0.544
9Nex-N2-Pro$1.00
10GPT-5.4 mini$1.69