The Inference Report

July 29, 2026

The SWE-rebench ranking underwent wholesale turnover, with five new models entering the top tier and five prior leaders removed entirely. AnthropicFable 5 now leads at 64.5% ±1.41%, displacing OpenAI's gpt-5.5-2026-04-23-xhigh (which scored 62.7% ±0.91% previously and has been dropped), while GrokGrok 4.5 and AnthropicOpus 5 occupy the second and third positions at 63.8% and 63.4% respectively. The prior top performer, OpenAI's gpt-5.5-2026-04-23-xhigh, along with four other models designated as medium or xhigh variants, disappeared from the rankings entirely, suggesting either a methodology change, model retirement, or evaluation discontinuation rather than performance degradation. Among models that remained in the benchmark, JunieJunieAgent advanced from #2 to #6 despite a negligible score change (61.6% to 61.8%), while OpenAICodexAgent dropped sharply from #3 to #8 (60.4% down to 58.0%), and CursorCursorAgent fell from #9 to #10 (53.0% to 51.7%). Lower-ranked models show more volatile movement: MiniMax M3 climbed from #17 to #11 (45.6% to 47.2%), and MiMo V2.5 Pro rose from #19 to #12 (42.4% to 46.5%), while Qwen3.6-27B and Qwen3.6-35B-A3B both declined substantially, suggesting either model-specific improvements in the new cohort or changes in test conditions rather than uniform capability shifts. The confidence intervals remain narrow (typically ±0.5% to ±1.8%), indicating stable measurement precision, but the simultaneous departure of five prior leaders and arrival of five new ones raises a methodological question: whether this reflects genuine capability advances or a recalibrated evaluation protocol.

Cole Brennan

Daily rankings from SWE-rebench, a benchmark designed to fairly compare LLM capabilities on real-world software engineering tasks. Unlike other evaluations, it uses a standardized scaffolding for all models, continuously updates its dataset to prevent contamination, and runs each model five times to account for stochastic variance.

#ModelScore
1AnthropicFable 5 [high]Model64.5%± 1.41%
2GrokGrok 4.5 [high]Model63.8%± 0.60%
3AnthropicOpus 5 [high]Model63.4%± 1.35%
4Z.aiGLM-5.2 [high]Model62.9%± 1.19%
5OpenAIGPT-5.6 Sol [medium]Model62.3%± 1.83%
6JunieJunieAgent61.8%± 0.54%
7AnthropicClaude CodeAgent60.4%± 1.03%
8OpenAICodexAgent58.0%± 1.29%
9AnthropicSonnet 5 [high]Model56.8%± 0.94%
10CursorCursorAgent51.7%± 0.84%

Artificial Analysis composite index across coding, math, and reasoning benchmarks.

#ModelScoretok/s$/1M
1Claude Opus 560.761$10.00
2Claude Fable 559.970$20.00
3GPT-5.6 Sol58.978$11.25
4Kimi K357.133$6.00
5Claude Opus 4.855.767$10.00
6GPT-5.6 Terra55156$5.63
7GPT-5.554.80$11.25
8Grok 4.553.861$3.00
9Claude Opus 4.753.50$10.00
10Claude Sonnet 553.491$4.00

Output tokens per second — higher is faster. Minimum intelligence score of 40.

#Modeltok/s
1Gemini 3.5 Flash264
2Gemini 3.6 Flash248
3GLM-5.2215
4GPT-5.6 Luna205
5Qwen3.7 Max202
6GPT-5.6 Terra156
7Gemini 3.1 Pro Preview146
8Nex-N2-Pro142
9Muse Spark 1.1135
10GPT-5.3 Codex128

Blended cost per 1M tokens (3:1 input/output) — lower is cheaper. Minimum intelligence score of 40.

#Model$/1M
1DeepSeek V4 Flash$0.175
2Hy3$0.241
3MiniMax-M3$0.525
4DeepSeek V4 Pro$0.544
5MiMo-V2.5-Pro$0.544
6Nex-N2-Pro$1.00
7GPT-5.4 mini$1.69
8Kimi K2.6$1.71
9Kimi K2.7 Code$1.71
10Muse Spark 1.1$2.00