The Inference Report

August 4, 2026

The SWE-rebench rankings remain unchanged at the top tier, with AnthropicFable 5 holding at 64.5 percent and the next eight positions identical to the previous cycle, but the Artificial Analysis leaderboard has undergone significant restructuring below the top 20. Claude Opus 5 and Claude Fable 5 have swapped positions, now ranking first and second respectively at 60.7 and 59.9, displacing GPT-5.6 Sol to third, while Kimi K3 enters the top five at 57.1. The most substantial movement occurs in the 80 to 420 entry range, where G9v3-39A5B debuts at position 88 on Artificial Analysis, pushing prior entries down by one rank through the tail of the list. Within the SWE-rebench benchmark, the consistency of scores and confidence intervals across the top performers suggests either a plateau in incremental gains or stable evaluation conditions, though the methodology of both benchmarks warrants scrutiny: SWE-rebench appears to measure code generation on pull request resolution tasks with relatively tight confidence bounds, while Artificial Analysis aggregates across a broader set of capabilities, making direct comparison between the two frameworks problematic. The data shows no evidence of breakthrough performance on either metric; the top SWE-rebench model at 64.5 percent leaves substantial room for improvement on a task designed to reflect real-world software engineering, and the Artificial Analysis rankings reflect primarily reshuffling rather than systematic score inflation across the board.

Cole Brennan

Daily rankings from SWE-rebench, a benchmark designed to fairly compare LLM capabilities on real-world software engineering tasks. Unlike other evaluations, it uses a standardized scaffolding for all models, continuously updates its dataset to prevent contamination, and runs each model five times to account for stochastic variance.

#ModelScore
1AnthropicFable 5 [high]Model64.5%± 1.41%
2GrokGrok 4.5 [high]Model63.8%± 0.60%
3AnthropicOpus 5 [high]Model63.4%± 1.35%
4Z.aiGLM-5.2 [high]Model62.9%± 1.19%
5OpenAIGPT-5.6 Sol [medium]Model62.3%± 1.83%
6JunieJunieAgent61.8%± 0.54%
7AnthropicClaude CodeAgent60.4%± 1.03%
8OpenAICodexAgent58.0%± 1.29%
9AnthropicSonnet 5 [high]Model56.8%± 0.94%
10CursorCursorAgent51.7%± 0.84%

Artificial Analysis composite index across coding, math, and reasoning benchmarks.

#ModelScoretok/s$/1M
1Claude Opus 560.760$10.00
2Claude Fable 559.967$20.00
3GPT-5.6 Sol58.980$11.25
4Kimi K357.134$6.00
5Claude Opus 4.855.70$10.00
6GPT-5.6 Terra55131$4.50
7GPT-5.554.80$11.25
8Grok 4.553.863$3.00
9Claude Opus 4.753.50$10.00
10Claude Sonnet 553.488$4.00

Output tokens per second — higher is faster. Minimum intelligence score of 40.

#Modeltok/s
1Gemini 3.5 Flash281
2Gemini 3.6 Flash231
3GLM-5.2212
4Qwen3.7 Max208
5Muse Spark 1.1206
6GPT-5.6 Luna172
7Nex-N2-Pro136
8Gemini 3.1 Pro Preview134
9GPT-5.6 Terra131
10Inkling Small131

Blended cost per 1M tokens (3:1 input/output) — lower is cheaper. Minimum intelligence score of 40.

#Model$/1M
1DeepSeek V4 Flash$0.171
2DeepSeek V4 Flash 0731$0.175
3Hy3$0.241
4GPT-5.6 Luna$0.45
5MiniMax-M3$0.525
6Inkling Small$0.525
7DeepSeek V4 Pro$0.544
8MiMo-V2.5-Pro$0.544
9Nex-N2-Pro$1.00
10GPT-5.4 mini$1.69