The Inference Report

August 7, 2026

The SWE-rebench standings show no movement at the top tier, with AnthropicFable 5 holding 64.5% and the next four positions (Grok 4.5, Opus 5, GLM-5.2, GPT-5.6 Sol) maintaining their rankings between 62.3% and 63.8%, all within margins that overlap their confidence intervals. The Artificial Analysis benchmark, by contrast, reveals substantial reordering across its full 422-entry leaderboard: Claude Opus 5 climbed to first place with 63.1 (up 2.4 points from 60.7), displacing Fable 5 to second at 62.1, while GPT-5.6 Sol moved from third to third with a 2.0-point gain to 60.9, and deeper in the list models like Muse Spark 1.2 jumped nine positions and 2.7 points to number 7. The divergence between these two benchmarks reflects their different methodologies: SWE-rebench measures coding agent performance on concrete repository tasks with tight confidence intervals that constrain apparent movement, while Artificial Analysis appears to evaluate broader model capabilities with larger score spreads. The top-tier SWE-rebench stability suggests either that the models tested are genuinely separated by meaningful gaps (the 1.1-point spread between first and fifth is real), or that the benchmark's variance is high enough to mask smaller changes, whereas Artificial Analysis's volatility in the 30-50 rank range and consistent 1-3 point gains across the board hints at either a recalibrated evaluation, a different test set, or systematic score inflation. Without knowing whether Artificial Analysis changed its methodology or simply re-ran the same test, the pattern is ambiguous: the gains could reflect genuine model improvement, measurement drift, or both.

Cole Brennan

Daily rankings from SWE-rebench, a benchmark designed to fairly compare LLM capabilities on real-world software engineering tasks. Unlike other evaluations, it uses a standardized scaffolding for all models, continuously updates its dataset to prevent contamination, and runs each model five times to account for stochastic variance.

#ModelScore
1AnthropicFable 5 [high]Model64.5%± 1.41%
2GrokGrok 4.5 [high]Model63.8%± 0.60%
3AnthropicOpus 5 [high]Model63.4%± 1.35%
4Z.aiGLM-5.2 [high]Model62.9%± 1.19%
5OpenAIGPT-5.6 Sol [medium]Model62.3%± 1.83%
6JunieJunieAgent61.8%± 0.54%
7AnthropicClaude CodeAgent60.4%± 1.03%
8OpenAICodexAgent58.0%± 1.29%
9AnthropicSonnet 5 [high]Model56.8%± 0.94%
10CursorCursorAgent51.7%± 0.84%

Artificial Analysis composite index across coding, math, and reasoning benchmarks.

#ModelScoretok/s$/1M
1Claude Opus 563.153$10.00
2Claude Fable 562.161$20.00
3GPT-5.6 Sol60.963$11.25
4Kimi K359.737$6.00
5Qwen3.8 Max58.169$3.00
6Claude Opus 4.857.30$10.00
7Muse Spark 1.256.80$2.00
8GPT-5.6 Terra56.6125$4.50
9GPT-5.556.30$11.25
10Grok 4.555.853$3.00

Output tokens per second — higher is faster. Minimum intelligence score of 40.

#Modeltok/s
1Gemini 3.5 Flash226
2Qwen3.7 Max196
3Muse Spark 1.1192
4Gemini 3.6 Flash189
5GPT-5.6 Luna169
6Gemini 3.1 Pro Preview128
7GPT-5.6 Terra125
8Nex-N2-Pro119
9Inkling Small119
10GLM-5.2115

Blended cost per 1M tokens (3:1 input/output) — lower is cheaper. Minimum intelligence score of 40.

#Model$/1M
1DeepSeek V4 Flash$0.168
2DeepSeek V4 Flash 0731$0.175
3Hy3$0.241
4GPT-5.6 Luna$0.45
5MiniMax-M3$0.525
6Inkling Small$0.525
7DeepSeek V4 Pro$0.544
8MiMo-V2.5-Pro$0.544
9Nex-N2-Pro$1.00
10Qwen3.6 Plus$1.13