The SWE-rebench leaderboard shows no movement in the top tier, with AnthropicFable 5 holding 64.5% and the next four models (Grok 4.5, Opus 5, GLM-5.2, GPT-5.6 Sol) maintaining their positions between 62.3% and 63.8%. The confidence intervals remain wide enough that none of these differences reach statistical significance, and the unchanged rankings suggest either genuine stability in agent performance on software engineering tasks or insufficient sensitivity in the benchmark to detect real improvements. The Artificial Analysis leaderboard, by contrast, shows substantial churn across the full 400+ model range, with entries like Motif 3 rising from 44.9 to 45.3 (moving from #27 to #26), Nova 2.0 Lite jumping from #157 with 19.2 to #145 with 20.8, and KAT Coder Pro V2 climbing from #81 at 33.9 to #74 at 34.5. This divergence is instructive: the SWE-rebench scores derive from controlled evaluation on real repositories with defined success criteria, while Artificial Analysis aggregates performance across broader language understanding benchmarks where small numerical shifts propagate into ranking changes. The stability at the top of SWE-rebench may reflect that frontier agent performance has plateaued on this particular task distribution, or that the evaluation's variance makes sub-1% score differences unreliable for ranking claims. Without methodological details on how Artificial Analysis scores are computed or weighted, it is difficult to assess whether its volatility represents genuine capability shifts or measurement noise.
Cole Brennan
Daily rankings from SWE-rebench, a benchmark designed to fairly compare LLM capabilities on real-world software engineering tasks. Unlike other evaluations, it uses a standardized scaffolding for all models, continuously updates its dataset to prevent contamination, and runs each model five times to account for stochastic variance.
| # | Model | Score |
|---|---|---|
| 1 | AnthropicFable 5 [high]Model | 64.5%± 1.41% |
| 2 | GrokGrok 4.5 [high]Model | 63.8%± 0.60% |
| 3 | AnthropicOpus 5 [high]Model | 63.4%± 1.35% |
| 4 | Z.aiGLM-5.2 [high]Model | 62.9%± 1.19% |
| 5 | OpenAIGPT-5.6 Sol [medium]Model | 62.3%± 1.83% |
| 6 | JunieJunieAgent | 61.8%± 0.54% |
| 7 | AnthropicClaude CodeAgent | 60.4%± 1.03% |
| 8 | OpenAICodexAgent | 58.0%± 1.29% |
| 9 | AnthropicSonnet 5 [high]Model | 56.8%± 0.94% |
| 10 | CursorCursorAgent | 51.7%± 0.84% |
Artificial Analysis composite index across coding, math, and reasoning benchmarks.
| # | Model | Score | tok/s | $/1M |
|---|---|---|---|---|
| 1 | Claude Opus 5 | 63.1 | 49 | $10.00 |
| 2 | Claude Fable 5 | 62.1 | 63 | $20.00 |
| 3 | GPT-5.6 Sol | 60.9 | 68 | $11.25 |
| 4 | Kimi K3 | 59.7 | 41 | $6.00 |
| 5 | Qwen3.8 Max | 58.1 | 54 | $3.00 |
| 6 | Claude Opus 4.8 | 57.3 | 0 | $10.00 |
| 7 | Muse Spark 1.2 | 56.8 | 0 | $2.00 |
| 8 | GPT-5.6 Terra | 56.6 | 121 | $4.50 |
| 9 | GPT-5.5 | 56.3 | 0 | $11.25 |
| 10 | Grok 4.5 | 55.8 | 49 | $3.00 |
Output tokens per second — higher is faster. Minimum intelligence score of 40.
| # | Model | tok/s |
|---|---|---|
| 1 | Gemini 3.6 Flash | 196 |
| 2 | GPT-5.6 Luna | 163 |
| 3 | Inkling Small | 145 |
| 4 | Nex-N2-Pro | 142 |
| 5 | GPT-5.3 Codex | 133 |
| 6 | Gemini 3.1 Pro Preview | 126 |
| 7 | GLM-5.2 | 124 |
| 8 | GPT-5.6 Terra | 121 |
| 9 | DeepSeek V4 Flash 0731 | 112 |
| 10 | MiniMax-M3 | 93 |
Blended cost per 1M tokens (3:1 input/output) — lower is cheaper. Minimum intelligence score of 40.
| # | Model | $/1M |
|---|---|---|
| 1 | DeepSeek V4 Flash | $0.168 |
| 2 | DeepSeek V4 Flash 0731 | $0.175 |
| 3 | Hy3 | $0.241 |
| 4 | GPT-5.6 Luna | $0.45 |
| 5 | MiniMax-M3 | $0.525 |
| 6 | Inkling Small | $0.525 |
| 7 | DeepSeek V4 Pro | $0.544 |
| 8 | MiMo-V2.5-Pro | $0.544 |
| 9 | Nex-N2-Pro | $1.00 |
| 10 | Qwen3.6 Plus | $1.13 |