The SWE-rebench rankings remain unchanged from the previous cycle, with Anthropic's Fable 5 holding the top position at 64.5 percent, followed by Grok 4.5 at 63.8 percent and Opus 5 at 63.4 percent; the stability across the top tier reflects the difficulty of moving beyond these performance plateaus, though the confidence intervals suggest room for meaningful separation if the variance decreases. The Artificial Analysis leaderboard, by contrast, shows considerable churn in the middle ranks, with DeepSeek V4 Flash 0731 entering at position 17 with a score of 49.9, displacing the previous entry and pushing Claude Sonnet 4.6 from 17th to 18th, yet this movement involves models scoring within a compressed band between 47 and 50 where differentiation becomes fragile. At the top of Artificial Analysis, Claude Opus 5 leads at 60.7, a position it does not hold on SWE-rebench where it ranks third; this discrepancy points to benchmark sensitivity in methodology or task composition, suggesting that neither leaderboard alone captures the full picture of coding capability. The SWE-rebench data carries reported confidence intervals, which provide some guard against false precision, whereas Artificial Analysis reports point estimates without uncertainty quantification, making it harder to assess whether observed shifts represent genuine performance changes or measurement noise. What stands out is not volatility at the extremes but rather the emergence of a dense cluster of models in the 40 to 50 range on both benchmarks, indicating that the frontier is broadening even as the absolute ceiling remains contested.
Cole Brennan
Daily rankings from SWE-rebench, a benchmark designed to fairly compare LLM capabilities on real-world software engineering tasks. Unlike other evaluations, it uses a standardized scaffolding for all models, continuously updates its dataset to prevent contamination, and runs each model five times to account for stochastic variance.
| # | Model | Score |
|---|---|---|
| 1 | AnthropicFable 5 [high]Model | 64.5%± 1.41% |
| 2 | GrokGrok 4.5 [high]Model | 63.8%± 0.60% |
| 3 | AnthropicOpus 5 [high]Model | 63.4%± 1.35% |
| 4 | Z.aiGLM-5.2 [high]Model | 62.9%± 1.19% |
| 5 | OpenAIGPT-5.6 Sol [medium]Model | 62.3%± 1.83% |
| 6 | JunieJunieAgent | 61.8%± 0.54% |
| 7 | AnthropicClaude CodeAgent | 60.4%± 1.03% |
| 8 | OpenAICodexAgent | 58.0%± 1.29% |
| 9 | AnthropicSonnet 5 [high]Model | 56.8%± 0.94% |
| 10 | CursorCursorAgent | 51.7%± 0.84% |
Artificial Analysis composite index across coding, math, and reasoning benchmarks.
| # | Model | Score | tok/s | $/1M |
|---|---|---|---|---|
| 1 | Claude Opus 5 | 60.7 | 55 | $10.00 |
| 2 | Claude Fable 5 | 59.9 | 59 | $20.00 |
| 3 | GPT-5.6 Sol | 58.9 | 67 | $11.25 |
| 4 | Kimi K3 | 57.1 | 33 | $6.00 |
| 5 | Claude Opus 4.8 | 55.7 | 52 | $10.00 |
| 6 | GPT-5.6 Terra | 55 | 128 | $4.50 |
| 7 | GPT-5.5 | 54.8 | 0 | $11.25 |
| 8 | Grok 4.5 | 53.8 | 55 | $3.00 |
| 9 | Claude Opus 4.7 | 53.5 | 0 | $10.00 |
| 10 | Claude Sonnet 5 | 53.4 | 85 | $4.00 |
Output tokens per second — higher is faster. Minimum intelligence score of 40.
| # | Model | tok/s |
|---|---|---|
| 1 | Gemini 3.5 Flash | 220 |
| 2 | Gemini 3.6 Flash | 207 |
| 3 | Qwen3.7 Max | 200 |
| 4 | GPT-5.6 Luna | 187 |
| 5 | Muse Spark 1.1 | 140 |
| 6 | Gemini 3.1 Pro Preview | 136 |
| 7 | Nex-N2-Pro | 133 |
| 8 | GPT-5.6 Terra | 128 |
| 9 | GPT-5.3 Codex | 108 |
| 10 | GLM-5.2 | 107 |
Blended cost per 1M tokens (3:1 input/output) — lower is cheaper. Minimum intelligence score of 40.
| # | Model | $/1M |
|---|---|---|
| 1 | DeepSeek V4 Flash 0731 | $0.175 |
| 2 | DeepSeek V4 Flash | $0.175 |
| 3 | Hy3 | $0.241 |
| 4 | GPT-5.6 Luna | $0.45 |
| 5 | MiniMax-M3 | $0.525 |
| 6 | Inkling Small | $0.525 |
| 7 | DeepSeek V4 Pro | $0.544 |
| 8 | MiMo-V2.5-Pro | $0.544 |
| 9 | Nex-N2-Pro | $1.00 |
| 10 | GPT-5.4 mini | $1.69 |