The SWE-rebench leaderboard shows no movement in the top tier, with AnthropicFable 5 maintaining 64.5% ± 1.41%, GrokGrok 4.5 holding 63.8% ± 0.60%, and AnthropicOpus 5 steady at 63.4% ± 1.35%. The Artificial Analysis benchmark, however, reveals more churn across its 417-entry roster, though the top positions remain occupied by Claude Opus 5 (60.7) and Claude Fable 5 (59.9). One new entry appears at rank 227 in Artificial Analysis: Celeris-1 at 11.8, which triggered a cascade of single-position shifts downward through the lower-middle tiers. The confidence intervals on SWE-rebench range from 0.54% to 1.83%, suggesting adequate precision for detecting real differences between adjacent models, yet the frozen top rankings across both benchmarks indicate either genuine convergence at the frontier or evaluation ceiling effects. The Artificial Analysis list, substantially longer and more volatile, shows that differentiation persists below the top 30, where scores drop from the 50s into the 40s and below. Without prior Artificial Analysis snapshots in the historical data, it is unclear whether the Celeris-1 insertion represents a new capability or a recalibration of the ranking system itself. SWE-rebench's stability at the top is consistent with the narrow margin between first and third place (1.1 percentage points), where confidence intervals overlap, making further separation difficult to resolve empirically.
Cole Brennan
Daily rankings from SWE-rebench, a benchmark designed to fairly compare LLM capabilities on real-world software engineering tasks. Unlike other evaluations, it uses a standardized scaffolding for all models, continuously updates its dataset to prevent contamination, and runs each model five times to account for stochastic variance.
| # | Model | Score |
|---|---|---|
| 1 | AnthropicFable 5 [high]Model | 64.5%± 1.41% |
| 2 | GrokGrok 4.5 [high]Model | 63.8%± 0.60% |
| 3 | AnthropicOpus 5 [high]Model | 63.4%± 1.35% |
| 4 | Z.aiGLM-5.2 [high]Model | 62.9%± 1.19% |
| 5 | OpenAIGPT-5.6 Sol [medium]Model | 62.3%± 1.83% |
| 6 | JunieJunieAgent | 61.8%± 0.54% |
| 7 | AnthropicClaude CodeAgent | 60.4%± 1.03% |
| 8 | OpenAICodexAgent | 58.0%± 1.29% |
| 9 | AnthropicSonnet 5 [high]Model | 56.8%± 0.94% |
| 10 | CursorCursorAgent | 51.7%± 0.84% |
Artificial Analysis composite index across coding, math, and reasoning benchmarks.
| # | Model | Score | tok/s | $/1M |
|---|---|---|---|---|
| 1 | Claude Opus 5 | 60.7 | 56 | $10.00 |
| 2 | Claude Fable 5 | 59.9 | 60 | $20.00 |
| 3 | GPT-5.6 Sol | 58.9 | 67 | $11.25 |
| 4 | Kimi K3 | 57.1 | 34 | $6.00 |
| 5 | Claude Opus 4.8 | 55.7 | 0 | $10.00 |
| 6 | GPT-5.6 Terra | 55 | 128 | $4.50 |
| 7 | GPT-5.5 | 54.8 | 0 | $11.25 |
| 8 | Grok 4.5 | 53.8 | 52 | $3.00 |
| 9 | Claude Opus 4.7 | 53.5 | 0 | $10.00 |
| 10 | Claude Sonnet 5 | 53.4 | 79 | $4.00 |
Output tokens per second — higher is faster. Minimum intelligence score of 40.
| # | Model | tok/s |
|---|---|---|
| 1 | Gemini 3.5 Flash | 220 |
| 2 | Gemini 3.6 Flash | 207 |
| 3 | Qwen3.7 Max | 200 |
| 4 | GPT-5.6 Luna | 172 |
| 5 | Nex-N2-Pro | 132 |
| 6 | Muse Spark 1.1 | 131 |
| 7 | Gemini 3.1 Pro Preview | 129 |
| 8 | GPT-5.6 Terra | 128 |
| 9 | GPT-5.3 Codex | 109 |
| 10 | DeepSeek V4 Flash | 102 |
Blended cost per 1M tokens (3:1 input/output) — lower is cheaper. Minimum intelligence score of 40.
| # | Model | $/1M |
|---|---|---|
| 1 | DeepSeek V4 Flash 0731 | $0.175 |
| 2 | DeepSeek V4 Flash | $0.175 |
| 3 | Hy3 | $0.241 |
| 4 | GPT-5.6 Luna | $0.45 |
| 5 | MiniMax-M3 | $0.525 |
| 6 | Inkling Small | $0.525 |
| 7 | DeepSeek V4 Pro | $0.544 |
| 8 | MiMo-V2.5-Pro | $0.544 |
| 9 | Nex-N2-Pro | $1.00 |
| 10 | GPT-5.4 mini | $1.69 |