The SWE-rebench rankings show no movement from the previous cycle: the top twelve entries remain in identical positions with matching scores and confidence intervals. AnthropicFable 5 holds first at 64.5% ± 1.41%, followed by GrokGrok 4.5 at 63.8% ± 0.60% and AnthropicOpus 5 at 63.4% ± 1.35%, with all downstream positions locked through rank seventeen. The stability across a full benchmark cycle suggests either that the test set has reached saturation or that the underlying model capabilities have plateaued relative to the task difficulty. On Artificial Analysis, by contrast, substantial reshuffling occurred in the top tier: Muse Spark 1.3 and GPT-6 Astra entered the ranking as new entries at positions three and five respectively, pushing prior occupants down, while Claude Fable 5.1 moved to first place at 65.7 from a prior position lower in the list. The broader Artificial Analysis leaderboard shows consistent but minor position shifts throughout the middle ranks, with no entries dropping off entirely and only two confirmed new additions in the visible window. The divergence between the two benchmarks, perfect stability on SWE-rebench versus active reordering on Artificial Analysis, raises a methodological question: SWE-rebench's narrow confidence intervals and lack of movement across its top entries suggest either tighter experimental controls or a ceiling effect that Artificial Analysis does not exhibit. Without access to sample sizes or retest protocols for either benchmark, it remains unclear whether the SWE-rebench plateau reflects genuine convergence in model performance or methodological constraints that limit discriminative power at the frontier.
Cole Brennan
Daily rankings from SWE-rebench, a benchmark designed to fairly compare LLM capabilities on real-world software engineering tasks. Unlike other evaluations, it uses a standardized scaffolding for all models, continuously updates its dataset to prevent contamination, and runs each model five times to account for stochastic variance.
| # | Model | Score |
|---|---|---|
| 1 | AnthropicFable 5 [high]Model | 64.5%± 1.41% |
| 2 | GrokGrok 4.5 [high]Model | 63.8%± 0.60% |
| 3 | AnthropicOpus 5 [high]Model | 63.4%± 1.35% |
| 4 | Z.aiGLM-5.2 [high]Model | 62.9%± 1.19% |
| 5 | OpenAIGPT-5.6 Sol [medium]Model | 62.3%± 1.83% |
| 6 | JunieJunieAgent | 61.8%± 0.54% |
| 7 | AnthropicClaude CodeAgent | 60.4%± 1.03% |
| 8 | OpenAICodexAgent | 58.0%± 1.29% |
| 9 | AnthropicSonnet 5 [high]Model | 56.8%± 0.94% |
| 10 | CursorCursorAgent | 51.7%± 0.84% |
Artificial Analysis composite index across coding, math, and reasoning benchmarks.
| # | Model | Score | tok/s | $/1M |
|---|---|---|---|---|
| 1 | Claude Fable 5.1 | 65.7 | 70 | $20.00 |
| 2 | Claude Opus 5 | 63.1 | 49 | $10.00 |
| 3 | Muse Spark 1.3 | 62.1 | 0 | $0.00 |
| 4 | Claude Fable 5 | 62.1 | 67 | $20.00 |
| 5 | GPT-6 Astra | 61.2 | 0 | $20.00 |
| 6 | GPT-5.6 Sol | 60.9 | 80 | $8.00 |
| 7 | Grok 4.6 | 60.9 | 60 | $3.00 |
| 8 | Kimi K3 | 59.7 | 38 | $6.00 |
| 9 | GLM-5.3 | 59.5 | 74 | $2.15 |
| 10 | Gemini 3.8 Flash | 58.7 | 312 | $1.50 |
Output tokens per second — higher is faster. Minimum intelligence score of 40.
| # | Model | tok/s |
|---|---|---|
| 1 | Gemini 3.8 Flash | 312 |
| 2 | Gemini 3.7 Flash | 292 |
| 3 | Muse Spark 1.2 | 201 |
| 4 | Gemini 3.6 Flash | 188 |
| 5 | Quasar 438B | 187 |
| 6 | Agnes 2.5 Pro Beta | 157 |
| 7 | Nex-N2-Pro | 134 |
| 8 | DeepSeek V4 Flash 0731 | 130 |
| 9 | GPT-5.3 Codex | 126 |
| 10 | Gemini 3.1 Pro Preview | 117 |
Blended cost per 1M tokens (3:1 input/output) — lower is cheaper. Minimum intelligence score of 40.
| # | Model | $/1M |
|---|---|---|
| 1 | Agnes 2.5 Pro Beta | $0.15 |
| 2 | DeepSeek V4 Flash | $0.168 |
| 3 | Qwen3.8-Flash-Next | $0.23 |
| 4 | GLM-5.3-Flash | $0.237 |
| 5 | Hy3 | $0.241 |
| 6 | GPT-5.6 Luna | $0.45 |
| 7 | MiniMax-M3 | $0.525 |
| 8 | Solar Pro 4 | $0.525 |
| 9 | Inkling Small | $0.525 |
| 10 | DeepSeek V4 Pro | $0.544 |