The SWE-rebench leaderboard shows no movement across the top tier, with AnthropicFable 5 maintaining 64.5±1.41%, GrokGrok 4.5 at 63.8±0.60%, and AnthropicOpus 5 at 63.4±1.35%, their confidence intervals overlapping sufficiently that ranking stability reflects measurement precision rather than performance divergence. The Artificial Analysis benchmark, by contrast, exhibits substantial churn below the top 100 entries, though the methodology underlying these scores remains opaque compared to SWE-rebench's transparent evaluation against real software engineering tasks. Claude Opus 5 and Claude Fable 5 swap positions at the top of Artificial Analysis (63.1 and 62.1 respectively), while G9v3-39A5B climbs from position 102 to 92, gaining 2.4 points in the process, and KAT Coder Pro V2, Kimi K2 Thinking, and o3-pro each shift up one rank. The mid-tier reshuffling involves models like QwQ 32B and Qwen3 VL 30B A3B trading positions 224 and 225, suggesting either evaluation variance or incremental improvements below the noise floor. SWE-rebench's stability across all 17 ranked models implies the coding task distribution has reached a plateau where current approaches plateau, while Artificial Analysis's frequent reordering raises questions about whether the benchmark captures genuine capability differences or reflects sensitivity to prompt variation and evaluation artifact. Without documented methodology for Artificial Analysis, the practical significance of mid-list movements remains ambiguous.
Cole Brennan
Daily rankings from SWE-rebench, a benchmark designed to fairly compare LLM capabilities on real-world software engineering tasks. Unlike other evaluations, it uses a standardized scaffolding for all models, continuously updates its dataset to prevent contamination, and runs each model five times to account for stochastic variance.
| # | Model | Score |
|---|---|---|
| 1 | AnthropicFable 5 [high]Model | 64.5%± 1.41% |
| 2 | GrokGrok 4.5 [high]Model | 63.8%± 0.60% |
| 3 | AnthropicOpus 5 [high]Model | 63.4%± 1.35% |
| 4 | Z.aiGLM-5.2 [high]Model | 62.9%± 1.19% |
| 5 | OpenAIGPT-5.6 Sol [medium]Model | 62.3%± 1.83% |
| 6 | JunieJunieAgent | 61.8%± 0.54% |
| 7 | AnthropicClaude CodeAgent | 60.4%± 1.03% |
| 8 | OpenAICodexAgent | 58.0%± 1.29% |
| 9 | AnthropicSonnet 5 [high]Model | 56.8%± 0.94% |
| 10 | CursorCursorAgent | 51.7%± 0.84% |
Artificial Analysis composite index across coding, math, and reasoning benchmarks.
| # | Model | Score | tok/s | $/1M |
|---|---|---|---|---|
| 1 | Claude Opus 5 | 63.1 | 56 | $10.00 |
| 2 | Claude Fable 5 | 62.1 | 68 | $20.00 |
| 3 | GPT-5.6 Sol | 60.9 | 75 | $11.25 |
| 4 | Grok 4.6 | 60.9 | 62 | $3.00 |
| 5 | Kimi K3 | 59.7 | 39 | $6.00 |
| 6 | GLM-5.3 | 59.5 | 79 | $2.15 |
| 7 | Qwen3.8 Max | 58.1 | 45 | $3.00 |
| 8 | Qwen3.8 2.4T A95B | 57.7 | 45 | $3.00 |
| 9 | Claude Opus 4.8 | 57.3 | 0 | $10.00 |
| 10 | Muse Spark 1.2 | 56.8 | 0 | $2.00 |
Output tokens per second — higher is faster. Minimum intelligence score of 40.
| # | Model | tok/s |
|---|---|---|
| 1 | Gemini 3.7 Flash | 358 |
| 2 | Gemini 3.6 Flash | 201 |
| 3 | GPT-5.6 Luna | 138 |
| 4 | Nex-N2-Pro | 135 |
| 5 | GPT-5.3 Codex | 125 |
| 6 | Gemini 3.1 Pro Preview | 124 |
| 7 | DeepSeek V4 Flash 0731 | 122 |
| 8 | GPT-5.6 Terra | 114 |
| 9 | MiniMax-M3 | 108 |
| 10 | Claude Sonnet 5 | 87 |
Blended cost per 1M tokens (3:1 input/output) — lower is cheaper. Minimum intelligence score of 40.
| # | Model | $/1M |
|---|---|---|
| 1 | DeepSeek V4 Flash | $0.168 |
| 2 | Hy3 | $0.241 |
| 3 | GPT-5.6 Luna | $0.45 |
| 4 | MiniMax-M3 | $0.525 |
| 5 | Solar Pro 4 | $0.525 |
| 6 | Inkling Small | $0.525 |
| 7 | DeepSeek V4 Pro | $0.544 |
| 8 | MiMo-V2.5-Pro | $0.544 |
| 9 | DeepSeek V4 Flash 0731 | $0.66 |
| 10 | Nex-N2-Pro | $1.00 |