The SWE-rebench leaderboard shows stability at the top tier, with no movement in the top ten positions. AnthropicFable 5 maintains 64.5% ± 1.41%, followed by GrokGrok 4.5 at 63.8% ± 0.60% and AnthropicOpus 5 at 63.4% ± 1.35%, margins well within their confidence intervals. The Artificial Analysis benchmark, by contrast, reveals substantial reshuffling across the middle and lower ranks, with Claude Opus 5 rising to #1 at 63.1 (from #2 at 63.1), Claude Fable 5 at #2 with 62.1, and notable drops for models like DeepSeek V4 Pro, which fell from #16 at 53.2 to #30 at 45.3. The discrepancy between these two benchmarks warrants scrutiny: SWE-rebench uses controlled problem-solving conditions with uncertainty quantification, while Artificial Analysis provides point scores without error bounds, making direct comparison difficult. Within SWE-rebench's narrower scope, the consistency suggests the benchmark has sufficient resolution to distinguish performance in the 40-65% range but may be reaching saturation at the frontier, where confidence intervals begin to overlap meaningfully. The Artificial Analysis shifts hint either at different evaluation methodologies, data drift, or sensitivity to model versioning that SWE-rebench does not capture. Until both benchmarks publish their evaluation protocols in detail, the divergence remains interpretable but not fully explainable.
Cole Brennan
Daily rankings from SWE-rebench, a benchmark designed to fairly compare LLM capabilities on real-world software engineering tasks. Unlike other evaluations, it uses a standardized scaffolding for all models, continuously updates its dataset to prevent contamination, and runs each model five times to account for stochastic variance.
| # | Model | Score |
|---|---|---|
| 1 | AnthropicFable 5 [high]Model | 64.5%± 1.41% |
| 2 | GrokGrok 4.5 [high]Model | 63.8%± 0.60% |
| 3 | AnthropicOpus 5 [high]Model | 63.4%± 1.35% |
| 4 | Z.aiGLM-5.2 [high]Model | 62.9%± 1.19% |
| 5 | OpenAIGPT-5.6 Sol [medium]Model | 62.3%± 1.83% |
| 6 | JunieJunieAgent | 61.8%± 0.54% |
| 7 | AnthropicClaude CodeAgent | 60.4%± 1.03% |
| 8 | OpenAICodexAgent | 58.0%± 1.29% |
| 9 | AnthropicSonnet 5 [high]Model | 56.8%± 0.94% |
| 10 | CursorCursorAgent | 51.7%± 0.84% |
Artificial Analysis composite index across coding, math, and reasoning benchmarks.
| # | Model | Score | tok/s | $/1M |
|---|---|---|---|---|
| 1 | Claude Opus 5 | 63.1 | 46 | $10.00 |
| 2 | Claude Fable 5 | 62.1 | 58 | $20.00 |
| 3 | GPT-5.6 Sol | 60.9 | 60 | $11.25 |
| 4 | Grok 4.6 | 60.9 | 55 | $3.00 |
| 5 | Kimi K3 | 59.7 | 36 | $6.00 |
| 6 | Qwen3.8 Max | 58.1 | 43 | $3.00 |
| 7 | Qwen3.8 2.4T A95B | 57.7 | 48 | $3.00 |
| 8 | Claude Opus 4.8 | 57.3 | 0 | $10.00 |
| 9 | Muse Spark 1.2 | 56.8 | 0 | $2.00 |
| 10 | GPT-5.6 Terra | 56.6 | 102 | $4.50 |
Output tokens per second — higher is faster. Minimum intelligence score of 40.
| # | Model | tok/s |
|---|---|---|
| 1 | Gemini 3.7 Flash | 515 |
| 2 | Gemini 3.6 Flash | 205 |
| 3 | GPT-5.6 Luna | 148 |
| 4 | Nex-N2-Pro | 142 |
| 5 | Gemini 3.1 Pro Preview | 128 |
| 6 | Inkling Small | 113 |
| 7 | GPT-5.3 Codex | 112 |
| 8 | GLM-5.2 | 104 |
| 9 | DeepSeek V4 Flash 0731 | 103 |
| 10 | GPT-5.6 Terra | 102 |
Blended cost per 1M tokens (3:1 input/output) — lower is cheaper. Minimum intelligence score of 40.
| # | Model | $/1M |
|---|---|---|
| 1 | DeepSeek V4 Flash | $0.168 |
| 2 | Hy3 | $0.241 |
| 3 | GPT-5.6 Luna | $0.45 |
| 4 | MiniMax-M3 | $0.525 |
| 5 | Solar Pro 4 | $0.525 |
| 6 | Inkling Small | $0.525 |
| 7 | DeepSeek V4 Pro | $0.544 |
| 8 | MiMo-V2.5-Pro | $0.544 |
| 9 | DeepSeek V4 Flash 0731 | $0.66 |
| 10 | Nex-N2-Pro | $1.00 |