The SWE-rebench results show no movement in the top tier, with Anthropic Fable 5 maintaining 64.5% ± 1.41% at rank one and the full top ten remaining locked in place. Grok 4.5 stays at 63.8% ± 0.60%, Opus 5 at 63.4% ± 1.35%, and the spread between positions one and ten spans 12.8 percentage points. The Artificial Analysis benchmark, by contrast, exhibits substantial churn across its 442 entries, with Claude Opus 5 leading at 63.1 and K-EXAONE 2.0 newly entering at rank 113 with 31.0, while K-EXAONE 2.0 0803 simultaneously dropped from that same position. This divergence signals a methodological gap worth noting: SWE-rebench isolates a narrow, high-signal task (software engineering problem resolution) where the top performers have reached a plateau tight enough that uncertainty intervals overlap, whereas Artificial Analysis aggregates across a broader evaluation surface where models shuffle constantly. The SWE-rebench stability is not stagnation but rather statistical saturation among frontier systems; further discrimination at this level would require either larger sample sizes to tighten confidence bands or harder problem instances. The Artificial Analysis flux, meanwhile, reflects sensitivity to model versioning and evaluation scope, making it less reliable for tracking real capability shifts in code-specific reasoning.
Cole Brennan
Daily rankings from SWE-rebench, a benchmark designed to fairly compare LLM capabilities on real-world software engineering tasks. Unlike other evaluations, it uses a standardized scaffolding for all models, continuously updates its dataset to prevent contamination, and runs each model five times to account for stochastic variance.
| # | Model | Score |
|---|---|---|
| 1 | AnthropicFable 5 [high]Model | 64.5%± 1.41% |
| 2 | GrokGrok 4.5 [high]Model | 63.8%± 0.60% |
| 3 | AnthropicOpus 5 [high]Model | 63.4%± 1.35% |
| 4 | Z.aiGLM-5.2 [high]Model | 62.9%± 1.19% |
| 5 | OpenAIGPT-5.6 Sol [medium]Model | 62.3%± 1.83% |
| 6 | JunieJunieAgent | 61.8%± 0.54% |
| 7 | AnthropicClaude CodeAgent | 60.4%± 1.03% |
| 8 | OpenAICodexAgent | 58.0%± 1.29% |
| 9 | AnthropicSonnet 5 [high]Model | 56.8%± 0.94% |
| 10 | CursorCursorAgent | 51.7%± 0.84% |
Artificial Analysis composite index across coding, math, and reasoning benchmarks.
| # | Model | Score | tok/s | $/1M |
|---|---|---|---|---|
| 1 | Claude Opus 5 | 63.1 | 53 | $10.00 |
| 2 | Claude Fable 5 | 62.1 | 65 | $20.00 |
| 3 | GPT-5.6 Sol | 60.9 | 81 | $8.00 |
| 4 | Grok 4.6 | 60.9 | 60 | $3.00 |
| 5 | Kimi K3 | 59.7 | 36 | $6.00 |
| 6 | GLM-5.3 | 59.5 | 75 | $2.15 |
| 7 | Qwen3.8 Max | 58.1 | 27 | $3.00 |
| 8 | Qwen3.8 2.4T A95B | 57.7 | 24 | $3.00 |
| 9 | GLM-5.3-Flash | 57.5 | 49 | $0.237 |
| 10 | Claude Opus 4.8 | 57.3 | 0 | $10.00 |
Output tokens per second — higher is faster. Minimum intelligence score of 40.
| # | Model | tok/s |
|---|---|---|
| 1 | Gemini 3.7 Flash | 376 |
| 2 | Gemini 3.6 Flash | 177 |
| 3 | Agnes 2.5 Pro Beta | 151 |
| 4 | Nex-N2-Pro | 137 |
| 5 | DeepSeek V4 Flash 0731 | 135 |
| 6 | GPT-5.3 Codex | 126 |
| 7 | GPT-5.6 Luna | 124 |
| 8 | Gemini 3.1 Pro Preview | 122 |
| 9 | DeepSeek V4 Flash Vision | 117 |
| 10 | GPT-5.6 Terra | 107 |
Blended cost per 1M tokens (3:1 input/output) — lower is cheaper. Minimum intelligence score of 40.
| # | Model | $/1M |
|---|---|---|
| 1 | Agnes 2.5 Pro Beta | $0.15 |
| 2 | DeepSeek V4 Flash | $0.168 |
| 3 | Qwen3.8-Flash-Next | $0.23 |
| 4 | GLM-5.3-Flash | $0.237 |
| 5 | Hy3 | $0.241 |
| 6 | GPT-5.6 Luna | $0.45 |
| 7 | MiniMax-M3 | $0.525 |
| 8 | Solar Pro 4 | $0.525 |
| 9 | Inkling Small | $0.525 |
| 10 | DeepSeek V4 Pro | $0.544 |