On SWE-rebench, the coding agent benchmark, the top rankings show no movement: AnthropicFable 5 holds first place at 64.5% ± 1.41%, followed by GrokGrok 4.5 at 63.8% ± 0.60% and AnthropicOpus 5 at 63.4% ± 1.35%, with confidence intervals that do not overlap meaningfully across the top five positions. The Artificial Analysis benchmark tells a different story, where two new entries appear in the top tier: Qwen3.8 2.4T A95B enters at #7 with 57.7, pushing Claude Opus 4.8 down one position, and Gemini 3.7 Flash debuts at #12 with 56.0. The broader rankings remain largely stable across 432 positions, with most models retaining their previous slots or shifting by single positions. The consistency across SWE-rebench contrasts with modest churn in Artificial Analysis, where new models occasionally displace prior entrants but without dramatic reshuffling. Neither benchmark shows the kind of compression at the top that would indicate a breakthrough in reasoning or coding capability; instead, the data reflects incremental refinement and occasional new model releases entering established hierarchies. The stability of the SWE-rebench top ten, despite its larger confidence intervals, suggests the evaluation has reached a plateau where further gains require architectural innovation rather than parameter tuning.
Cole Brennan
Daily rankings from SWE-rebench, a benchmark designed to fairly compare LLM capabilities on real-world software engineering tasks. Unlike other evaluations, it uses a standardized scaffolding for all models, continuously updates its dataset to prevent contamination, and runs each model five times to account for stochastic variance.
| # | Model | Score |
|---|---|---|
| 1 | AnthropicFable 5 [high]Model | 64.5%± 1.41% |
| 2 | GrokGrok 4.5 [high]Model | 63.8%± 0.60% |
| 3 | AnthropicOpus 5 [high]Model | 63.4%± 1.35% |
| 4 | Z.aiGLM-5.2 [high]Model | 62.9%± 1.19% |
| 5 | OpenAIGPT-5.6 Sol [medium]Model | 62.3%± 1.83% |
| 6 | JunieJunieAgent | 61.8%± 0.54% |
| 7 | AnthropicClaude CodeAgent | 60.4%± 1.03% |
| 8 | OpenAICodexAgent | 58.0%± 1.29% |
| 9 | AnthropicSonnet 5 [high]Model | 56.8%± 0.94% |
| 10 | CursorCursorAgent | 51.7%± 0.84% |
Artificial Analysis composite index across coding, math, and reasoning benchmarks.
| # | Model | Score | tok/s | $/1M |
|---|---|---|---|---|
| 1 | Claude Opus 5 | 63.1 | 50 | $10.00 |
| 2 | Claude Fable 5 | 62.1 | 60 | $20.00 |
| 3 | GPT-5.6 Sol | 60.9 | 61 | $11.25 |
| 4 | Grok 4.6 | 60.9 | 54 | $3.00 |
| 5 | Kimi K3 | 59.7 | 37 | $6.00 |
| 6 | Qwen3.8 Max | 58.1 | 44 | $3.00 |
| 7 | Qwen3.8 2.4T A95B | 57.7 | 50 | $3.00 |
| 8 | Claude Opus 4.8 | 57.3 | 0 | $10.00 |
| 9 | Muse Spark 1.2 | 56.8 | 0 | $2.00 |
| 10 | GPT-5.6 Terra | 56.6 | 106 | $4.50 |
Output tokens per second — higher is faster. Minimum intelligence score of 40.
| # | Model | tok/s |
|---|---|---|
| 1 | Gemini 3.7 Flash | 515 |
| 2 | Gemini 3.6 Flash | 207 |
| 3 | GPT-5.6 Luna | 156 |
| 4 | Nex-N2-Pro | 139 |
| 5 | Gemini 3.1 Pro Preview | 127 |
| 6 | Inkling Small | 125 |
| 7 | GLM-5.2 | 110 |
| 8 | DeepSeek V4 Flash 0731 | 109 |
| 9 | GPT-5.3 Codex | 109 |
| 10 | GPT-5.6 Terra | 106 |
Blended cost per 1M tokens (3:1 input/output) — lower is cheaper. Minimum intelligence score of 40.
| # | Model | $/1M |
|---|---|---|
| 1 | DeepSeek V4 Flash | $0.168 |
| 2 | DeepSeek V4 Flash 0731 | $0.175 |
| 3 | Hy3 | $0.241 |
| 4 | GPT-5.6 Luna | $0.45 |
| 5 | MiniMax-M3 | $0.525 |
| 6 | Solar Pro 4 | $0.525 |
| 7 | Inkling Small | $0.525 |
| 8 | DeepSeek V4 Pro | $0.544 |
| 9 | MiMo-V2.5-Pro | $0.544 |
| 10 | Nex-N2-Pro | $1.00 |