SWE-rebench shows no movement in the top tier, with AnthropicFable 5 holding 64.5% and the next four models within a 2.2 percentage point band, all within their stated confidence intervals. Artificial Analysis reveals a different story: GLM-5.3 enters at rank 6 with 59.5, displacing Qwen3.8 Max to rank 7, a reordering that reflects genuine competition in the 59 to 60 range where four models now cluster. DeepSeek V3.1 Terminus climbs three positions to 104 from 107 with a 0.3-point gain (31.4 vs. 31.1), a marginal improvement that may signal refinement rather than breakthrough. The broader pattern across both benchmarks is consolidation rather than disruption: the top five on SWE-rebench remain unchanged, and Artificial Analysis shows mostly positional shuffles among models scoring below 35, where entry and exit are frequent. Neither benchmark exhibits the kind of score inflation or collapse that would suggest methodological drift, though the divergence between SWE-rebench (which measures code completion on real GitHub issues with stricter evaluation) and Artificial Analysis (which appears to use a different test set or scoring regime) deserves scrutiny. The confidence intervals on SWE-rebench, particularly the ±1.41% for Fable 5, are large enough that small reorderings in the 56 to 64 range would not be statistically meaningful; Artificial Analysis provides no error bounds, making it impossible to assess whether its movements reflect real capability changes or sampling variance.
Cole Brennan
Daily rankings from SWE-rebench, a benchmark designed to fairly compare LLM capabilities on real-world software engineering tasks. Unlike other evaluations, it uses a standardized scaffolding for all models, continuously updates its dataset to prevent contamination, and runs each model five times to account for stochastic variance.
| # | Model | Score |
|---|---|---|
| 1 | AnthropicFable 5 [high]Model | 64.5%± 1.41% |
| 2 | GrokGrok 4.5 [high]Model | 63.8%± 0.60% |
| 3 | AnthropicOpus 5 [high]Model | 63.4%± 1.35% |
| 4 | Z.aiGLM-5.2 [high]Model | 62.9%± 1.19% |
| 5 | OpenAIGPT-5.6 Sol [medium]Model | 62.3%± 1.83% |
| 6 | JunieJunieAgent | 61.8%± 0.54% |
| 7 | AnthropicClaude CodeAgent | 60.4%± 1.03% |
| 8 | OpenAICodexAgent | 58.0%± 1.29% |
| 9 | AnthropicSonnet 5 [high]Model | 56.8%± 0.94% |
| 10 | CursorCursorAgent | 51.7%± 0.84% |
Artificial Analysis composite index across coding, math, and reasoning benchmarks.
| # | Model | Score | tok/s | $/1M |
|---|---|---|---|---|
| 1 | Claude Opus 5 | 63.1 | 54 | $10.00 |
| 2 | Claude Fable 5 | 62.1 | 67 | $20.00 |
| 3 | GPT-5.6 Sol | 60.9 | 69 | $11.25 |
| 4 | Grok 4.6 | 60.9 | 66 | $3.00 |
| 5 | Kimi K3 | 59.7 | 39 | $6.00 |
| 6 | GLM-5.3 | 59.5 | 80 | $2.15 |
| 7 | Qwen3.8 Max | 58.1 | 47 | $3.00 |
| 8 | Qwen3.8 2.4T A95B | 57.7 | 47 | $3.00 |
| 9 | Claude Opus 4.8 | 57.3 | 0 | $10.00 |
| 10 | Muse Spark 1.2 | 56.8 | 0 | $2.00 |
Output tokens per second — higher is faster. Minimum intelligence score of 40.
| # | Model | tok/s |
|---|---|---|
| 1 | Gemini 3.7 Flash | 339 |
| 2 | Gemini 3.6 Flash | 187 |
| 3 | GPT-5.6 Luna | 165 |
| 4 | Nex-N2-Pro | 140 |
| 5 | GPT-5.3 Codex | 132 |
| 6 | Gemini 3.1 Pro Preview | 126 |
| 7 | GPT-5.6 Terra | 117 |
| 8 | DeepSeek V4 Flash 0731 | 115 |
| 9 | MiniMax-M3 | 104 |
| 10 | Claude Sonnet 5 | 82 |
Blended cost per 1M tokens (3:1 input/output) — lower is cheaper. Minimum intelligence score of 40.
| # | Model | $/1M |
|---|---|---|
| 1 | DeepSeek V4 Flash | $0.168 |
| 2 | Hy3 | $0.241 |
| 3 | GPT-5.6 Luna | $0.45 |
| 4 | MiniMax-M3 | $0.525 |
| 5 | Solar Pro 4 | $0.525 |
| 6 | Inkling Small | $0.525 |
| 7 | DeepSeek V4 Pro | $0.544 |
| 8 | MiMo-V2.5-Pro | $0.544 |
| 9 | DeepSeek V4 Flash 0731 | $0.66 |
| 10 | Nex-N2-Pro | $1.00 |