AnthropicFable 5 holds the top position on SWE-rebench at 64.5 percent, unchanged from the prior round, with confidence intervals tight enough to rule out meaningful movement at the top tier. GrokGrok 4.5 remains at 63.8 percent and AnthropicOpus 5 at 63.4 percent, preserving the same three-model hierarchy. The interval widths across the top six entries (ranging from 0.54 to 1.83 percentage points) suggest that real separation exists between the first tier and everything below rank six, though the 1.41 percent margin around Fable 5 leaves room for natural variation on re-evaluation. On Artificial Analysis, the leaderboard shows broader movement within the 50 to 65 percent band where most capable models cluster, with Claude Opus 5 at 63.1 and Claude Fable 5 at 62.1, yet these two benchmarks measure different problem sets and evaluation protocols, making direct score comparison misleading. The SWE-rebench top ten remains stable in composition and order, suggesting that the coding task distribution and model capability gaps have not shifted materially. Lower-ranked models show more scatter, particularly between positions 220 and 430 on Artificial Analysis where single-point score differences become common and ranking precision degrades. The lack of disruption at the summit indicates that incremental model releases have not yet displaced the established leaders, and the consistency of the top six across both evaluation windows points to reliable separation rather than noise. Whether this stability reflects genuine plateaus in coding task performance or merely the lag between training and evaluation remains an open question.
Cole Brennan
Daily rankings from SWE-rebench, a benchmark designed to fairly compare LLM capabilities on real-world software engineering tasks. Unlike other evaluations, it uses a standardized scaffolding for all models, continuously updates its dataset to prevent contamination, and runs each model five times to account for stochastic variance.
| # | Model | Score |
|---|---|---|
| 1 | AnthropicFable 5 [high]Model | 64.5%± 1.41% |
| 2 | GrokGrok 4.5 [high]Model | 63.8%± 0.60% |
| 3 | AnthropicOpus 5 [high]Model | 63.4%± 1.35% |
| 4 | Z.aiGLM-5.2 [high]Model | 62.9%± 1.19% |
| 5 | OpenAIGPT-5.6 Sol [medium]Model | 62.3%± 1.83% |
| 6 | JunieJunieAgent | 61.8%± 0.54% |
| 7 | AnthropicClaude CodeAgent | 60.4%± 1.03% |
| 8 | OpenAICodexAgent | 58.0%± 1.29% |
| 9 | AnthropicSonnet 5 [high]Model | 56.8%± 0.94% |
| 10 | CursorCursorAgent | 51.7%± 0.84% |
Artificial Analysis composite index across coding, math, and reasoning benchmarks.
| # | Model | Score | tok/s | $/1M |
|---|---|---|---|---|
| 1 | Claude Opus 5 | 63.1 | 49 | $10.00 |
| 2 | Claude Fable 5 | 62.1 | 60 | $20.00 |
| 3 | GPT-5.6 Sol | 60.9 | 61 | $11.25 |
| 4 | Grok 4.6 | 60.9 | 58 | $3.00 |
| 5 | Kimi K3 | 59.7 | 37 | $6.00 |
| 6 | Qwen3.8 Max | 58.1 | 44 | $3.00 |
| 7 | Qwen3.8 2.4T A95B | 57.7 | 48 | $3.00 |
| 8 | Claude Opus 4.8 | 57.3 | 0 | $10.00 |
| 9 | Muse Spark 1.2 | 56.8 | 0 | $2.00 |
| 10 | GPT-5.6 Terra | 56.6 | 103 | $4.50 |
Output tokens per second — higher is faster. Minimum intelligence score of 40.
| # | Model | tok/s |
|---|---|---|
| 1 | Gemini 3.6 Flash | 207 |
| 2 | GPT-5.6 Luna | 160 |
| 3 | Nex-N2-Pro | 143 |
| 4 | Gemini 3.1 Pro Preview | 131 |
| 5 | GPT-5.3 Codex | 120 |
| 6 | GLM-5.2 | 117 |
| 7 | DeepSeek V4 Flash 0731 | 104 |
| 8 | GPT-5.6 Terra | 103 |
| 9 | MiniMax-M3 | 97 |
| 10 | Inkling | 79 |
Blended cost per 1M tokens (3:1 input/output) — lower is cheaper. Minimum intelligence score of 40.
| # | Model | $/1M |
|---|---|---|
| 1 | DeepSeek V4 Flash | $0.168 |
| 2 | Hy3 | $0.241 |
| 3 | GPT-5.6 Luna | $0.45 |
| 4 | MiniMax-M3 | $0.525 |
| 5 | Solar Pro 4 | $0.525 |
| 6 | Inkling Small | $0.525 |
| 7 | DeepSeek V4 Pro | $0.544 |
| 8 | MiMo-V2.5-Pro | $0.544 |
| 9 | DeepSeek V4 Flash 0731 | $0.66 |
| 10 | Nex-N2-Pro | $1.00 |