The SWE-rebench leaderboard holds stable at the top, with AnthropicFable 5 maintaining 64.5% and the top five models clustered within 2.2 percentage points, confidence intervals overlapping substantially enough that ranking shifts between them would not represent genuine performance differences. The Artificial Analysis benchmark, by contrast, shows considerable churn across its 422-entry list, with models shuffling positions frequently even when scores differ by 0.1 to 0.3 points, a pattern that raises questions about whether such fine-grained distinctions reflect real capability gaps or measurement noise. Notable movements include Ling 3.0 Flash dropping from rank 59 (37.8) to unranked status, NVIDIA Nemotron 3 Super 120B falling from 119 (25.7) to unranked, Ling 3.0 Tiny climbing from 125 (23.9) to 121 (24.5), and Phi-4 Mini Instruct rising from 326 (5.7) to 316 (6.2). The SWE-rebench methodology, which tests agents on actual software engineering tasks with controlled evaluation conditions, produces the more interpretable signal: the top tier of models genuinely solves 56 to 65 percent of problems, with uncertainty bands that reflect real variability. The Artificial Analysis scores, spanning from 64.5 down to 1.0 across hundreds of entries, compress evaluation into a single dimension that may conflate different failure modes or reflect task-specific brittleness rather than general capability. Without visibility into Artificial Analysis's evaluation protocol, whether it uses held-out test sets, how it handles edge cases, or whether scoring is normalized, the frequent micro-movements and aggressive differentiation at the tail end suggest the ranking may be sensitive to evaluation artifacts rather than tracking reproducible differences in model behavior.
Cole Brennan
Daily rankings from SWE-rebench, a benchmark designed to fairly compare LLM capabilities on real-world software engineering tasks. Unlike other evaluations, it uses a standardized scaffolding for all models, continuously updates its dataset to prevent contamination, and runs each model five times to account for stochastic variance.
| # | Model | Score |
|---|---|---|
| 1 | AnthropicFable 5 [high]Model | 64.5%± 1.41% |
| 2 | GrokGrok 4.5 [high]Model | 63.8%± 0.60% |
| 3 | AnthropicOpus 5 [high]Model | 63.4%± 1.35% |
| 4 | Z.aiGLM-5.2 [high]Model | 62.9%± 1.19% |
| 5 | OpenAIGPT-5.6 Sol [medium]Model | 62.3%± 1.83% |
| 6 | JunieJunieAgent | 61.8%± 0.54% |
| 7 | AnthropicClaude CodeAgent | 60.4%± 1.03% |
| 8 | OpenAICodexAgent | 58.0%± 1.29% |
| 9 | AnthropicSonnet 5 [high]Model | 56.8%± 0.94% |
| 10 | CursorCursorAgent | 51.7%± 0.84% |
Artificial Analysis composite index across coding, math, and reasoning benchmarks.
| # | Model | Score | tok/s | $/1M |
|---|---|---|---|---|
| 1 | Claude Opus 5 | 63.1 | 48 | $10.00 |
| 2 | Claude Fable 5 | 62.1 | 58 | $20.00 |
| 3 | GPT-5.6 Sol | 60.9 | 63 | $11.25 |
| 4 | Kimi K3 | 59.7 | 37 | $6.00 |
| 5 | Qwen3.8 Max | 58.1 | 69 | $3.00 |
| 6 | Claude Opus 4.8 | 57.3 | 0 | $10.00 |
| 7 | Muse Spark 1.2 | 56.8 | 0 | $2.00 |
| 8 | GPT-5.6 Terra | 56.6 | 115 | $4.50 |
| 9 | GPT-5.5 | 56.3 | 0 | $11.25 |
| 10 | Grok 4.5 | 55.8 | 51 | $3.00 |
Output tokens per second — higher is faster. Minimum intelligence score of 40.
| # | Model | tok/s |
|---|---|---|
| 1 | Gemini 3.5 Flash | 219 |
| 2 | Qwen3.7 Max | 195 |
| 3 | Muse Spark 1.1 | 191 |
| 4 | Gemini 3.6 Flash | 189 |
| 5 | GPT-5.6 Luna | 176 |
| 6 | Inkling Small | 130 |
| 7 | Gemini 3.1 Pro Preview | 127 |
| 8 | GLM-5.2 | 119 |
| 9 | Nex-N2-Pro | 118 |
| 10 | GPT-5.6 Terra | 115 |
Blended cost per 1M tokens (3:1 input/output) — lower is cheaper. Minimum intelligence score of 40.
| # | Model | $/1M |
|---|---|---|
| 1 | DeepSeek V4 Flash | $0.168 |
| 2 | DeepSeek V4 Flash 0731 | $0.175 |
| 3 | Hy3 | $0.241 |
| 4 | GPT-5.6 Luna | $0.45 |
| 5 | MiniMax-M3 | $0.525 |
| 6 | Inkling Small | $0.525 |
| 7 | DeepSeek V4 Pro | $0.544 |
| 8 | MiMo-V2.5-Pro | $0.544 |
| 9 | Nex-N2-Pro | $1.00 |
| 10 | Qwen3.6 Plus | $1.13 |