The SWE-rebench rankings show no movement from the previous cycle, with AnthropicFable 5 holding 64.5% ± 1.41% at the top, followed by GrokGrok 4.5 at 63.8% ± 0.60% and AnthropicOpus 5 at 63.4% ± 1.35%, but the Artificial Analysis benchmark reveals substantial churn across its 421-model roster, with two new entries disrupting the upper tier: Qwen3.8 Max enters at number 5 with 56.2, pushing Claude Opus 4.8 down one slot to 6, and Muse Spark 1.2 debuts at number 9 with 54.1, while Ling-3.0-flash appears at number 57 with 37.4. The stability in SWE-rebench contrasts sharply with the Artificial Analysis churn and raises a methodological question: SWE-rebench uses confidence intervals (ranging from ± 0.54% to ± 1.83%) suggesting repeated trials or statistical sampling, whereas Artificial Analysis reports single point scores without uncertainty estimates, making it unclear whether the apparent volatility reflects genuine model improvements, evaluation methodology shifts, or simply different statistical rigor between benchmarks. The top-tier SWE-rebench models cluster tightly between 62.3% and 64.5%, overlapping within their confidence bands, which means the ranking order itself carries limited discriminative power at that level. Without historical Artificial Analysis rankings to confirm whether these entries are genuinely new models or data collection artifacts, the apparent movement cannot be cleanly separated from benchmark drift.
Cole Brennan
Daily rankings from SWE-rebench, a benchmark designed to fairly compare LLM capabilities on real-world software engineering tasks. Unlike other evaluations, it uses a standardized scaffolding for all models, continuously updates its dataset to prevent contamination, and runs each model five times to account for stochastic variance.
| # | Model | Score |
|---|---|---|
| 1 | AnthropicFable 5 [high]Model | 64.5%± 1.41% |
| 2 | GrokGrok 4.5 [high]Model | 63.8%± 0.60% |
| 3 | AnthropicOpus 5 [high]Model | 63.4%± 1.35% |
| 4 | Z.aiGLM-5.2 [high]Model | 62.9%± 1.19% |
| 5 | OpenAIGPT-5.6 Sol [medium]Model | 62.3%± 1.83% |
| 6 | JunieJunieAgent | 61.8%± 0.54% |
| 7 | AnthropicClaude CodeAgent | 60.4%± 1.03% |
| 8 | OpenAICodexAgent | 58.0%± 1.29% |
| 9 | AnthropicSonnet 5 [high]Model | 56.8%± 0.94% |
| 10 | CursorCursorAgent | 51.7%± 0.84% |
Artificial Analysis composite index across coding, math, and reasoning benchmarks.
| # | Model | Score | tok/s | $/1M |
|---|---|---|---|---|
| 1 | Claude Opus 5 | 60.7 | 58 | $10.00 |
| 2 | Claude Fable 5 | 59.9 | 71 | $20.00 |
| 3 | GPT-5.6 Sol | 58.9 | 70 | $11.25 |
| 4 | Kimi K3 | 57.1 | 37 | $6.00 |
| 5 | Qwen3.8 Max | 56.2 | 53 | $3.00 |
| 6 | Claude Opus 4.8 | 55.7 | 0 | $10.00 |
| 7 | GPT-5.6 Terra | 55 | 127 | $4.50 |
| 8 | GPT-5.5 | 54.8 | 0 | $11.25 |
| 9 | Muse Spark 1.2 | 54.1 | 0 | $2.00 |
| 10 | Grok 4.5 | 53.8 | 59 | $3.00 |
Output tokens per second — higher is faster. Minimum intelligence score of 40.
| # | Model | tok/s |
|---|---|---|
| 1 | Gemini 3.5 Flash | 255 |
| 2 | Gemini 3.6 Flash | 236 |
| 3 | Qwen3.7 Max | 201 |
| 4 | Muse Spark 1.1 | 192 |
| 5 | GPT-5.6 Luna | 166 |
| 6 | GLM-5.2 | 160 |
| 7 | Nex-N2-Pro | 138 |
| 8 | Gemini 3.1 Pro Preview | 134 |
| 9 | GPT-5.6 Terra | 127 |
| 10 | Inkling Small | 118 |
Blended cost per 1M tokens (3:1 input/output) — lower is cheaper. Minimum intelligence score of 40.
| # | Model | $/1M |
|---|---|---|
| 1 | DeepSeek V4 Flash | $0.171 |
| 2 | DeepSeek V4 Flash 0731 | $0.175 |
| 3 | Hy3 | $0.241 |
| 4 | GPT-5.6 Luna | $0.45 |
| 5 | MiniMax-M3 | $0.525 |
| 6 | Inkling Small | $0.525 |
| 7 | DeepSeek V4 Pro | $0.544 |
| 8 | MiMo-V2.5-Pro | $0.544 |
| 9 | Nex-N2-Pro | $1.00 |
| 10 | GPT-5.4 mini | $1.69 |