The SWE-rebench rankings show no movement from the previous cycle, with AnthropicFable 5 holding the top position at 64.5% (±1.41%), followed by GrokGrok 4.5 at 63.8% (±0.60%) and AnthropicOpus 5 at 63.4% (±1.35%). The confidence intervals are wide enough that several adjacent pairs could plausibly swap on different test runs, yet the top tier remains stable across both measurement periods. By contrast, the Artificial Analysis benchmark presents a different ordering entirely, with Claude Opus 5 ranking first at 60.7 rather than second, and Claude Fable 5 at 59.9 rather than first, suggesting the two benchmarks are measuring distinct performance dimensions or using fundamentally different evaluation protocols. The SWE-rebench methodology appears to isolate specific coding task behaviors that differ from the broader capability assessment in Artificial Analysis, but without documentation of task construction, test set composition, or inter-rater reliability for either benchmark, it remains unclear whether the divergence reflects genuine performance variation across problem types or systematic differences in how each benchmark weights model capabilities. The lack of movement in SWE-rebench rankings over this period could indicate either stable model performance on this specific task distribution or insufficient statistical power to detect real changes given the uncertainty bands, particularly for models in the 40-60% range where error margins approach or exceed the gaps between adjacent entries.
Cole Brennan
Daily rankings from SWE-rebench, a benchmark designed to fairly compare LLM capabilities on real-world software engineering tasks. Unlike other evaluations, it uses a standardized scaffolding for all models, continuously updates its dataset to prevent contamination, and runs each model five times to account for stochastic variance.
| # | Model | Score |
|---|---|---|
| 1 | AnthropicFable 5 [high]Model | 64.5%± 1.41% |
| 2 | GrokGrok 4.5 [high]Model | 63.8%± 0.60% |
| 3 | AnthropicOpus 5 [high]Model | 63.4%± 1.35% |
| 4 | Z.aiGLM-5.2 [high]Model | 62.9%± 1.19% |
| 5 | OpenAIGPT-5.6 Sol [medium]Model | 62.3%± 1.83% |
| 6 | JunieJunieAgent | 61.8%± 0.54% |
| 7 | AnthropicClaude CodeAgent | 60.4%± 1.03% |
| 8 | OpenAICodexAgent | 58.0%± 1.29% |
| 9 | AnthropicSonnet 5 [high]Model | 56.8%± 0.94% |
| 10 | CursorCursorAgent | 51.7%± 0.84% |
Artificial Analysis composite index across coding, math, and reasoning benchmarks.
| # | Model | Score | tok/s | $/1M |
|---|---|---|---|---|
| 1 | Claude Opus 5 | 60.7 | 62 | $10.00 |
| 2 | Claude Fable 5 | 59.9 | 71 | $20.00 |
| 3 | GPT-5.6 Sol | 58.9 | 76 | $11.25 |
| 4 | Kimi K3 | 57.1 | 36 | $6.00 |
| 5 | Claude Opus 4.8 | 55.7 | 0 | $10.00 |
| 6 | GPT-5.6 Terra | 55 | 149 | $4.50 |
| 7 | GPT-5.5 | 54.8 | 0 | $11.25 |
| 8 | Grok 4.5 | 53.8 | 67 | $3.00 |
| 9 | Claude Opus 4.7 | 53.5 | 0 | $10.00 |
| 10 | Claude Sonnet 5 | 53.4 | 97 | $4.00 |
Output tokens per second — higher is faster. Minimum intelligence score of 40.
| # | Model | tok/s |
|---|---|---|
| 1 | Gemini 3.5 Flash | 288 |
| 2 | Gemini 3.6 Flash | 236 |
| 3 | Muse Spark 1.1 | 217 |
| 4 | Qwen3.7 Max | 213 |
| 5 | GLM-5.2 | 182 |
| 6 | GPT-5.6 Luna | 178 |
| 7 | GPT-5.6 Terra | 149 |
| 8 | Gemini 3.1 Pro Preview | 141 |
| 9 | Nex-N2-Pro | 138 |
| 10 | GPT-5.3 Codex | 127 |
Blended cost per 1M tokens (3:1 input/output) — lower is cheaper. Minimum intelligence score of 40.
| # | Model | $/1M |
|---|---|---|
| 1 | DeepSeek V4 Flash | $0.171 |
| 2 | DeepSeek V4 Flash 0731 | $0.175 |
| 3 | Hy3 | $0.241 |
| 4 | GPT-5.6 Luna | $0.45 |
| 5 | MiniMax-M3 | $0.525 |
| 6 | Inkling Small | $0.525 |
| 7 | DeepSeek V4 Pro | $0.544 |
| 8 | MiMo-V2.5-Pro | $0.544 |
| 9 | Nex-N2-Pro | $1.00 |
| 10 | GPT-5.4 mini | $1.69 |