The SWE-rebench standings remain frozen at their previous positions, with AnthropicFable 5 holding 64.5% ± 1.41%, GrokGrok 4.5 at 63.8% ± 0.60%, and AnthropicOpus 5 at 63.4% ± 1.35% across the top three slots, suggesting either a pause in model iteration or an evaluation cycle that hasn't refreshed since the last report. The Artificial Analysis benchmark, by contrast, shows substantial churn throughout its 415-entry roster, with Claude Opus 5 now leading at 60.7 (up from prior position), Claude Fable 5 at 59.9, and GPT-5.6 Sol at 58.9, though the methodology underlying these scores remains opaque and the entries lack confidence intervals that would allow assessment of whether observed ranking shifts exceed noise. A notable addition appears at position 36: Inkling Small enters the Artificial Analysis rankings at 40.2, bumping prior entries down, while Mistral Medium 3.1 enters at 188 with a score of 14.7, suggesting either new model releases or expanded evaluation coverage. The two benchmarks diverge sharply in their top performers, SWE-rebench privileges Anthropic and Grok models at the frontier, while Artificial Analysis splits leadership across Anthropic and OpenAI, a divergence that likely reflects different problem distributions and evaluation rigor rather than a meaningful signal about absolute capability. Without documentation of Artificial Analysis's evaluation protocol, sample sizes, or inter-rater agreement, the ranking movements cannot be distinguished from reordering noise, making the SWE-rebench's stable top tier the more reliable reference point despite its slower refresh cycle.
Cole Brennan
Daily rankings from SWE-rebench, a benchmark designed to fairly compare LLM capabilities on real-world software engineering tasks. Unlike other evaluations, it uses a standardized scaffolding for all models, continuously updates its dataset to prevent contamination, and runs each model five times to account for stochastic variance.
| # | Model | Score |
|---|---|---|
| 1 | AnthropicFable 5 [high]Model | 64.5%± 1.41% |
| 2 | GrokGrok 4.5 [high]Model | 63.8%± 0.60% |
| 3 | AnthropicOpus 5 [high]Model | 63.4%± 1.35% |
| 4 | Z.aiGLM-5.2 [high]Model | 62.9%± 1.19% |
| 5 | OpenAIGPT-5.6 Sol [medium]Model | 62.3%± 1.83% |
| 6 | JunieJunieAgent | 61.8%± 0.54% |
| 7 | AnthropicClaude CodeAgent | 60.4%± 1.03% |
| 8 | OpenAICodexAgent | 58.0%± 1.29% |
| 9 | AnthropicSonnet 5 [high]Model | 56.8%± 0.94% |
| 10 | CursorCursorAgent | 51.7%± 0.84% |
Artificial Analysis composite index across coding, math, and reasoning benchmarks.
| # | Model | Score | tok/s | $/1M |
|---|---|---|---|---|
| 1 | Claude Opus 5 | 60.7 | 55 | $10.00 |
| 2 | Claude Fable 5 | 59.9 | 58 | $20.00 |
| 3 | GPT-5.6 Sol | 58.9 | 67 | $11.25 |
| 4 | Kimi K3 | 57.1 | 32 | $6.00 |
| 5 | Claude Opus 4.8 | 55.7 | 60 | $10.00 |
| 6 | GPT-5.6 Terra | 55 | 135 | $4.50 |
| 7 | GPT-5.5 | 54.8 | 0 | $11.25 |
| 8 | Grok 4.5 | 53.8 | 55 | $3.00 |
| 9 | Claude Opus 4.7 | 53.5 | 0 | $10.00 |
| 10 | Claude Sonnet 5 | 53.4 | 85 | $4.00 |
Output tokens per second — higher is faster. Minimum intelligence score of 40.
| # | Model | tok/s |
|---|---|---|
| 1 | Gemini 3.5 Flash | 220 |
| 2 | Gemini 3.6 Flash | 214 |
| 3 | Qwen3.7 Max | 202 |
| 4 | GPT-5.6 Luna | 187 |
| 5 | Gemini 3.1 Pro Preview | 136 |
| 6 | GPT-5.6 Terra | 135 |
| 7 | Muse Spark 1.1 | 135 |
| 8 | Nex-N2-Pro | 133 |
| 9 | GLM-5.2 | 113 |
| 10 | DeepSeek V4 Flash | 112 |
Blended cost per 1M tokens (3:1 input/output) — lower is cheaper. Minimum intelligence score of 40.
| # | Model | $/1M |
|---|---|---|
| 1 | DeepSeek V4 Flash | $0.175 |
| 2 | Hy3 | $0.241 |
| 3 | GPT-5.6 Luna | $0.45 |
| 4 | MiniMax-M3 | $0.525 |
| 5 | Inkling Small | $0.525 |
| 6 | DeepSeek V4 Pro | $0.544 |
| 7 | MiMo-V2.5-Pro | $0.544 |
| 8 | Nex-N2-Pro | $1.00 |
| 9 | GPT-5.4 mini | $1.69 |
| 10 | Kimi K2.6 | $1.71 |