The SWE-rebench rankings hold steady at the top, with AnthropicFable 5 retaining first place at 64.5% and the next five positions unchanged through GPT-5.6 Sol. The confidence intervals remain wide enough that the top tier shows genuine separation: Fable 5's 1.41% margin and Grok 4.5's 0.60% margin indicate different levels of measurement precision, yet both models occupy their positions with confidence bounds that don't overlap meaningfully with adjacent ranks. Below the top six, however, the data becomes noisier. JunieAgent holds sixth at 61.8% with a tight 0.54% margin, but Claude CodeAgent and OpenAICodexAgent follow at 60.4% and 58.0%, suggesting a drop-off in performance that the error bars don't fully resolve. The Artificial Analysis rankings tell a different story. Claude Opus 5 leads there at 60.7, while AnthropicFable 5 ranks second at 59.9, a reversal of the SWE-rebench hierarchy that points to benchmark-specific strengths rather than universal capability differences. GPT-5.6 Sol appears at position three in Artificial Analysis (58.9) but holds fifth on SWE-rebench (62.3%), indicating the two evaluation frameworks weight code-generation and general reasoning tasks differently. The divergence widens further down: DeepSeek-V4 Pro scores 40.2% on SWE-rebench but only 44.3 on Artificial Analysis, placing it 23 positions lower in the latter ranking. These inversions suggest neither benchmark captures the full picture of model capability, and that agentic code performance (measured by SWE-rebench) does not correlate linearly with broader reasoning benchmarks.
Cole Brennan
Daily rankings from SWE-rebench, a benchmark designed to fairly compare LLM capabilities on real-world software engineering tasks. Unlike other evaluations, it uses a standardized scaffolding for all models, continuously updates its dataset to prevent contamination, and runs each model five times to account for stochastic variance.
| # | Model | Score |
|---|---|---|
| 1 | AnthropicFable 5 [high]Model | 64.5%± 1.41% |
| 2 | GrokGrok 4.5 [high]Model | 63.8%± 0.60% |
| 3 | AnthropicOpus 5 [high]Model | 63.4%± 1.35% |
| 4 | Z.aiGLM-5.2 [high]Model | 62.9%± 1.19% |
| 5 | OpenAIGPT-5.6 Sol [medium]Model | 62.3%± 1.83% |
| 6 | JunieJunieAgent | 61.8%± 0.54% |
| 7 | AnthropicClaude CodeAgent | 60.4%± 1.03% |
| 8 | OpenAICodexAgent | 58.0%± 1.29% |
| 9 | AnthropicSonnet 5 [high]Model | 56.8%± 0.94% |
| 10 | CursorCursorAgent | 51.7%± 0.84% |
Artificial Analysis composite index across coding, math, and reasoning benchmarks.
| # | Model | Score | tok/s | $/1M |
|---|---|---|---|---|
| 1 | Claude Opus 5 | 60.7 | 59 | $10.00 |
| 2 | Claude Fable 5 | 59.9 | 71 | $20.00 |
| 3 | GPT-5.6 Sol | 58.9 | 77 | $11.25 |
| 4 | Kimi K3 | 57.1 | 34 | $6.00 |
| 5 | Claude Opus 4.8 | 55.7 | 0 | $10.00 |
| 6 | GPT-5.6 Terra | 55 | 136 | $4.50 |
| 7 | GPT-5.5 | 54.8 | 0 | $11.25 |
| 8 | Grok 4.5 | 53.8 | 56 | $3.00 |
| 9 | Claude Opus 4.7 | 53.5 | 0 | $10.00 |
| 10 | Claude Sonnet 5 | 53.4 | 86 | $4.00 |
Output tokens per second — higher is faster. Minimum intelligence score of 40.
| # | Model | tok/s |
|---|---|---|
| 1 | Gemini 3.5 Flash | 268 |
| 2 | Gemini 3.6 Flash | 230 |
| 3 | Qwen3.7 Max | 204 |
| 4 | GLM-5.2 | 187 |
| 5 | GPT-5.6 Luna | 174 |
| 6 | Muse Spark 1.1 | 172 |
| 7 | GPT-5.6 Terra | 136 |
| 8 | Gemini 3.1 Pro Preview | 129 |
| 9 | GPT-5.3 Codex | 129 |
| 10 | Nex-N2-Pro | 128 |
Blended cost per 1M tokens (3:1 input/output) — lower is cheaper. Minimum intelligence score of 40.
| # | Model | $/1M |
|---|---|---|
| 1 | DeepSeek V4 Flash | $0.171 |
| 2 | DeepSeek V4 Flash 0731 | $0.175 |
| 3 | Hy3 | $0.241 |
| 4 | GPT-5.6 Luna | $0.45 |
| 5 | MiniMax-M3 | $0.525 |
| 6 | Inkling Small | $0.525 |
| 7 | DeepSeek V4 Pro | $0.544 |
| 8 | MiMo-V2.5-Pro | $0.544 |
| 9 | Nex-N2-Pro | $1.00 |
| 10 | GPT-5.4 mini | $1.69 |