The SWE-rebench rankings show no movement in the top tier, with Anthropic's Fable 5 holding at 64.5%, Grok 4.5 at 63.8%, and Opus 5 at 63.4%, all within their confidence intervals and unchanged from the previous cycle. The Artificial Analysis benchmark, by contrast, shows substantial reshuffling: Nemotron 3 Ultra 550B A55B enters at position 128, pushing Gemini 2.5 Pro Preview down one slot, a cascade that ripples through 297 positions below. Claude Opus 5 leads Artificial Analysis at 63.1, a 0.8-point gap from Fable 5 at 62.1, reversing the SWE-rebench ordering and suggesting the benchmarks measure different problem spaces or that the evaluation protocols diverge in their sensitivity to model capability. The SWE-rebench data carries tighter confidence bounds (Grok 4.5 at 0.60%, Junie Agent at 0.54%) compared to the broader spread in Artificial Analysis, indicating more controlled experimental conditions for the coding task, though the lack of movement across 17 entries raises questions about whether the test set has reached saturation or whether the evaluation is insensitive to recent model improvements. Artificial Analysis's dense 424-entry ranking with models clustered in the 1.0 to 3.7 range at the bottom suggests either a different difficulty calibration or inclusion of older or smaller-parameter models that SWE-rebench does not test, making direct comparison between the two benchmarks unreliable for assessing true performance shifts.
Cole Brennan
Daily rankings from SWE-rebench, a benchmark designed to fairly compare LLM capabilities on real-world software engineering tasks. Unlike other evaluations, it uses a standardized scaffolding for all models, continuously updates its dataset to prevent contamination, and runs each model five times to account for stochastic variance.
| # | Model | Score |
|---|---|---|
| 1 | AnthropicFable 5 [high]Model | 64.5%± 1.41% |
| 2 | GrokGrok 4.5 [high]Model | 63.8%± 0.60% |
| 3 | AnthropicOpus 5 [high]Model | 63.4%± 1.35% |
| 4 | Z.aiGLM-5.2 [high]Model | 62.9%± 1.19% |
| 5 | OpenAIGPT-5.6 Sol [medium]Model | 62.3%± 1.83% |
| 6 | JunieJunieAgent | 61.8%± 0.54% |
| 7 | AnthropicClaude CodeAgent | 60.4%± 1.03% |
| 8 | OpenAICodexAgent | 58.0%± 1.29% |
| 9 | AnthropicSonnet 5 [high]Model | 56.8%± 0.94% |
| 10 | CursorCursorAgent | 51.7%± 0.84% |
Artificial Analysis composite index across coding, math, and reasoning benchmarks.
| # | Model | Score | tok/s | $/1M |
|---|---|---|---|---|
| 1 | Claude Opus 5 | 63.1 | 53 | $10.00 |
| 2 | Claude Fable 5 | 62.1 | 66 | $20.00 |
| 3 | GPT-5.6 Sol | 60.9 | 65 | $11.25 |
| 4 | Kimi K3 | 59.7 | 40 | $6.00 |
| 5 | Qwen3.8 Max | 58.1 | 51 | $3.00 |
| 6 | Claude Opus 4.8 | 57.3 | 0 | $10.00 |
| 7 | Muse Spark 1.2 | 56.8 | 0 | $2.00 |
| 8 | GPT-5.6 Terra | 56.6 | 121 | $4.50 |
| 9 | GPT-5.5 | 56.3 | 0 | $11.25 |
| 10 | Grok 4.5 | 55.8 | 50 | $3.00 |
Output tokens per second — higher is faster. Minimum intelligence score of 40.
| # | Model | tok/s |
|---|---|---|
| 1 | Gemini 3.6 Flash | 206 |
| 2 | GPT-5.6 Luna | 164 |
| 3 | Nex-N2-Pro | 144 |
| 4 | Gemini 3.1 Pro Preview | 126 |
| 5 | GLM-5.2 | 125 |
| 6 | Inkling Small | 125 |
| 7 | GPT-5.3 Codex | 124 |
| 8 | GPT-5.6 Terra | 121 |
| 9 | DeepSeek V4 Flash 0731 | 118 |
| 10 | MiniMax-M3 | 96 |
Blended cost per 1M tokens (3:1 input/output) — lower is cheaper. Minimum intelligence score of 40.
| # | Model | $/1M |
|---|---|---|
| 1 | DeepSeek V4 Flash | $0.168 |
| 2 | DeepSeek V4 Flash 0731 | $0.175 |
| 3 | Hy3 | $0.241 |
| 4 | GPT-5.6 Luna | $0.45 |
| 5 | MiniMax-M3 | $0.525 |
| 6 | Inkling Small | $0.525 |
| 7 | DeepSeek V4 Pro | $0.544 |
| 8 | MiMo-V2.5-Pro | $0.544 |
| 9 | Nex-N2-Pro | $1.00 |
| 10 | Qwen3.6 Plus | $1.13 |