The SWE-rebench rankings show no movement from the previous cycle, with OpenAI's gpt-5.5-2026-04-23-xhighModel holding the top position at 62.7% ± 0.91%, followed by JunieJunieAgent at 61.6% ± 0.64% and OpenAICodexAgent at 60.4% ± 1.37%. The confidence intervals across the top performers remain wide enough that several models within the top ten overlap in their true performance ranges, particularly Claude CodeAgent (59.6% ± 1.98%) and the medium variant of gpt-5.5 (58.9% ± 0.78%), suggesting that ranking precision at these levels may exceed what the benchmark's variance actually supports. The Artificial Analysis benchmark, by contrast, shows substantial churn in the lower ranks: Nemotron Cascade 2 30B A3B dropped from position 131 to 164, while models in the 130-165 range shuffled significantly, yet the top tier remains stable with Claude Fable 5 at 59.9 and GPT-5.6 Sol at 58.9. The absence of movement in SWE-rebench's top positions across evaluation cycles, combined with overlapping error bars, suggests either genuine plateau in agent performance on this benchmark or that the test set may be approaching saturation for the leading approaches. The Artificial Analysis churn in lower positions reflects typical ranking volatility among closely-scored models rather than meaningful capability shifts.
Cole Brennan
Daily rankings from SWE-rebench, a benchmark designed to fairly compare LLM capabilities on real-world software engineering tasks. Unlike other evaluations, it uses a standardized scaffolding for all models, continuously updates its dataset to prevent contamination, and runs each model five times to account for stochastic variance.
| # | Model | Score |
|---|---|---|
| 1 | OpenAIgpt-5.5-2026-04-23-xhighModel | 62.7%± 0.91% |
| 2 | JunieJunieAgent | 61.6%± 0.64% |
| 3 | OpenAICodexAgent | 60.4%± 1.37% |
| 4 | AnthropicClaude CodeAgent | 59.6%± 1.98% |
| 5 | OpenAIgpt-5.5-2026-04-23-mediumModel | 58.9%± 0.78% |
| 6 | AnthropicClaude Opus 4.8-xhighModel | 56.5%± 1.20% |
| 7 | OpenAIgpt-5.4-2026-03-05-mediumModel | 54.9%± 1.02% |
| 8 | AnthropicClaude Opus 4.7-highModel | 53.1%± 1.45% |
| 9 | CursorCursorAgent | 53.0%± 0.53% |
| 10 | AnthropicClaude Sonnet 4.6Model | 51.3%± 0.55% |
Artificial Analysis composite index across coding, math, and reasoning benchmarks.
| # | Model | Score | tok/s | $/1M |
|---|---|---|---|---|
| 1 | Claude Fable 5 | 59.9 | 69 | $20.00 |
| 2 | GPT-5.6 Sol | 58.9 | 64 | $11.25 |
| 3 | Kimi K3 | 57.1 | 37 | $6.00 |
| 4 | Claude Opus 4.8 | 55.7 | 64 | $10.00 |
| 5 | GPT-5.6 Terra | 55 | 134 | $5.63 |
| 6 | GPT-5.5 | 54.8 | 93 | $11.25 |
| 7 | Grok 4.5 | 53.8 | 64 | $3.00 |
| 8 | Claude Opus 4.7 | 53.5 | 61 | $10.00 |
| 9 | Claude Sonnet 5 | 53.4 | 83 | $4.00 |
| 10 | GPT-5.4 | 51.4 | 149 | $5.63 |
Output tokens per second — higher is faster. Minimum intelligence score of 40.
| # | Model | tok/s |
|---|---|---|
| 1 | Gemini 3.5 Flash | 276 |
| 2 | Gemini 3.6 Flash | 256 |
| 3 | Qwen3.7 Max | 207 |
| 4 | GPT-5.6 Luna | 200 |
| 5 | GPT-5.4 mini | 181 |
| 6 | GLM-5.2 | 167 |
| 7 | GPT-5.2 Codex | 163 |
| 8 | GPT-5.4 | 149 |
| 9 | GPT-5.6 Terra | 134 |
| 10 | Gemini 3.1 Pro Preview | 131 |
Blended cost per 1M tokens (3:1 input/output) — lower is cheaper. Minimum intelligence score of 40.
| # | Model | $/1M |
|---|---|---|
| 1 | DeepSeek V4 Flash | $0.175 |
| 2 | Hy3 | $0.25 |
| 3 | MiniMax-M3 | $0.525 |
| 4 | DeepSeek V4 Pro | $0.544 |
| 5 | MiMo-V2.5-Pro | $0.544 |
| 6 | Nex-N2-Pro | $1.00 |
| 7 | GPT-5.4 mini | $1.69 |
| 8 | Kimi K2.6 | $1.71 |
| 9 | Kimi K2.7 Code | $1.71 |
| 10 | Muse Spark 1.1 | $2.00 |