The SWE-rebench rankings remain unchanged from the previous snapshot, with OpenAI's gpt-5.5-2026-04-23-xhighModel holding first place at 62.7% ± 0.91%, followed by JunieJunieAgent at 61.6% ± 0.64% and OpenAICodexAgent at 60.4% ± 1.37%. The confidence intervals overlap substantially across the top tier, suggesting the performance gap between positions one through five is within measurement noise. Artificial Analysis data shows broader dispersion across 408 models, with Claude Fable 5 leading at 59.9 and models dropping below 2% in the long tail, but this benchmark uses different evaluation methodology and does not track software engineering task completion in the same way as SWE-rebench. The absence of movement in the SWE-rebench leaderboard across two reporting cycles could indicate either genuine stability in the coding agent landscape or insufficient statistical power to detect meaningful shifts given the modest confidence intervals. Without prior historical snapshots beyond these two identical readings, it remains unclear whether the SWE-rebench results represent a genuine plateau in agent performance or simply reflect the inherent variance of the evaluation itself.
Cole Brennan
Daily rankings from SWE-rebench, a benchmark designed to fairly compare LLM capabilities on real-world software engineering tasks. Unlike other evaluations, it uses a standardized scaffolding for all models, continuously updates its dataset to prevent contamination, and runs each model five times to account for stochastic variance.
| # | Model | Score |
|---|---|---|
| 1 | OpenAIgpt-5.5-2026-04-23-xhighModel | 62.7%± 0.91% |
| 2 | JunieJunieAgent | 61.6%± 0.64% |
| 3 | OpenAICodexAgent | 60.4%± 1.37% |
| 4 | AnthropicClaude CodeAgent | 59.6%± 1.98% |
| 5 | OpenAIgpt-5.5-2026-04-23-mediumModel | 58.9%± 0.78% |
| 6 | AnthropicClaude Opus 4.8-xhighModel | 56.5%± 1.20% |
| 7 | OpenAIgpt-5.4-2026-03-05-mediumModel | 54.9%± 1.02% |
| 8 | AnthropicClaude Opus 4.7-highModel | 53.1%± 1.45% |
| 9 | CursorCursorAgent | 53.0%± 0.53% |
| 10 | AnthropicClaude Sonnet 4.6Model | 51.3%± 0.55% |
Artificial Analysis composite index across coding, math, and reasoning benchmarks.
| # | Model | Score | tok/s | $/1M |
|---|---|---|---|---|
| 1 | Claude Fable 5 | 59.9 | 66 | $20.00 |
| 2 | GPT-5.6 Sol | 58.9 | 71 | $11.25 |
| 3 | Kimi K3 | 57.1 | 0 | $6.00 |
| 4 | Claude Opus 4.8 | 55.7 | 61 | $10.00 |
| 5 | GPT-5.6 Terra | 55 | 140 | $5.63 |
| 6 | GPT-5.5 | 54.8 | 86 | $11.25 |
| 7 | Grok 4.5 | 53.8 | 73 | $3.00 |
| 8 | Claude Opus 4.7 | 53.5 | 53 | $10.00 |
| 9 | Claude Sonnet 5 | 53.4 | 81 | $4.00 |
| 10 | GPT-5.4 | 51.4 | 165 | $5.63 |
Output tokens per second — higher is faster. Minimum intelligence score of 40.
| # | Model | tok/s |
|---|---|---|
| 1 | Gemini 3.5 Flash | 291 |
| 2 | GPT-5.6 Luna | 209 |
| 3 | Qwen3.7 Max | 205 |
| 4 | GLM-5.2 | 185 |
| 5 | GPT-5.4 mini | 178 |
| 6 | GPT-5.4 | 165 |
| 7 | GPT-5.2 Codex | 154 |
| 8 | Nex-N2-Pro | 142 |
| 9 | GPT-5.6 Terra | 140 |
| 10 | Gemini 3.1 Pro Preview | 137 |
Blended cost per 1M tokens (3:1 input/output) — lower is cheaper. Minimum intelligence score of 40.
| # | Model | $/1M |
|---|---|---|
| 1 | DeepSeek V4 Flash | $0.175 |
| 2 | MiniMax-M3 | $0.525 |
| 3 | DeepSeek V4 Pro | $0.544 |
| 4 | MiMo-V2.5-Pro | $0.544 |
| 5 | Nex-N2-Pro | $1.00 |
| 6 | GPT-5.4 mini | $1.69 |
| 7 | Kimi K2.6 | $1.71 |
| 8 | Kimi K2.7 Code | $1.71 |
| 9 | Muse Spark 1.1 | $2.00 |
| 10 | GLM-5.2 | $2.15 |