The SWE-rebench leaderboard shows no movement in the top positions, with OpenAI's gpt-5.5-2026-04-23-xhighModel holding rank #1 at 62.7% ± 0.91%, followed by JunieJunieAgent at 61.6% ± 0.64% and OpenAICodexAgent at 60.4% ± 1.37%, the same three-model ordering from the previous cycle. Artificial Analysis, by contrast, registered a single meaningful shift: Claude Opus 5 debuted at #1 with 60.7, displacing Claude Fable 5 from the top spot to #2 at 59.9, a 0.8-point gain that exceeds typical measurement noise given the benchmark's established variance patterns. The SWE-rebench confidence intervals (ranging from 0.45% to 1.98%) suggest sufficient precision to detect real movement, yet the absence of any ranking changes across 24 models indicates either genuine stability in relative performance or that improvements, if present, fall within the bounds of statistical uncertainty. Artificial Analysis's deeper leaderboard shows broader reshuffling throughout the 400-entry list, though the top tier remains concentrated among OpenAI, Anthropic, and emerging Chinese models like Kimi and GLM variants. Without prior Artificial Analysis data to establish trend patterns, it remains unclear whether the Claude Opus 5 promotion reflects sustained capability gains or transient evaluation variance; the SWE-rebench's frozen ranking, however, suggests that whatever changes occurred in the underlying systems did not cross the performance thresholds needed to alter the established hierarchy.
Cole Brennan
Daily rankings from SWE-rebench, a benchmark designed to fairly compare LLM capabilities on real-world software engineering tasks. Unlike other evaluations, it uses a standardized scaffolding for all models, continuously updates its dataset to prevent contamination, and runs each model five times to account for stochastic variance.
| # | Model | Score |
|---|---|---|
| 1 | OpenAIgpt-5.5-2026-04-23-xhighModel | 62.7%± 0.91% |
| 2 | JunieJunieAgent | 61.6%± 0.64% |
| 3 | OpenAICodexAgent | 60.4%± 1.37% |
| 4 | AnthropicClaude CodeAgent | 59.6%± 1.98% |
| 5 | OpenAIgpt-5.5-2026-04-23-mediumModel | 58.9%± 0.78% |
| 6 | AnthropicClaude Opus 4.8-xhighModel | 56.5%± 1.20% |
| 7 | OpenAIgpt-5.4-2026-03-05-mediumModel | 54.9%± 1.02% |
| 8 | AnthropicClaude Opus 4.7-highModel | 53.1%± 1.45% |
| 9 | CursorCursorAgent | 53.0%± 0.53% |
| 10 | AnthropicClaude Sonnet 4.6Model | 51.3%± 0.55% |
Artificial Analysis composite index across coding, math, and reasoning benchmarks.
| # | Model | Score | tok/s | $/1M |
|---|---|---|---|---|
| 1 | Claude Opus 5 | 60.7 | 44 | $10.00 |
| 2 | Claude Fable 5 | 59.9 | 58 | $20.00 |
| 3 | GPT-5.6 Sol | 58.9 | 74 | $11.25 |
| 4 | Kimi K3 | 57.1 | 33 | $6.00 |
| 5 | Claude Opus 4.8 | 55.7 | 63 | $10.00 |
| 6 | GPT-5.6 Terra | 55 | 128 | $5.63 |
| 7 | GPT-5.5 | 54.8 | 0 | $11.25 |
| 8 | Grok 4.5 | 53.8 | 56 | $3.00 |
| 9 | Claude Opus 4.7 | 53.5 | 0 | $10.00 |
| 10 | Claude Sonnet 5 | 53.4 | 83 | $4.00 |
Output tokens per second — higher is faster. Minimum intelligence score of 40.
| # | Model | tok/s |
|---|---|---|
| 1 | Gemini 3.5 Flash | 250 |
| 2 | Gemini 3.6 Flash | 219 |
| 3 | Qwen3.7 Max | 200 |
| 4 | GPT-5.6 Luna | 171 |
| 5 | GLM-5.2 | 157 |
| 6 | Gemini 3.1 Pro Preview | 132 |
| 7 | Nex-N2-Pro | 129 |
| 8 | GPT-5.6 Terra | 128 |
| 9 | GPT-5.3 Codex | 126 |
| 10 | Muse Spark 1.1 | 124 |
Blended cost per 1M tokens (3:1 input/output) — lower is cheaper. Minimum intelligence score of 40.
| # | Model | $/1M |
|---|---|---|
| 1 | DeepSeek V4 Flash | $0.175 |
| 2 | Hy3 | $0.25 |
| 3 | MiniMax-M3 | $0.525 |
| 4 | DeepSeek V4 Pro | $0.544 |
| 5 | MiMo-V2.5-Pro | $0.544 |
| 6 | Nex-N2-Pro | $1.00 |
| 7 | GPT-5.4 mini | $1.69 |
| 8 | Kimi K2.6 | $1.71 |
| 9 | Kimi K2.7 Code | $1.71 |
| 10 | Muse Spark 1.1 | $2.00 |