The SWE-rebench leaderboard shows no movement from the previous snapshot, with OpenAI's gpt-5.5-2026-04-23-xhighModel holding 62.7% (±0.91%) at the top, followed by JunieJunieAgent at 61.6% (±0.64%) and OpenAICodexAgent at 60.4% (±1.37%). The Artificial Analysis benchmark, however, exhibits substantial churn across its 409 entries, with Motif 3 entering at rank 22 (44.1 points) and Mercury 2 dropping from rank 109 to 128 (25.3 to 21.4 points), suggesting either model updates or changes in evaluation methodology that warrant clarification. The lack of overlap between top performers on the two benchmarks, Claude Fable 5 leads Artificial Analysis at 59.9 while ranking nowhere on SWE-rebench's top 24, indicates these measure distinct problem spaces rather than equivalent capabilities. SWE-rebench's narrow confidence intervals (0.45% to 1.98%) suggest tighter experimental control than typical leaderboards, though the static rankings across both snapshots raise questions about evaluation frequency and whether these represent live benchmarks or periodic releases. Without details on SWE-rebench's test set composition, task distribution, or whether agents can call external tools, it remains unclear whether the 1.1-point gap between first and second place reflects genuine capability differences or measurement noise within the confidence bounds.
Cole Brennan
Daily rankings from SWE-rebench, a benchmark designed to fairly compare LLM capabilities on real-world software engineering tasks. Unlike other evaluations, it uses a standardized scaffolding for all models, continuously updates its dataset to prevent contamination, and runs each model five times to account for stochastic variance.
| # | Model | Score |
|---|---|---|
| 1 | OpenAIgpt-5.5-2026-04-23-xhighModel | 62.7%± 0.91% |
| 2 | JunieJunieAgent | 61.6%± 0.64% |
| 3 | OpenAICodexAgent | 60.4%± 1.37% |
| 4 | AnthropicClaude CodeAgent | 59.6%± 1.98% |
| 5 | OpenAIgpt-5.5-2026-04-23-mediumModel | 58.9%± 0.78% |
| 6 | AnthropicClaude Opus 4.8-xhighModel | 56.5%± 1.20% |
| 7 | OpenAIgpt-5.4-2026-03-05-mediumModel | 54.9%± 1.02% |
| 8 | AnthropicClaude Opus 4.7-highModel | 53.1%± 1.45% |
| 9 | CursorCursorAgent | 53.0%± 0.53% |
| 10 | AnthropicClaude Sonnet 4.6Model | 51.3%± 0.55% |
Artificial Analysis composite index across coding, math, and reasoning benchmarks.
| # | Model | Score | tok/s | $/1M |
|---|---|---|---|---|
| 1 | Claude Fable 5 | 59.9 | 70 | $20.00 |
| 2 | GPT-5.6 Sol | 58.9 | 68 | $11.25 |
| 3 | Kimi K3 | 57.1 | 39 | $6.00 |
| 4 | Claude Opus 4.8 | 55.7 | 62 | $10.00 |
| 5 | GPT-5.6 Terra | 55 | 154 | $5.63 |
| 6 | GPT-5.5 | 54.8 | 86 | $11.25 |
| 7 | Grok 4.5 | 53.8 | 75 | $3.00 |
| 8 | Claude Opus 4.7 | 53.5 | 60 | $10.00 |
| 9 | Claude Sonnet 5 | 53.4 | 91 | $4.00 |
| 10 | GPT-5.4 | 51.4 | 163 | $5.63 |
Output tokens per second — higher is faster. Minimum intelligence score of 40.
| # | Model | tok/s |
|---|---|---|
| 1 | Gemini 3.5 Flash | 294 |
| 2 | GPT-5.6 Luna | 212 |
| 3 | Qwen3.7 Max | 208 |
| 4 | GLM-5.2 | 198 |
| 5 | GPT-5.4 mini | 180 |
| 6 | GPT-5.4 | 163 |
| 7 | GPT-5.2 Codex | 163 |
| 8 | GPT-5.6 Terra | 154 |
| 9 | Nex-N2-Pro | 138 |
| 10 | Gemini 3.1 Pro Preview | 137 |
Blended cost per 1M tokens (3:1 input/output) — lower is cheaper. Minimum intelligence score of 40.
| # | Model | $/1M |
|---|---|---|
| 1 | DeepSeek V4 Flash | $0.175 |
| 2 | MiniMax-M3 | $0.525 |
| 3 | DeepSeek V4 Pro | $0.544 |
| 4 | MiMo-V2.5-Pro | $0.544 |
| 5 | Nex-N2-Pro | $1.00 |
| 6 | GPT-5.4 mini | $1.69 |
| 7 | Kimi K2.6 | $1.71 |
| 8 | Kimi K2.7 Code | $1.71 |
| 9 | Muse Spark 1.1 | $2.00 |
| 10 | GLM-5.2 | $2.15 |