The SWE-rebench and Artificial Analysis rankings show stability at the top with minor consolidation below. On SWE-rebench, the top four remain unchanged: OpenAI gpt-5.5-2026-04-23-xhighModel at 62.7%, JunieJunieAgent at 61.6%, OpenAICodexAgent at 60.4%, and AnthropicClaude CodeAgent at 59.6%, with confidence intervals suggesting these gaps are meaningful given the ±0.64% to ±1.98% uncertainty bands. The Artificial Analysis benchmark shows more churn in the 40-180 range, where Agnes 2.5 Pro Alpha enters at position 45 with 38.8 points, pushing prior entries down one slot, and K-EXAONE drops from 113 to 125 (24.7 to 22.1), while G9v3-3B enters at 176 with 16.1 points. Neither benchmark exhibits dramatic shifts at the frontier; the SWE-rebench top performers remain separated by gaps of 0.7 to 1.1 percentage points that exceed most reported error margins, indicating the ranking reflects genuine capability differences rather than noise. Artificial Analysis shows more granular differentiation across its extended tail, but the lack of confidence intervals there limits interpretation of whether movements represent real changes or sampling variation. The stability of the top positions across both benchmarks suggests the leading agent architectures have reached a performance plateau on these tasks, at least relative to the measurement precision available.
Cole Brennan
Daily rankings from SWE-rebench, a benchmark designed to fairly compare LLM capabilities on real-world software engineering tasks. Unlike other evaluations, it uses a standardized scaffolding for all models, continuously updates its dataset to prevent contamination, and runs each model five times to account for stochastic variance.
| # | Model | Score |
|---|---|---|
| 1 | OpenAIgpt-5.5-2026-04-23-xhighModel | 62.7%± 0.91% |
| 2 | JunieJunieAgent | 61.6%± 0.64% |
| 3 | OpenAICodexAgent | 60.4%± 1.37% |
| 4 | AnthropicClaude CodeAgent | 59.6%± 1.98% |
| 5 | OpenAIgpt-5.5-2026-04-23-mediumModel | 58.9%± 0.78% |
| 6 | AnthropicClaude Opus 4.8-xhighModel | 56.5%± 1.20% |
| 7 | OpenAIgpt-5.4-2026-03-05-mediumModel | 54.9%± 1.02% |
| 8 | AnthropicClaude Opus 4.7-highModel | 53.1%± 1.45% |
| 9 | CursorCursorAgent | 53.0%± 0.53% |
| 10 | AnthropicClaude Sonnet 4.6Model | 51.3%± 0.55% |
Artificial Analysis composite index across coding, math, and reasoning benchmarks.
| # | Model | Score | tok/s | $/1M |
|---|---|---|---|---|
| 1 | Claude Fable 5 | 59.9 | 59 | $20.00 |
| 2 | GPT-5.6 Sol | 58.9 | 62 | $11.25 |
| 3 | Kimi K3 | 57.1 | 33 | $6.00 |
| 4 | Claude Opus 4.8 | 55.7 | 61 | $10.00 |
| 5 | GPT-5.6 Terra | 55 | 122 | $5.63 |
| 6 | GPT-5.5 | 54.8 | 0 | $11.25 |
| 7 | Grok 4.5 | 53.8 | 56 | $3.00 |
| 8 | Claude Opus 4.7 | 53.5 | 0 | $10.00 |
| 9 | Claude Sonnet 5 | 53.4 | 81 | $4.00 |
| 10 | GPT-5.4 | 51.4 | 0 | $5.63 |
Output tokens per second — higher is faster. Minimum intelligence score of 40.
| # | Model | tok/s |
|---|---|---|
| 1 | Gemini 3.6 Flash | 255 |
| 2 | Gemini 3.5 Flash | 252 |
| 3 | Qwen3.7 Max | 200 |
| 4 | GPT-5.6 Luna | 170 |
| 5 | GLM-5.2 | 164 |
| 6 | Nex-N2-Pro | 130 |
| 7 | Gemini 3.1 Pro Preview | 128 |
| 8 | Muse Spark 1.1 | 124 |
| 9 | GPT-5.6 Terra | 122 |
| 10 | DeepSeek V4 Flash | 120 |
Blended cost per 1M tokens (3:1 input/output) — lower is cheaper. Minimum intelligence score of 40.
| # | Model | $/1M |
|---|---|---|
| 1 | DeepSeek V4 Flash | $0.175 |
| 2 | Hy3 | $0.25 |
| 3 | MiniMax-M3 | $0.525 |
| 4 | DeepSeek V4 Pro | $0.544 |
| 5 | MiMo-V2.5-Pro | $0.544 |
| 6 | Nex-N2-Pro | $1.00 |
| 7 | GPT-5.4 mini | $1.69 |
| 8 | Kimi K2.6 | $1.71 |
| 9 | Kimi K2.7 Code | $1.71 |
| 10 | Muse Spark 1.1 | $2.00 |