The two benchmark systems show stability at the top tier but divergence in methodology that warrants scrutiny. On SWE-rebench, the leader remains OpenAI gpt-5.5-2026-04-23-xhighModel at 62.7% ± 0.91%, with positions two through ten unchanged from the prior snapshot: JunieJunieAgent (61.6%), OpenAICodexAgent (60.4%), AnthropicClaude CodeAgent (59.6%), and five others holding their ranks. The confidence intervals are tight enough that these orderings reflect real performance separation, though the 1.1-point gap between first and second suggests the frontier remains contested. Artificial Analysis data tells a different story about which models matter. Claude Opus 5 leads there at 60.7, ahead of Claude Fable 5 (59.9) and GPT-5.6 Sol (58.9), a ranking that clusters Anthropic and OpenAI variants rather than privileging agent-based configurations. The two datasets agree on little: SWE-rebench elevates specialized agent frameworks (JunieAgent, Cursor) to positions two and nine, while Artificial Analysis buries agent variants in the mid-tier or lower. This gap reflects fundamental differences in evaluation scope. SWE-rebench appears narrowly calibrated to software engineering tasks with reproducible metrics and error bars; Artificial Analysis spans broader capability assessment and lacks reported confidence bounds, making score movements harder to interpret. Neither system has shifted meaningfully since the previous snapshot, suggesting either genuine stability or evaluation lag. The practical signal is clearest on SWE-rebench's top twenty, where the spread from 62.7% to 38.4% separates systems with meaningfully different engineering performance. Below that, the Artificial Analysis tail extends to models scoring 1.0, a floor that lacks discriminative power and suggests saturation in the evaluation methodology rather than genuine parity.
Cole Brennan
Daily rankings from SWE-rebench, a benchmark designed to fairly compare LLM capabilities on real-world software engineering tasks. Unlike other evaluations, it uses a standardized scaffolding for all models, continuously updates its dataset to prevent contamination, and runs each model five times to account for stochastic variance.
| # | Model | Score |
|---|---|---|
| 1 | OpenAIgpt-5.5-2026-04-23-xhighModel | 62.7%± 0.91% |
| 2 | JunieJunieAgent | 61.6%± 0.64% |
| 3 | OpenAICodexAgent | 60.4%± 1.37% |
| 4 | AnthropicClaude CodeAgent | 59.6%± 1.98% |
| 5 | OpenAIgpt-5.5-2026-04-23-mediumModel | 58.9%± 0.78% |
| 6 | AnthropicClaude Opus 4.8-xhighModel | 56.5%± 1.20% |
| 7 | OpenAIgpt-5.4-2026-03-05-mediumModel | 54.9%± 1.02% |
| 8 | AnthropicClaude Opus 4.7-highModel | 53.1%± 1.45% |
| 9 | CursorCursorAgent | 53.0%± 0.53% |
| 10 | AnthropicClaude Sonnet 4.6Model | 51.3%± 0.55% |
Artificial Analysis composite index across coding, math, and reasoning benchmarks.
| # | Model | Score | tok/s | $/1M |
|---|---|---|---|---|
| 1 | Claude Opus 5 | 60.7 | 63 | $10.00 |
| 2 | Claude Fable 5 | 59.9 | 71 | $20.00 |
| 3 | GPT-5.6 Sol | 58.9 | 90 | $11.25 |
| 4 | Kimi K3 | 57.1 | 33 | $6.00 |
| 5 | Claude Opus 4.8 | 55.7 | 65 | $10.00 |
| 6 | GPT-5.6 Terra | 55 | 165 | $5.63 |
| 7 | GPT-5.5 | 54.8 | 0 | $11.25 |
| 8 | Grok 4.5 | 53.8 | 61 | $3.00 |
| 9 | Claude Opus 4.7 | 53.5 | 0 | $10.00 |
| 10 | Claude Sonnet 5 | 53.4 | 86 | $4.00 |
Output tokens per second — higher is faster. Minimum intelligence score of 40.
| # | Model | tok/s |
|---|---|---|
| 1 | Gemini 3.5 Flash | 266 |
| 2 | Gemini 3.6 Flash | 255 |
| 3 | GPT-5.6 Luna | 220 |
| 4 | GLM-5.2 | 219 |
| 5 | Qwen3.7 Max | 203 |
| 6 | GPT-5.6 Terra | 165 |
| 7 | GPT-5.3 Codex | 148 |
| 8 | Gemini 3.1 Pro Preview | 147 |
| 9 | Nex-N2-Pro | 142 |
| 10 | Muse Spark 1.1 | 137 |
Blended cost per 1M tokens (3:1 input/output) — lower is cheaper. Minimum intelligence score of 40.
| # | Model | $/1M |
|---|---|---|
| 1 | DeepSeek V4 Flash | $0.175 |
| 2 | Hy3 | $0.25 |
| 3 | MiniMax-M3 | $0.525 |
| 4 | DeepSeek V4 Pro | $0.544 |
| 5 | MiMo-V2.5-Pro | $0.544 |
| 6 | Nex-N2-Pro | $1.00 |
| 7 | GPT-5.4 mini | $1.69 |
| 8 | Kimi K2.6 | $1.71 |
| 9 | Kimi K2.7 Code | $1.71 |
| 10 | Muse Spark 1.1 | $2.00 |