The SWE-rebench ranking underwent wholesale turnover, with five new models entering the top tier and five prior leaders removed entirely. AnthropicFable 5 now leads at 64.5% ±1.41%, displacing OpenAI's gpt-5.5-2026-04-23-xhigh (which scored 62.7% ±0.91% previously and has been dropped), while GrokGrok 4.5 and AnthropicOpus 5 occupy the second and third positions at 63.8% and 63.4% respectively. The prior top performer, OpenAI's gpt-5.5-2026-04-23-xhigh, along with four other models designated as medium or xhigh variants, disappeared from the rankings entirely, suggesting either a methodology change, model retirement, or evaluation discontinuation rather than performance degradation. Among models that remained in the benchmark, JunieJunieAgent advanced from #2 to #6 despite a negligible score change (61.6% to 61.8%), while OpenAICodexAgent dropped sharply from #3 to #8 (60.4% down to 58.0%), and CursorCursorAgent fell from #9 to #10 (53.0% to 51.7%). Lower-ranked models show more volatile movement: MiniMax M3 climbed from #17 to #11 (45.6% to 47.2%), and MiMo V2.5 Pro rose from #19 to #12 (42.4% to 46.5%), while Qwen3.6-27B and Qwen3.6-35B-A3B both declined substantially, suggesting either model-specific improvements in the new cohort or changes in test conditions rather than uniform capability shifts. The confidence intervals remain narrow (typically ±0.5% to ±1.8%), indicating stable measurement precision, but the simultaneous departure of five prior leaders and arrival of five new ones raises a methodological question: whether this reflects genuine capability advances or a recalibrated evaluation protocol.
Cole Brennan
Daily rankings from SWE-rebench, a benchmark designed to fairly compare LLM capabilities on real-world software engineering tasks. Unlike other evaluations, it uses a standardized scaffolding for all models, continuously updates its dataset to prevent contamination, and runs each model five times to account for stochastic variance.
| # | Model | Score |
|---|---|---|
| 1 | AnthropicFable 5 [high]Model | 64.5%± 1.41% |
| 2 | GrokGrok 4.5 [high]Model | 63.8%± 0.60% |
| 3 | AnthropicOpus 5 [high]Model | 63.4%± 1.35% |
| 4 | Z.aiGLM-5.2 [high]Model | 62.9%± 1.19% |
| 5 | OpenAIGPT-5.6 Sol [medium]Model | 62.3%± 1.83% |
| 6 | JunieJunieAgent | 61.8%± 0.54% |
| 7 | AnthropicClaude CodeAgent | 60.4%± 1.03% |
| 8 | OpenAICodexAgent | 58.0%± 1.29% |
| 9 | AnthropicSonnet 5 [high]Model | 56.8%± 0.94% |
| 10 | CursorCursorAgent | 51.7%± 0.84% |
Artificial Analysis composite index across coding, math, and reasoning benchmarks.
| # | Model | Score | tok/s | $/1M |
|---|---|---|---|---|
| 1 | Claude Opus 5 | 60.7 | 61 | $10.00 |
| 2 | Claude Fable 5 | 59.9 | 70 | $20.00 |
| 3 | GPT-5.6 Sol | 58.9 | 78 | $11.25 |
| 4 | Kimi K3 | 57.1 | 33 | $6.00 |
| 5 | Claude Opus 4.8 | 55.7 | 67 | $10.00 |
| 6 | GPT-5.6 Terra | 55 | 156 | $5.63 |
| 7 | GPT-5.5 | 54.8 | 0 | $11.25 |
| 8 | Grok 4.5 | 53.8 | 61 | $3.00 |
| 9 | Claude Opus 4.7 | 53.5 | 0 | $10.00 |
| 10 | Claude Sonnet 5 | 53.4 | 91 | $4.00 |
Output tokens per second — higher is faster. Minimum intelligence score of 40.
| # | Model | tok/s |
|---|---|---|
| 1 | Gemini 3.5 Flash | 264 |
| 2 | Gemini 3.6 Flash | 248 |
| 3 | GLM-5.2 | 215 |
| 4 | GPT-5.6 Luna | 205 |
| 5 | Qwen3.7 Max | 202 |
| 6 | GPT-5.6 Terra | 156 |
| 7 | Gemini 3.1 Pro Preview | 146 |
| 8 | Nex-N2-Pro | 142 |
| 9 | Muse Spark 1.1 | 135 |
| 10 | GPT-5.3 Codex | 128 |
Blended cost per 1M tokens (3:1 input/output) — lower is cheaper. Minimum intelligence score of 40.
| # | Model | $/1M |
|---|---|---|
| 1 | DeepSeek V4 Flash | $0.175 |
| 2 | Hy3 | $0.241 |
| 3 | MiniMax-M3 | $0.525 |
| 4 | DeepSeek V4 Pro | $0.544 |
| 5 | MiMo-V2.5-Pro | $0.544 |
| 6 | Nex-N2-Pro | $1.00 |
| 7 | GPT-5.4 mini | $1.69 |
| 8 | Kimi K2.6 | $1.71 |
| 9 | Kimi K2.7 Code | $1.71 |
| 10 | Muse Spark 1.1 | $2.00 |