The SWE-rebench rankings remain stable at the top, with AnthropicFable 5 holding 64.5% and the next five models clustered within 1.5 percentage points, all within their reported confidence intervals. The consistency across both measurement points suggests the gap between frontier code agents has plateaued: Fable 5, Grok 4.5, and Opus 5 occupy positions 1 through 3 with no reordering, and the confidence bands are tight enough that sub-point movements carry little weight. Below rank 10, however, the Artificial Analysis benchmark reveals a different picture. Claude Opus 5 leads that leaderboard at 63.1, beating Claude Fable 5 by a full point despite Fable 5's dominance on SWE-rebench. This divergence suggests the two benchmarks measure different problem classes or evaluation conditions: SWE-rebench may weight integration and repository-level reasoning more heavily, while Artificial Analysis may favor broader reasoning or code understanding. DeepSeek-V4 Pro ranks 14th on SWE-rebench at 40.2% but falls to 30th on Artificial Analysis at 45.3, a reversal that points to distinct test distributions rather than measurement error. The Artificial Analysis list extends to 432 entries, including models scoring 1.0 that likely represent floor effects or incomplete evaluation, whereas SWE-rebench stops at 17 entries with meaningful differentiation. Neither benchmark shows discontinuous jumps or evidence of retesting volatility; the movement is meaningful only insofar as it confirms that code-solving ability measured in controlled repository environments does not correlate perfectly with performance on broader reasoning tasks, and that frontier models have reached a capability ceiling on SWE-rebench that differentiates them from second-tier systems but not from each other.
Cole Brennan
Daily rankings from SWE-rebench, a benchmark designed to fairly compare LLM capabilities on real-world software engineering tasks. Unlike other evaluations, it uses a standardized scaffolding for all models, continuously updates its dataset to prevent contamination, and runs each model five times to account for stochastic variance.
| # | Model | Score |
|---|---|---|
| 1 | AnthropicFable 5 [high]Model | 64.5%± 1.41% |
| 2 | GrokGrok 4.5 [high]Model | 63.8%± 0.60% |
| 3 | AnthropicOpus 5 [high]Model | 63.4%± 1.35% |
| 4 | Z.aiGLM-5.2 [high]Model | 62.9%± 1.19% |
| 5 | OpenAIGPT-5.6 Sol [medium]Model | 62.3%± 1.83% |
| 6 | JunieJunieAgent | 61.8%± 0.54% |
| 7 | AnthropicClaude CodeAgent | 60.4%± 1.03% |
| 8 | OpenAICodexAgent | 58.0%± 1.29% |
| 9 | AnthropicSonnet 5 [high]Model | 56.8%± 0.94% |
| 10 | CursorCursorAgent | 51.7%± 0.84% |
Artificial Analysis composite index across coding, math, and reasoning benchmarks.
| # | Model | Score | tok/s | $/1M |
|---|---|---|---|---|
| 1 | Claude Opus 5 | 63.1 | 49 | $10.00 |
| 2 | Claude Fable 5 | 62.1 | 65 | $20.00 |
| 3 | GPT-5.6 Sol | 60.9 | 65 | $11.25 |
| 4 | Grok 4.6 | 60.9 | 62 | $3.00 |
| 5 | Kimi K3 | 59.7 | 40 | $6.00 |
| 6 | Qwen3.8 Max | 58.1 | 47 | $3.00 |
| 7 | Qwen3.8 2.4T A95B | 57.7 | 47 | $3.00 |
| 8 | Claude Opus 4.8 | 57.3 | 0 | $10.00 |
| 9 | Muse Spark 1.2 | 56.8 | 0 | $2.00 |
| 10 | GPT-5.6 Terra | 56.6 | 109 | $4.50 |
Output tokens per second — higher is faster. Minimum intelligence score of 40.
| # | Model | tok/s |
|---|---|---|
| 1 | Gemini 3.6 Flash | 208 |
| 2 | GPT-5.6 Luna | 166 |
| 3 | GLM-5.2 | 144 |
| 4 | Nex-N2-Pro | 142 |
| 5 | Gemini 3.1 Pro Preview | 133 |
| 6 | GPT-5.3 Codex | 132 |
| 7 | DeepSeek V4 Flash 0731 | 116 |
| 8 | GPT-5.6 Terra | 109 |
| 9 | MiniMax-M3 | 97 |
| 10 | Inkling | 81 |
Blended cost per 1M tokens (3:1 input/output) — lower is cheaper. Minimum intelligence score of 40.
| # | Model | $/1M |
|---|---|---|
| 1 | DeepSeek V4 Flash | $0.168 |
| 2 | Hy3 | $0.241 |
| 3 | GPT-5.6 Luna | $0.45 |
| 4 | MiniMax-M3 | $0.525 |
| 5 | Solar Pro 4 | $0.525 |
| 6 | Inkling Small | $0.525 |
| 7 | DeepSeek V4 Pro | $0.544 |
| 8 | MiMo-V2.5-Pro | $0.544 |
| 9 | DeepSeek V4 Flash 0731 | $0.66 |
| 10 | Nex-N2-Pro | $1.00 |