The SWE-rebench rankings show no movement from the previous cycle, with AnthropicFable 5 holding first place at 64.5 percent and the entire top 17 models maintaining their positions unchanged. Across the Artificial Analysis benchmark, the top tier remains similarly static: Claude Opus 5 leads at 63.1, followed by Claude Fable 5 at 62.1, though these scores differ slightly from the SWE-rebench ordering, suggesting the two benchmarks measure overlapping but distinct capabilities. Within the broader Artificial Analysis leaderboard, a single new entry appears at rank 21, Qwen3.8 27B scoring 52.0, displacing previous models down by one position each through the middle ranks, while the lower tiers below position 300 show only minor reordering with no meaningful score changes. The consistency across both benchmarks indicates either genuine stability in model performance or measurement limitations that prevent detection of incremental gains; the SWE-rebench confidence intervals, ranging from 0.54 to 1.83 percentage points, are wide enough that most apparent differences between adjacent models fall within noise, and without access to the evaluation methodology details, it remains unclear whether the benchmark is sufficiently sensitive to capture real progress or whether the field has genuinely plateaued at this performance level.
Cole Brennan
Daily rankings from SWE-rebench, a benchmark designed to fairly compare LLM capabilities on real-world software engineering tasks. Unlike other evaluations, it uses a standardized scaffolding for all models, continuously updates its dataset to prevent contamination, and runs each model five times to account for stochastic variance.
| # | Model | Score |
|---|---|---|
| 1 | AnthropicFable 5 [high]Model | 64.5%± 1.41% |
| 2 | GrokGrok 4.5 [high]Model | 63.8%± 0.60% |
| 3 | AnthropicOpus 5 [high]Model | 63.4%± 1.35% |
| 4 | Z.aiGLM-5.2 [high]Model | 62.9%± 1.19% |
| 5 | OpenAIGPT-5.6 Sol [medium]Model | 62.3%± 1.83% |
| 6 | JunieJunieAgent | 61.8%± 0.54% |
| 7 | AnthropicClaude CodeAgent | 60.4%± 1.03% |
| 8 | OpenAICodexAgent | 58.0%± 1.29% |
| 9 | AnthropicSonnet 5 [high]Model | 56.8%± 0.94% |
| 10 | CursorCursorAgent | 51.7%± 0.84% |
Artificial Analysis composite index across coding, math, and reasoning benchmarks.
| # | Model | Score | tok/s | $/1M |
|---|---|---|---|---|
| 1 | Claude Opus 5 | 63.1 | 52 | $10.00 |
| 2 | Claude Fable 5 | 62.1 | 67 | $20.00 |
| 3 | GPT-5.6 Sol | 60.9 | 70 | $11.25 |
| 4 | Grok 4.6 | 60.9 | 63 | $3.00 |
| 5 | Kimi K3 | 59.7 | 40 | $6.00 |
| 6 | Qwen3.8 Max | 58.1 | 47 | $3.00 |
| 7 | Qwen3.8 2.4T A95B | 57.7 | 47 | $3.00 |
| 8 | Claude Opus 4.8 | 57.3 | 0 | $10.00 |
| 9 | Muse Spark 1.2 | 56.8 | 0 | $2.00 |
| 10 | GPT-5.6 Terra | 56.6 | 118 | $4.50 |
Output tokens per second — higher is faster. Minimum intelligence score of 40.
| # | Model | tok/s |
|---|---|---|
| 1 | Gemini 3.7 Flash | 297 |
| 2 | Gemini 3.6 Flash | 209 |
| 3 | GPT-5.6 Luna | 169 |
| 4 | GLM-5.2 | 149 |
| 5 | Nex-N2-Pro | 143 |
| 6 | Gemini 3.1 Pro Preview | 136 |
| 7 | GPT-5.3 Codex | 132 |
| 8 | DeepSeek V4 Flash 0731 | 123 |
| 9 | GPT-5.6 Terra | 118 |
| 10 | MiniMax-M3 | 95 |
Blended cost per 1M tokens (3:1 input/output) — lower is cheaper. Minimum intelligence score of 40.
| # | Model | $/1M |
|---|---|---|
| 1 | DeepSeek V4 Flash | $0.168 |
| 2 | Hy3 | $0.241 |
| 3 | GPT-5.6 Luna | $0.45 |
| 4 | MiniMax-M3 | $0.525 |
| 5 | Solar Pro 4 | $0.525 |
| 6 | Inkling Small | $0.525 |
| 7 | DeepSeek V4 Pro | $0.544 |
| 8 | MiMo-V2.5-Pro | $0.544 |
| 9 | DeepSeek V4 Flash 0731 | $0.66 |
| 10 | Nex-N2-Pro | $1.00 |