Claude Fable 5.1 enters the Artificial Analysis leaderboard at the top with 65.7, displacing Claude Opus 5 to second place at 63.1, while the SWE-rebench rankings show no movement at all from the previous cycle. The two benchmarks tell divergent stories: Artificial Analysis reflects a competitive landscape where Anthropic's newest model edges past its predecessor and where the top tier remains tightly clustered within a 5-point spread across the first dozen entries, yet SWE-rebench presents a static snapshot with AnthropicFable 5 holding 64.5% ± 1.41% and the same models occupying the same positions with confidence intervals that overlap substantially across adjacent ranks. The methodological gap between these datasets is worth noting. SWE-rebench measures code generation against repository-scale engineering tasks with explicit success criteria and uncertainty quantification; Artificial Analysis appears to report single-point estimates across a far larger model roster without visible error bounds, making it unclear whether observed differences like the 2.6-point gap between Fable 5.1 and Opus 5 reflect genuine capability separation or evaluation variance. Within SWE-rebench's narrower cohort, the top models cluster so tightly that confidence intervals frequently overlap, which means ranking shifts require either methodological changes or performance moves larger than current uncertainty allows. The appearance of two new entries in Artificial Analysis (Apodex 1.1 at rank 41 and Quasar 438B at rank 43) alongside unchanged SWE-rebench results suggests the two leaderboards operate on different update cycles or evaluation protocols rather than measuring the same phenomenon. Neither benchmark movement is particularly meaningful without understanding whether scores reflect genuine advances or measurement drift.
Cole Brennan
Daily rankings from SWE-rebench, a benchmark designed to fairly compare LLM capabilities on real-world software engineering tasks. Unlike other evaluations, it uses a standardized scaffolding for all models, continuously updates its dataset to prevent contamination, and runs each model five times to account for stochastic variance.
| # | Model | Score |
|---|---|---|
| 1 | AnthropicFable 5 [high]Model | 64.5%± 1.41% |
| 2 | GrokGrok 4.5 [high]Model | 63.8%± 0.60% |
| 3 | AnthropicOpus 5 [high]Model | 63.4%± 1.35% |
| 4 | Z.aiGLM-5.2 [high]Model | 62.9%± 1.19% |
| 5 | OpenAIGPT-5.6 Sol [medium]Model | 62.3%± 1.83% |
| 6 | JunieJunieAgent | 61.8%± 0.54% |
| 7 | AnthropicClaude CodeAgent | 60.4%± 1.03% |
| 8 | OpenAICodexAgent | 58.0%± 1.29% |
| 9 | AnthropicSonnet 5 [high]Model | 56.8%± 0.94% |
| 10 | CursorCursorAgent | 51.7%± 0.84% |
Artificial Analysis composite index across coding, math, and reasoning benchmarks.
| # | Model | Score | tok/s | $/1M |
|---|---|---|---|---|
| 1 | Claude Fable 5.1 | 65.7 | 68 | $20.00 |
| 2 | Claude Opus 5 | 63.1 | 48 | $10.00 |
| 3 | Claude Fable 5 | 62.1 | 59 | $20.00 |
| 4 | GPT-5.6 Sol | 60.9 | 77 | $8.00 |
| 5 | Grok 4.6 | 60.9 | 51 | $3.00 |
| 6 | Kimi K3 | 59.7 | 40 | $6.00 |
| 7 | GLM-5.3 | 59.5 | 73 | $2.15 |
| 8 | Qwen3.8 Max | 58.1 | 40 | $3.00 |
| 9 | Qwen3.8 2.4T A95B | 57.7 | 40 | $3.00 |
| 10 | GLM-5.3-Flash | 57.5 | 42 | $0.237 |
Output tokens per second — higher is faster. Minimum intelligence score of 40.
| # | Model | tok/s |
|---|---|---|
| 1 | Gemini 3.7 Flash | 307 |
| 2 | Quasar 438B | 187 |
| 3 | Gemini 3.6 Flash | 161 |
| 4 | Agnes 2.5 Pro Beta | 152 |
| 5 | Nex-N2-Pro | 135 |
| 6 | GPT-5.6 Luna | 122 |
| 7 | GPT-5.3 Codex | 117 |
| 8 | Gemini 3.1 Pro Preview | 113 |
| 9 | DeepSeek V4 Flash Vision | 111 |
| 10 | MiniMax-M3 | 111 |
Blended cost per 1M tokens (3:1 input/output) — lower is cheaper. Minimum intelligence score of 40.
| # | Model | $/1M |
|---|---|---|
| 1 | Agnes 2.5 Pro Beta | $0.15 |
| 2 | DeepSeek V4 Flash | $0.168 |
| 3 | Qwen3.8-Flash-Next | $0.23 |
| 4 | GLM-5.3-Flash | $0.237 |
| 5 | Hy3 | $0.241 |
| 6 | GPT-5.6 Luna | $0.45 |
| 7 | MiniMax-M3 | $0.525 |
| 8 | Solar Pro 4 | $0.525 |
| 9 | Inkling Small | $0.525 |
| 10 | DeepSeek V4 Pro | $0.544 |