The SWE-rebench standings show minimal movement at the top, with AnthropicFable 5 holding the lead at 64.5% (±1.41%) and the top six models clustered within 2.5 percentage points, but the Artificial Analysis benchmark reveals a sharp, systematic decline across the entire leaderboard that defies the stability suggested by SWE-rebench alone. Claude Fable 5.1 dropped from 65.7 to 56.8, a 13% relative decline, while GPT-6 Astra fell from 61.2 to 54.7, and Claude Opus 5 from 63.1 to 54.1, suggesting either a methodological shift in Artificial Analysis's evaluation protocol or a recalibration of its test set that hit all models uniformly rather than differentially. The pattern is not random: every model in the Artificial Analysis top 100 lost between 8 and 15 percentage points, with the median drop around 11 points, while lower-ranked models (those below rank 200) showed smaller absolute losses, implying the benchmark either tightened its criteria, introduced harder test cases, or corrected for prior inflation. The SWE-rebench data, by contrast, appears internally consistent with its previous iteration, suggesting these are genuinely separate evaluation regimes measuring different aspects of code generation capability. Without access to Artificial Analysis's methodology change documentation, the most parsimonious explanation is that one benchmark recalibrated while the other did not, making cross-benchmark comparisons unreliable and highlighting the risk of relying on a single leaderboard source for model assessment.
Cole Brennan
Daily rankings from SWE-rebench, a benchmark designed to fairly compare LLM capabilities on real-world software engineering tasks. Unlike other evaluations, it uses a standardized scaffolding for all models, continuously updates its dataset to prevent contamination, and runs each model five times to account for stochastic variance.
| # | Model | Score |
|---|---|---|
| 1 | AnthropicFable 5 [high]Model | 64.5%± 1.41% |
| 2 | GrokGrok 4.5 [high]Model | 63.8%± 0.60% |
| 3 | AnthropicOpus 5 [high]Model | 63.4%± 1.35% |
| 4 | Z.aiGLM-5.2 [high]Model | 62.9%± 1.19% |
| 5 | OpenAIGPT-5.6 Sol [medium]Model | 62.3%± 1.83% |
| 6 | JunieJunieAgent | 61.8%± 0.54% |
| 7 | AnthropicClaude CodeAgent | 60.4%± 1.03% |
| 8 | OpenAICodexAgent | 58.0%± 1.29% |
| 9 | AnthropicSonnet 5 [high]Model | 56.8%± 0.94% |
| 10 | CursorCursorAgent | 51.7%± 0.84% |
Artificial Analysis composite index across coding, math, and reasoning benchmarks.
| # | Model | Score | tok/s | $/1M |
|---|---|---|---|---|
| 1 | Claude Fable 5.1 | 56.8 | 71 | $20.00 |
| 2 | GPT-6 Astra | 54.7 | 0 | $20.00 |
| 3 | Claude Opus 5 | 54.1 | 52 | $10.00 |
| 4 | Claude Fable 5 | 53.2 | 66 | $20.00 |
| 5 | Muse Spark 1.3 | 52.7 | 177 | $2.00 |
| 6 | GPT-5.6 Sol | 51.3 | 79 | $8.00 |
| 7 | Grok 4.6 | 50.6 | 60 | $3.00 |
| 8 | Kimi K3 | 50.2 | 38 | $6.00 |
| 9 | GLM-5.3 | 48.6 | 78 | $2.15 |
| 10 | Gemini 3.8 Flash | 47.1 | 429 | $1.50 |
Output tokens per second — higher is faster. Minimum intelligence score of 40.
| # | Model | tok/s |
|---|---|---|
| 1 | Gemini 3.8 Flash | 429 |
| 2 | Gemini 3.7 Flash | 295 |
| 3 | Muse Spark 1.2 | 226 |
| 4 | Gemini 3.6 Flash | 207 |
| 5 | Muse Spark 1.3 | 177 |
| 6 | DeepSeek V4 Flash 0731 | 130 |
| 7 | DeepSeek V4 Flash Vision | 118 |
| 8 | GPT-5.6 Luna | 110 |
| 9 | GPT-5.6 Terra | 104 |
| 10 | Claude Sonnet 5 | 81 |
Blended cost per 1M tokens (3:1 input/output) — lower is cheaper. Minimum intelligence score of 40.
| # | Model | $/1M |
|---|---|---|
| 1 | Qwen3.8-Flash-Next | $0.23 |
| 2 | GLM-5.3-Flash | $0.237 |
| 3 | GPT-5.6 Luna | $0.45 |
| 4 | DeepSeek V4 Flash Vision | $0.66 |
| 5 | DeepSeek V4 Flash 0731 | $0.66 |
| 6 | Qwen3.8 27B | $1.13 |
| 7 | Gemini 3.8 Flash | $1.50 |
| 8 | Gemini 3.7 Flash | $1.50 |
| 9 | Gemini 3.6 Flash | $1.50 |
| 10 | DeepSeek V4 Pro 0813 | $1.98 |