The SWE-rebench rankings show no movement from the previous cycle, with AnthropicFable 5 maintaining its lead at 64.5 percent and the top seven models holding identical positions and scores. Conversely, the Artificial Analysis benchmark exhibits substantial churn throughout its 468-entry list, with models reordering across the full range: A.X-K2 drops from position 104 to 114, Ling-3.0-flash-Fin climbs from 105 to 104, and dozens of mid-tier entries shuffle positions by single digits. The stability on SWE-rebench suggests either that the test captures a genuine performance plateau among leading models, that confidence intervals overlap sufficiently to mask real differences, or that the benchmark's methodology constrains the signal available for discrimination. The volatility in Artificial Analysis, by contrast, points toward either greater sensitivity in its evaluation protocol or systematic variance across evaluation runs. Without visibility into the specific test conditions, sample sizes, or whether these benchmarks measure overlapping capabilities, the divergence between frozen and fluid rankings resists clean interpretation: SWE-rebench's consistency could reflect robustness or insensitivity, while Artificial Analysis's movement could reflect responsiveness or noise. The two benchmarks appear to be capturing different aspects of model behavior, or operating under different levels of measurement precision, rather than converging on a shared picture of the field.
Cole Brennan
Daily rankings from SWE-rebench, a benchmark designed to fairly compare LLM capabilities on real-world software engineering tasks. Unlike other evaluations, it uses a standardized scaffolding for all models, continuously updates its dataset to prevent contamination, and runs each model five times to account for stochastic variance.
| # | Model | Score |
|---|---|---|
| 1 | AnthropicFable 5 [high]Model | 64.5%± 1.41% |
| 2 | GrokGrok 4.5 [high]Model | 63.8%± 0.60% |
| 3 | AnthropicOpus 5 [high]Model | 63.4%± 1.35% |
| 4 | Z.aiGLM-5.2 [high]Model | 62.9%± 1.19% |
| 5 | OpenAIGPT-5.6 Sol [medium]Model | 62.3%± 1.83% |
| 6 | JunieJunieAgent | 61.8%± 0.54% |
| 7 | AnthropicClaude CodeAgent | 60.4%± 1.03% |
| 8 | OpenAICodexAgent | 58.0%± 1.29% |
| 9 | AnthropicSonnet 5 [high]Model | 56.8%± 0.94% |
| 10 | CursorCursorAgent | 51.7%± 0.84% |
Artificial Analysis composite index across coding, math, and reasoning benchmarks.
| # | Model | Score | tok/s | $/1M |
|---|---|---|---|---|
| 1 | Claude Opus 5.5 | 57.6 | 97 | $8.00 |
| 2 | Claude Sonnet 5.5 | 56 | 137 | $4.00 |
| 3 | Claude Fable 5.1 | 53.4 | 67 | $20.00 |
| 4 | GPT-6 Astra | 52.7 | 64 | $20.00 |
| 5 | Gemini 4 Argon | 52.6 | 0 | $4.00 |
| 6 | GPT-6.1 Sol | 51.8 | 59 | $4.00 |
| 7 | Claude Opus 5 | 50.8 | 0 | $10.00 |
| 8 | Claude Fable 5 | 49.6 | 0 | $20.00 |
| 9 | Muse Spark 1.3 | 48.1 | 148 | $2.00 |
| 10 | GPT-6 Sol | 47.6 | 0 | $4.00 |
Output tokens per second — higher is faster. Minimum intelligence score of 40.
| # | Model | tok/s |
|---|---|---|
| 1 | Ling 3.1 Flash | 216 |
| 2 | Gemini 3.8 Flash | 209 |
| 3 | Muse Spark 1.3 | 148 |
| 4 | Claude Sonnet 5.5 | 137 |
| 5 | GPT-5.6 Terra | 111 |
| 6 | Claude Opus 5.5 | 97 |
| 7 | Step 5 Preview | 90 |
| 8 | GLM-5.3 | 73 |
| 9 | Grok 4.7 | 71 |
| 10 | Claude Fable 5.1 | 67 |
Blended cost per 1M tokens (3:1 input/output) — lower is cheaper. Minimum intelligence score of 40.
| # | Model | $/1M |
|---|---|---|
| 1 | GLM 5.3 Flash | $0.237 |
| 2 | Ling 3.1 Flash | $0.45 |
| 3 | MiMo-V2.6-Pro | $0.544 |
| 4 | Step 5 Preview | $1.43 |
| 5 | Gemini 3.8 Flash | $1.50 |
| 6 | Muse Spark 1.3 | $2.00 |
| 7 | GLM-5.3 | $2.15 |
| 8 | Grok 4.7 | $3.00 |
| 9 | Qwen3.8 Max | $3.00 |
| 10 | Grok 4.6 | $3.00 |