The SWE-rebench rankings remain static across both snapshots, with AnthropicFable 5 holding first place at 64.5% and the top seventeen models unchanged in position and score. The Artificial Analysis benchmark, by contrast, shows wholesale reshuffling: Grok 4.7 and MiMo-V2.6-Pro enter the top ten as new entries at positions 7 and 8, displacing Qwen3.8 Max and GLM-5.3 downward by two spots each, while every other model in the 300-entry list shifts position despite most scores remaining identical to the previous cycle. This pattern suggests the Artificial Analysis ranking operates on a different methodology than SWE-rebench, likely incorporating recency weighting, model release date, or other temporal factors that cause constant reordering without corresponding score changes, whereas SWE-rebench appears to measure against a fixed problem set with genuine performance stability at the top. The stability of SWE-rebench scores across the full range, paired with tight confidence intervals (most under 1.5%), indicates controlled experimental conditions and reproducible results; Artificial Analysis's perpetual repositioning of identical-scoring models raises questions about whether the ranking reflects meaningful differentiation or administrative sorting unmoored from the underlying evaluation data. Without visibility into how Artificial Analysis recalculates rankings when scores don't change, treating its movement as a signal of capability shift would be premature.
Cole Brennan
Daily rankings from SWE-rebench, a benchmark designed to fairly compare LLM capabilities on real-world software engineering tasks. Unlike other evaluations, it uses a standardized scaffolding for all models, continuously updates its dataset to prevent contamination, and runs each model five times to account for stochastic variance.
| # | Model | Score |
|---|---|---|
| 1 | AnthropicFable 5 [high]Model | 64.5%± 1.41% |
| 2 | GrokGrok 4.5 [high]Model | 63.8%± 0.60% |
| 3 | AnthropicOpus 5 [high]Model | 63.4%± 1.35% |
| 4 | Z.aiGLM-5.2 [high]Model | 62.9%± 1.19% |
| 5 | OpenAIGPT-5.6 Sol [medium]Model | 62.3%± 1.83% |
| 6 | JunieJunieAgent | 61.8%± 0.54% |
| 7 | AnthropicClaude CodeAgent | 60.4%± 1.03% |
| 8 | OpenAICodexAgent | 58.0%± 1.29% |
| 9 | AnthropicSonnet 5 [high]Model | 56.8%± 0.94% |
| 10 | CursorCursorAgent | 51.7%± 0.84% |
Artificial Analysis composite index across coding, math, and reasoning benchmarks.
| # | Model | Score | tok/s | $/1M |
|---|---|---|---|---|
| 1 | Claude Fable 5.1 | 53.4 | 71 | $20.00 |
| 2 | GPT-6 Astra | 52.7 | 62 | $20.00 |
| 3 | Claude Opus 5 | 50.8 | 57 | $10.00 |
| 4 | Claude Fable 5 | 49.6 | 0 | $20.00 |
| 5 | Muse Spark 1.3 | 48.1 | 281 | $2.00 |
| 6 | GPT-5.6 Sol | 47 | 73 | $8.00 |
| 7 | Grok 4.7 | 46.4 | 44 | $3.00 |
| 8 | MiMo-V2.6-Pro | 46.3 | 116 | $0.544 |
| 9 | Qwen3.8 Max | 45.4 | 42 | $3.00 |
| 10 | GLM-5.3 | 44.8 | 65 | $2.15 |
Output tokens per second — higher is faster. Minimum intelligence score of 40.
| # | Model | tok/s |
|---|---|---|
| 1 | Gemini 3.8 Flash | 365 |
| 2 | Muse Spark 1.3 | 281 |
| 3 | MiMo-V2.6-Pro | 116 |
| 4 | GPT-5.6 Terra | 109 |
| 5 | Step 5 Preview | 93 |
| 6 | GLM 5.3 Flash | 79 |
| 7 | GPT-5.6 Sol | 73 |
| 8 | Claude Fable 5.1 | 71 |
| 9 | Grok 4.6 | 70 |
| 10 | GLM-5.3 | 65 |
Blended cost per 1M tokens (3:1 input/output) — lower is cheaper. Minimum intelligence score of 40.
| # | Model | $/1M |
|---|---|---|
| 1 | GLM 5.3 Flash | $0.237 |
| 2 | MiMo-V2.6-Pro | $0.544 |
| 3 | Step 5 Preview | $1.43 |
| 4 | Gemini 3.8 Flash | $1.50 |
| 5 | Muse Spark 1.3 | $2.00 |
| 6 | GLM-5.3 | $2.15 |
| 7 | Grok 4.7 | $3.00 |
| 8 | Qwen3.8 Max | $3.00 |
| 9 | Grok 4.6 | $3.00 |
| 10 | GPT-5.6 Terra | $4.50 |