The SWE-rebench rankings remain static at the top tier, with AnthropicFable 5 holding 64.5% and the next five models clustered within 2.2 percentage points, all showing confidence intervals that overlap substantially with their neighbors. Artificial Analysis data, by contrast, reveals significant churn across its 470-entry leaderboard, particularly in the middle ranks where Granite 4.2 30B has shifted from position 163 to 193, and several DeepSeek variants have moved multiple positions, though the magnitude of these shifts often falls within typical measurement noise for scores in the 12-15 range. The two benchmarks measure different problem spaces: SWE-rebench evaluates real-world software engineering tasks with controlled conditions and reported error margins, while Artificial Analysis appears to sample a broader cross-section of models with no disclosed methodology or uncertainty quantification. At the SWE-rebench top, the stability reflects either genuine convergence in capability or plateau in discrimination power; the absence of new entries in the top ten positions suggests the frontier has settled. Artificial Analysis's fluidity at ranks 160-200 warrants skepticism about whether those shifts represent meaningful performance changes or noise from unmeasured experimental variance. Neither benchmark clarifies whether the apparent separation between Fable 5 (64.5%) and the field reflects a genuine algorithmic advance or saturation effects on the task distribution itself.
Cole Brennan
Daily rankings from SWE-rebench, a benchmark designed to fairly compare LLM capabilities on real-world software engineering tasks. Unlike other evaluations, it uses a standardized scaffolding for all models, continuously updates its dataset to prevent contamination, and runs each model five times to account for stochastic variance.
| # | Model | Score |
|---|---|---|
| 1 | AnthropicFable 5 [high]Model | 64.5%± 1.41% |
| 2 | GrokGrok 4.5 [high]Model | 63.8%± 0.60% |
| 3 | AnthropicOpus 5 [high]Model | 63.4%± 1.35% |
| 4 | Z.aiGLM-5.2 [high]Model | 62.9%± 1.19% |
| 5 | OpenAIGPT-5.6 Sol [medium]Model | 62.3%± 1.83% |
| 6 | JunieJunieAgent | 61.8%± 0.54% |
| 7 | AnthropicClaude CodeAgent | 60.4%± 1.03% |
| 8 | OpenAICodexAgent | 58.0%± 1.29% |
| 9 | AnthropicSonnet 5 [high]Model | 56.8%± 0.94% |
| 10 | CursorCursorAgent | 51.7%± 0.84% |
Artificial Analysis composite index across coding, math, and reasoning benchmarks.
| # | Model | Score | tok/s | $/1M |
|---|---|---|---|---|
| 1 | Claude Opus 5.5 | 57.6 | 97 | $8.00 |
| 2 | Claude Sonnet 5.5 | 56 | 137 | $4.00 |
| 3 | Claude Fable 5.1 | 53.4 | 72 | $20.00 |
| 4 | GPT-6 Astra | 52.7 | 47 | $20.00 |
| 5 | Gemini 4 Argon | 52.6 | 0 | $4.00 |
| 6 | GPT-6.1 Sol | 51.8 | 58 | $4.00 |
| 7 | Claude Opus 5 | 50.8 | 0 | $10.00 |
| 8 | Claude Fable 5 | 49.6 | 0 | $20.00 |
| 9 | Muse Spark 1.3 | 48.1 | 118 | $2.00 |
| 10 | GPT-6 Sol | 47.6 | 0 | $4.00 |
Output tokens per second — higher is faster. Minimum intelligence score of 40.
| # | Model | tok/s |
|---|---|---|
| 1 | Claude Haiku 5.5 | 235 |
| 2 | Ling 3.1 Flash | 215 |
| 3 | Claude Sonnet 5.5 | 137 |
| 4 | Gemini 3.8 Flash | 131 |
| 5 | Muse Spark 1.3 | 118 |
| 6 | GPT-5.6 Terra | 111 |
| 7 | Claude Opus 5.5 | 97 |
| 8 | Step 5 Preview | 90 |
| 9 | Claude Fable 5.1 | 72 |
| 10 | Grok 4.7 | 70 |
Blended cost per 1M tokens (3:1 input/output) — lower is cheaper. Minimum intelligence score of 40.
| # | Model | $/1M |
|---|---|---|
| 1 | Claude Haiku 5.5 | $0.20 |
| 2 | GLM-5.3-Flash | $0.237 |
| 3 | Ling 3.1 Flash | $0.45 |
| 4 | MiMo-V2.6-Pro | $0.544 |
| 5 | Step 5 Preview | $1.43 |
| 6 | Gemini 3.8 Flash | $1.50 |
| 7 | Muse Spark 1.3 | $2.00 |
| 8 | GLM-5.3 | $2.15 |
| 9 | Grok 4.7 | $3.00 |
| 10 | Qwen3.8 Max | $3.00 |