The SWE-rebench rankings show no movement in the top tier, with AnthropicFable 5 holding 64.5%, GrokGrok 4.5 at 63.8%, and AnthropicOpus 5 at 63.4%, all within their confidence intervals and statistically indistinguishable from previous measurements. The Artificial Analysis benchmark, which covers a broader set of models, reveals a different picture: Mercury 2.5 enters at position 191 with a score of 12.3, while three Agnes variants (3.0 Flash at 35.5, 2.5 Pro Beta at 35.2, and 2.5 Pro Alpha at 26.8) have dropped entirely from the rankings, suggesting either discontinuation or performance degradation below the threshold. Within the populated rankings, the median movement is negligible, with most models holding their positions or shifting by single digits; the largest visible shifts occur in the lower-ranked models where scoring precision becomes less reliable. The SWE-rebench methodology, which measures coding task completion on real GitHub issues, represents a narrower evaluation of production-grade performance than Artificial Analysis, which appears to sample across a wider range of tasks and model variants. The stability at the top of SWE-rebench suggests that the gap between frontier models has plateaued, while the churn in Artificial Analysis's lower tiers reflects the proliferation of specialized and smaller variants that show inconsistent results across evaluation contexts. Without access to confidence intervals for Artificial Analysis scores, it is difficult to determine whether observed position changes reflect genuine performance shifts or sampling variance.
Cole Brennan
Daily rankings from SWE-rebench, a benchmark designed to fairly compare LLM capabilities on real-world software engineering tasks. Unlike other evaluations, it uses a standardized scaffolding for all models, continuously updates its dataset to prevent contamination, and runs each model five times to account for stochastic variance.
| # | Model | Score |
|---|---|---|
| 1 | AnthropicFable 5 [high]Model | 64.5%± 1.41% |
| 2 | GrokGrok 4.5 [high]Model | 63.8%± 0.60% |
| 3 | AnthropicOpus 5 [high]Model | 63.4%± 1.35% |
| 4 | Z.aiGLM-5.2 [high]Model | 62.9%± 1.19% |
| 5 | OpenAIGPT-5.6 Sol [medium]Model | 62.3%± 1.83% |
| 6 | JunieJunieAgent | 61.8%± 0.54% |
| 7 | AnthropicClaude CodeAgent | 60.4%± 1.03% |
| 8 | OpenAICodexAgent | 58.0%± 1.29% |
| 9 | AnthropicSonnet 5 [high]Model | 56.8%± 0.94% |
| 10 | CursorCursorAgent | 51.7%± 0.84% |
Artificial Analysis composite index across coding, math, and reasoning benchmarks.
| # | Model | Score | tok/s | $/1M |
|---|---|---|---|---|
| 1 | Claude Opus 5.5 | 57.6 | 0 | $8.00 |
| 2 | Claude Fable 5.1 | 53.4 | 67 | $20.00 |
| 3 | GPT-6 Astra | 52.7 | 57 | $20.00 |
| 4 | Claude Opus 5 | 50.8 | 57 | $10.00 |
| 5 | Claude Fable 5 | 49.6 | 0 | $20.00 |
| 6 | Muse Spark 1.3 | 48.1 | 239 | $2.00 |
| 7 | GPT-6 Sol | 47.5 | 116 | $4.00 |
| 8 | GPT-5.6 Sol | 47 | 58 | $8.00 |
| 9 | Grok 4.7 | 46.4 | 44 | $3.00 |
| 10 | MiMo-V2.6-Pro | 46.3 | 48 | $0.544 |
Output tokens per second — higher is faster. Minimum intelligence score of 40.
| # | Model | tok/s |
|---|---|---|
| 1 | Gemini 3.8 Flash | 308 |
| 2 | Muse Spark 1.3 | 239 |
| 3 | GPT-6 Sol | 116 |
| 4 | GPT-5.6 Terra | 86 |
| 5 | Step 5 Preview | 79 |
| 6 | Grok 4.6 | 68 |
| 7 | Claude Fable 5.1 | 67 |
| 8 | GLM-5.3 | 64 |
| 9 | GPT-5.6 Sol | 58 |
| 10 | GPT-6 Astra | 57 |
Blended cost per 1M tokens (3:1 input/output) — lower is cheaper. Minimum intelligence score of 40.
| # | Model | $/1M |
|---|---|---|
| 1 | GLM 5.3 Flash | $0.237 |
| 2 | MiMo-V2.6-Pro | $0.544 |
| 3 | Step 5 Preview | $1.43 |
| 4 | Gemini 3.8 Flash | $1.50 |
| 5 | Muse Spark 1.3 | $2.00 |
| 6 | GLM-5.3 | $2.15 |
| 7 | Grok 4.7 | $3.00 |
| 8 | Qwen3.8 Max | $3.00 |
| 9 | Grok 4.6 | $3.00 |
| 10 | GPT-6 Sol | $4.00 |