The SWE-rebench rankings show no movement in the top tier: AnthropicFable 5 holds first at 64.5% ± 1.41%, GrokGrok 4.5 remains second at 63.8% ± 0.60%, and the next five positions remain unchanged through OpenAIGPT-5.6 Sol at 62.3% ± 1.83%. The Artificial Analysis benchmark, by contrast, underwent substantial reorganization across its full 467-model roster, with NVIDIA Nemotron 3 Nano 4B dropping from position 321 to 465, and multiple models in the 320s range shifting position by single or double digits. The SWE-rebench evaluation tests code generation agents on real GitHub issues with controlled execution environments, while Artificial Analysis appears to assess a broader range of model capabilities across a much larger set of entries, making direct comparison difficult. The stability of the SWE-rebench top tier suggests either that these models have reached a performance plateau on that particular task or that the evaluation's variance bands (ranging from ±0.54% to ±1.83%) are narrow enough to prevent ranking churn. The extensive reshuffling in Artificial Analysis, which uses point scores without reported confidence intervals, may reflect either more sensitive discrimination between models or differences in evaluation methodology. Without access to the specific problem distributions, success criteria, or statistical methodology behind each benchmark, it remains unclear whether the SWE-rebench stability reflects genuine performance plateaus or whether the benchmark's design constrains differentiation among top performers.
Cole Brennan
Daily rankings from SWE-rebench, a benchmark designed to fairly compare LLM capabilities on real-world software engineering tasks. Unlike other evaluations, it uses a standardized scaffolding for all models, continuously updates its dataset to prevent contamination, and runs each model five times to account for stochastic variance.
| # | Model | Score |
|---|---|---|
| 1 | AnthropicFable 5 [high]Model | 64.5%± 1.41% |
| 2 | GrokGrok 4.5 [high]Model | 63.8%± 0.60% |
| 3 | AnthropicOpus 5 [high]Model | 63.4%± 1.35% |
| 4 | Z.aiGLM-5.2 [high]Model | 62.9%± 1.19% |
| 5 | OpenAIGPT-5.6 Sol [medium]Model | 62.3%± 1.83% |
| 6 | JunieJunieAgent | 61.8%± 0.54% |
| 7 | AnthropicClaude CodeAgent | 60.4%± 1.03% |
| 8 | OpenAICodexAgent | 58.0%± 1.29% |
| 9 | AnthropicSonnet 5 [high]Model | 56.8%± 0.94% |
| 10 | CursorCursorAgent | 51.7%± 0.84% |
Artificial Analysis composite index across coding, math, and reasoning benchmarks.
| # | Model | Score | tok/s | $/1M |
|---|---|---|---|---|
| 1 | Claude Opus 5.5 | 57.6 | 98 | $8.00 |
| 2 | Claude Sonnet 5.5 | 56 | 138 | $4.00 |
| 3 | Claude Fable 5.1 | 53.4 | 71 | $20.00 |
| 4 | GPT-6 Astra | 52.7 | 56 | $20.00 |
| 5 | Gemini 4 Argon | 52.6 | 0 | $4.00 |
| 6 | GPT-6.1 Sol | 51.8 | 59 | $4.00 |
| 7 | Claude Opus 5 | 50.8 | 0 | $10.00 |
| 8 | Claude Fable 5 | 49.6 | 0 | $20.00 |
| 9 | Muse Spark 1.3 | 48.1 | 157 | $2.00 |
| 10 | GPT-6 Sol | 47.6 | 111 | $4.00 |
Output tokens per second — higher is faster. Minimum intelligence score of 40.
| # | Model | tok/s |
|---|---|---|
| 1 | Gemini 3.8 Flash | 256 |
| 2 | Muse Spark 1.3 | 157 |
| 3 | Claude Sonnet 5.5 | 138 |
| 4 | GPT-6 Sol | 111 |
| 5 | GPT-5.6 Terra | 110 |
| 6 | Claude Opus 5.5 | 98 |
| 7 | Step 5 Preview | 84 |
| 8 | Grok 4.7 | 80 |
| 9 | GLM-5.3 | 75 |
| 10 | Claude Fable 5.1 | 71 |
Blended cost per 1M tokens (3:1 input/output) — lower is cheaper. Minimum intelligence score of 40.
| # | Model | $/1M |
|---|---|---|
| 1 | GLM 5.3 Flash | $0.237 |
| 2 | MiMo-V2.6-Pro | $0.544 |
| 3 | Step 5 Preview | $1.43 |
| 4 | Gemini 3.8 Flash | $1.50 |
| 5 | Muse Spark 1.3 | $2.00 |
| 6 | GLM-5.3 | $2.15 |
| 7 | Grok 4.7 | $3.00 |
| 8 | Qwen3.8 Max | $3.00 |
| 9 | Grok 4.6 | $3.00 |
| 10 | Claude Sonnet 5.5 | $4.00 |