The SWE-rebench leaderboard holds steady at the top with Anthropic Fable 5 maintaining 64.5% ± 1.41%, followed by Grok 4.5 at 63.8% ± 0.60% and Anthropic Opus 5 at 63.4% ± 1.35%, showing no movement in the top tier. Notably, the confidence intervals on SWE-rebench remain tight across the tested models, suggesting controlled evaluation conditions, though the gap between the leader and tenth-place Cursor Agent (51.7% ± 0.84%) spans 12.8 percentage points, indicating clear stratification in code-solving capability. On the Artificial Analysis benchmark, Gemini 4 Argon enters at number five with 52.6 points as a new entrant, displacing prior entries by one position, while GPT-6 Luna climbed from #34 to #33 with a score increase from 37.3 to 38.1. Solar Mini 4 also appears as a new entry at #91 with 24.1 points. The broader Artificial Analysis leaderboard shows tight clustering in the 5 to 10 point range across positions 380 to 450, where parameter count and model size variations produce minimal score differentiation, raising questions about whether that benchmark's resolution distinguishes meaningfully between smaller models. The SWE-rebench results align with the intuition that coding tasks reward architectural sophistication and training data quality more than raw parameters, whereas Artificial Analysis's crowded lower tier suggests saturation or floor effects in its evaluation methodology.
Cole Brennan
Daily rankings from SWE-rebench, a benchmark designed to fairly compare LLM capabilities on real-world software engineering tasks. Unlike other evaluations, it uses a standardized scaffolding for all models, continuously updates its dataset to prevent contamination, and runs each model five times to account for stochastic variance.
| # | Model | Score |
|---|---|---|
| 1 | AnthropicFable 5 [high]Model | 64.5%± 1.41% |
| 2 | GrokGrok 4.5 [high]Model | 63.8%± 0.60% |
| 3 | AnthropicOpus 5 [high]Model | 63.4%± 1.35% |
| 4 | Z.aiGLM-5.2 [high]Model | 62.9%± 1.19% |
| 5 | OpenAIGPT-5.6 Sol [medium]Model | 62.3%± 1.83% |
| 6 | JunieJunieAgent | 61.8%± 0.54% |
| 7 | AnthropicClaude CodeAgent | 60.4%± 1.03% |
| 8 | OpenAICodexAgent | 58.0%± 1.29% |
| 9 | AnthropicSonnet 5 [high]Model | 56.8%± 0.94% |
| 10 | CursorCursorAgent | 51.7%± 0.84% |
Artificial Analysis composite index across coding, math, and reasoning benchmarks.
| # | Model | Score | tok/s | $/1M |
|---|---|---|---|---|
| 1 | Claude Opus 5.5 | 57.6 | 96 | $8.00 |
| 2 | Claude Sonnet 5.5 | 56 | 145 | $4.00 |
| 3 | Claude Fable 5.1 | 53.4 | 70 | $20.00 |
| 4 | GPT-6 Astra | 52.7 | 54 | $20.00 |
| 5 | Gemini 4 Argon | 52.6 | 0 | $4.00 |
| 6 | GPT-6.1 Sol | 51.8 | 73 | $4.00 |
| 7 | Claude Opus 5 | 50.8 | 0 | $10.00 |
| 8 | Claude Fable 5 | 49.6 | 0 | $20.00 |
| 9 | Muse Spark 1.3 | 48.1 | 177 | $2.00 |
| 10 | GPT-6 Sol | 47.6 | 83 | $4.00 |
Output tokens per second — higher is faster. Minimum intelligence score of 40.
| # | Model | tok/s |
|---|---|---|
| 1 | Gemini 3.8 Flash | 238 |
| 2 | Muse Spark 1.3 | 177 |
| 3 | Claude Sonnet 5.5 | 145 |
| 4 | GPT-5.6 Terra | 106 |
| 5 | Claude Opus 5.5 | 96 |
| 6 | Step 5 Preview | 88 |
| 7 | GPT-6 Sol | 83 |
| 8 | Grok 4.7 | 83 |
| 9 | GPT-6.1 Sol | 73 |
| 10 | Claude Fable 5.1 | 70 |
Blended cost per 1M tokens (3:1 input/output) — lower is cheaper. Minimum intelligence score of 40.
| # | Model | $/1M |
|---|---|---|
| 1 | GLM 5.3 Flash | $0.237 |
| 2 | MiMo-V2.6-Pro | $0.544 |
| 3 | Step 5 Preview | $1.43 |
| 4 | Gemini 3.8 Flash | $1.50 |
| 5 | Muse Spark 1.3 | $2.00 |
| 6 | GLM-5.3 | $2.15 |
| 7 | Grok 4.7 | $3.00 |
| 8 | Qwen3.8 Max | $3.00 |
| 9 | Grok 4.6 | $3.00 |
| 10 | Claude Sonnet 5.5 | $4.00 |