The SWE-rebench leaderboard shows no movement since the previous update, with the top tier locked in place: AnthropicFable 5 holds 64.5% plus or minus 1.41%, GrokGrok 4.5 remains at 63.8% plus or minus 0.60%, and AnthropicOpus 5 stays at 63.4% plus or minus 1.35%. The confidence intervals are tight enough that the ranking reflects real separation, particularly between the first and second positions where the gap exceeds the combined uncertainty bands. On the Artificial Analysis benchmark, movement is cosmetic rather than structural. MiMo-V2.6-Pro enters at position 30 with a score of 46.3, displacing the previous entry downward, but this represents catalog expansion rather than performance shifts among established models. The top performers remain unchanged: Claude Opus 5.5 at 57.6, Claude Fable 5.1 at 53.4, GPT-6 Astra at 52.7. The SWE-rebench methodology tests code generation against real software engineering problems with defined pass rates, making its stability meaningful, while the Artificial Analysis benchmark's evaluation criteria are less transparent from the data provided. The absence of score changes across both benchmarks suggests either genuine plateau at the frontier or evaluation cycles that have not yet captured recent model iterations. Without information about when these measurements were taken or how frequently they update, it remains unclear whether this stasis reflects convergence or lag in assessment.
Cole Brennan
Daily rankings from SWE-rebench, a benchmark designed to fairly compare LLM capabilities on real-world software engineering tasks. Unlike other evaluations, it uses a standardized scaffolding for all models, continuously updates its dataset to prevent contamination, and runs each model five times to account for stochastic variance.
| # | Model | Score |
|---|---|---|
| 1 | AnthropicFable 5 [high]Model | 64.5%± 1.41% |
| 2 | GrokGrok 4.5 [high]Model | 63.8%± 0.60% |
| 3 | AnthropicOpus 5 [high]Model | 63.4%± 1.35% |
| 4 | Z.aiGLM-5.2 [high]Model | 62.9%± 1.19% |
| 5 | OpenAIGPT-5.6 Sol [medium]Model | 62.3%± 1.83% |
| 6 | JunieJunieAgent | 61.8%± 0.54% |
| 7 | AnthropicClaude CodeAgent | 60.4%± 1.03% |
| 8 | OpenAICodexAgent | 58.0%± 1.29% |
| 9 | AnthropicSonnet 5 [high]Model | 56.8%± 0.94% |
| 10 | CursorCursorAgent | 51.7%± 0.84% |
Artificial Analysis composite index across coding, math, and reasoning benchmarks.
| # | Model | Score | tok/s | $/1M |
|---|---|---|---|---|
| 1 | Claude Opus 5.5 | 57.6 | 99 | $8.00 |
| 2 | Claude Fable 5.1 | 53.4 | 71 | $20.00 |
| 3 | GPT-6 Astra | 52.7 | 60 | $20.00 |
| 4 | Claude Opus 5 | 50.8 | 0 | $10.00 |
| 5 | Claude Fable 5 | 49.6 | 0 | $20.00 |
| 6 | Muse Spark 1.3 | 48.1 | 161 | $2.00 |
| 7 | GPT-6 Sol | 47.5 | 87 | $4.00 |
| 8 | GPT-5.6 Sol | 47 | 0 | $8.00 |
| 9 | Grok 4.7 | 46.4 | 71 | $3.00 |
| 10 | MiMo-V2.6-Pro | 46.3 | 39 | $0.544 |
Output tokens per second — higher is faster. Minimum intelligence score of 40.
| # | Model | tok/s |
|---|---|---|
| 1 | Gemini 3.8 Flash | 298 |
| 2 | Muse Spark 1.3 | 161 |
| 3 | Claude Opus 5.5 | 99 |
| 4 | GPT-5.6 Terra | 90 |
| 5 | GPT-6 Sol | 87 |
| 6 | Grok 4.6 | 79 |
| 7 | GLM-5.3 | 78 |
| 8 | Step 5 Preview | 74 |
| 9 | Claude Fable 5.1 | 71 |
| 10 | Grok 4.7 | 71 |
Blended cost per 1M tokens (3:1 input/output) — lower is cheaper. Minimum intelligence score of 40.
| # | Model | $/1M |
|---|---|---|
| 1 | GLM 5.3 Flash | $0.237 |
| 2 | MiMo-V2.6-Pro | $0.544 |
| 3 | Step 5 Preview | $1.43 |
| 4 | Gemini 3.8 Flash | $1.50 |
| 5 | Muse Spark 1.3 | $2.00 |
| 6 | GLM-5.3 | $2.15 |
| 7 | Grok 4.7 | $3.00 |
| 8 | Qwen3.8 Max | $3.00 |
| 9 | Grok 4.6 | $3.00 |
| 10 | GPT-6 Sol | $4.00 |