The SWE-rebench leaderboard shows no movement at the top nine positions, with AnthropicFable 5 maintaining 64.5% and the tier below it unchanged through OpenAICodexAgent at 58.0%. The consistency across these scores, identical to the previous snapshot, reflects either a stabilization in performance or a lack of new evaluations on this benchmark. The confidence intervals remain tight enough that the ranking would require meaningful score shifts to alter, particularly in the 0.5-1.4 percentage point range where several models cluster. On Artificial Analysis, which measures a broader set of coding tasks, the ordering shows minor churn in the 45-67 position band, with Apodex 1.1 dropping from #47 to #66 on a score decline from 30.4 to 26.4, while models like Claude Sonnet 4.6 and Gemini 3.1 Pro Preview each gained one position. These shifts are modest and do not reflect fundamental reordering of capability tiers. The divergence between benchmarks is notable: models ranked highest on SWE-rebench (Fable 5, Grok 4.5, Opus 5) occupy positions 5, 27, and 1 respectively on Artificial Analysis, suggesting the benchmarks measure different aspects of coding performance or that SWE-rebench may emphasize repository-level software engineering tasks where agentic orchestration matters more than raw instruction-following ability. Without methodological transparency on how SWE-rebench constructs its evaluation, particularly whether it tests tool use, multi-step reasoning, or purely code generation, the stability of its rankings is difficult to interpret as either meaningful stasis or measurement artifact.
Cole Brennan
Daily rankings from SWE-rebench, a benchmark designed to fairly compare LLM capabilities on real-world software engineering tasks. Unlike other evaluations, it uses a standardized scaffolding for all models, continuously updates its dataset to prevent contamination, and runs each model five times to account for stochastic variance.
| # | Model | Score |
|---|---|---|
| 1 | AnthropicFable 5 [high]Model | 64.5%± 1.41% |
| 2 | GrokGrok 4.5 [high]Model | 63.8%± 0.60% |
| 3 | AnthropicOpus 5 [high]Model | 63.4%± 1.35% |
| 4 | Z.aiGLM-5.2 [high]Model | 62.9%± 1.19% |
| 5 | OpenAIGPT-5.6 Sol [medium]Model | 62.3%± 1.83% |
| 6 | JunieJunieAgent | 61.8%± 0.54% |
| 7 | AnthropicClaude CodeAgent | 60.4%± 1.03% |
| 8 | OpenAICodexAgent | 58.0%± 1.29% |
| 9 | AnthropicSonnet 5 [high]Model | 56.8%± 0.94% |
| 10 | CursorCursorAgent | 51.7%± 0.84% |
Artificial Analysis composite index across coding, math, and reasoning benchmarks.
| # | Model | Score | tok/s | $/1M |
|---|---|---|---|---|
| 1 | Claude Opus 5.5 | 57.6 | 102 | $8.00 |
| 2 | Claude Fable 5.1 | 53.4 | 70 | $20.00 |
| 3 | GPT-6 Astra | 52.7 | 54 | $20.00 |
| 4 | Claude Opus 5 | 50.8 | 51 | $10.00 |
| 5 | Claude Fable 5 | 49.6 | 0 | $20.00 |
| 6 | Muse Spark 1.3 | 48.1 | 191 | $2.00 |
| 7 | GPT-6 Sol | 47.5 | 107 | $4.00 |
| 8 | GPT-5.6 Sol | 47 | 58 | $8.00 |
| 9 | Grok 4.7 | 46.4 | 51 | $3.00 |
| 10 | MiMo-V2.6-Pro | 46.3 | 37 | $0.544 |
Output tokens per second — higher is faster. Minimum intelligence score of 40.
| # | Model | tok/s |
|---|---|---|
| 1 | Gemini 3.8 Flash | 267 |
| 2 | Muse Spark 1.3 | 191 |
| 3 | GPT-6 Sol | 107 |
| 4 | Claude Opus 5.5 | 102 |
| 5 | GPT-5.6 Terra | 87 |
| 6 | Step 5 Preview | 73 |
| 7 | Claude Fable 5.1 | 70 |
| 8 | Grok 4.6 | 65 |
| 9 | GLM-5.3 | 64 |
| 10 | GPT-5.6 Sol | 58 |
Blended cost per 1M tokens (3:1 input/output) — lower is cheaper. Minimum intelligence score of 40.
| # | Model | $/1M |
|---|---|---|
| 1 | GLM 5.3 Flash | $0.237 |
| 2 | MiMo-V2.6-Pro | $0.544 |
| 3 | Step 5 Preview | $1.43 |
| 4 | Gemini 3.8 Flash | $1.50 |
| 5 | Muse Spark 1.3 | $2.00 |
| 6 | GLM-5.3 | $2.15 |
| 7 | Grok 4.7 | $3.00 |
| 8 | Qwen3.8 Max | $3.00 |
| 9 | Grok 4.6 | $3.00 |
| 10 | GPT-6 Sol | $4.00 |