The SWE-rebench leaderboard shows complete stability across the top tier, with AnthropicFable 5 holding 64.5% ± 1.41%, GrokGrok 4.5 at 63.8% ± 0.60%, and AnthropicOpus 5 at 63.4% ± 1.35%, identical to the previous cycle. The confidence intervals are tight enough that these positions reflect genuine performance differences rather than noise, yet the lack of movement suggests either that the test set has reached saturation among frontier models or that incremental improvements are genuinely stalled. The gap between rank 5 (OpenAIGPT-5.6 Sol at 62.3% ± 1.83%) and rank 14 (DeepSeekDeepSeek-V4 Pro at 40.2% ± 1.29%) is substantial and persistent, indicating a clear separation between high-capability and mid-tier systems. On the Artificial Analysis benchmark, the data shows comprehensive churn throughout the ranking, with Ling 3.1 Flash entering at rank 22 as a new entry, pushing all subsequent models down by one position. This suggests Artificial Analysis tracks a broader ecosystem with more frequent model releases, whereas SWE-rebench's frozen leaderboard may reflect either fewer new submissions or the difficulty of breaking into a plateau dominated by three Anthropic variants. Neither benchmark shows evidence of methodological weakness in the raw numbers, though SWE-rebench's lack of movement warrants scrutiny into whether the test remains sensitive to real capability differences or has become a ceiling effect masking progress in specialized domains.
Cole Brennan
Daily rankings from SWE-rebench, a benchmark designed to fairly compare LLM capabilities on real-world software engineering tasks. Unlike other evaluations, it uses a standardized scaffolding for all models, continuously updates its dataset to prevent contamination, and runs each model five times to account for stochastic variance.
| # | Model | Score |
|---|---|---|
| 1 | AnthropicFable 5 [high]Model | 64.5%± 1.41% |
| 2 | GrokGrok 4.5 [high]Model | 63.8%± 0.60% |
| 3 | AnthropicOpus 5 [high]Model | 63.4%± 1.35% |
| 4 | Z.aiGLM-5.2 [high]Model | 62.9%± 1.19% |
| 5 | OpenAIGPT-5.6 Sol [medium]Model | 62.3%± 1.83% |
| 6 | JunieJunieAgent | 61.8%± 0.54% |
| 7 | AnthropicClaude CodeAgent | 60.4%± 1.03% |
| 8 | OpenAICodexAgent | 58.0%± 1.29% |
| 9 | AnthropicSonnet 5 [high]Model | 56.8%± 0.94% |
| 10 | CursorCursorAgent | 51.7%± 0.84% |
Artificial Analysis composite index across coding, math, and reasoning benchmarks.
| # | Model | Score | tok/s | $/1M |
|---|---|---|---|---|
| 1 | Claude Opus 5.5 | 57.6 | 97 | $8.00 |
| 2 | Claude Sonnet 5.5 | 56 | 138 | $4.00 |
| 3 | Claude Fable 5.1 | 53.4 | 69 | $20.00 |
| 4 | GPT-6 Astra | 52.7 | 63 | $20.00 |
| 5 | Gemini 4 Argon | 52.6 | 0 | $4.00 |
| 6 | GPT-6.1 Sol | 51.8 | 59 | $4.00 |
| 7 | Claude Opus 5 | 50.8 | 0 | $10.00 |
| 8 | Claude Fable 5 | 49.6 | 0 | $20.00 |
| 9 | Muse Spark 1.3 | 48.1 | 162 | $2.00 |
| 10 | GPT-6 Sol | 47.6 | 101 | $4.00 |
Output tokens per second — higher is faster. Minimum intelligence score of 40.
| # | Model | tok/s |
|---|---|---|
| 1 | Gemini 3.8 Flash | 247 |
| 2 | Ling 3.1 Flash | 216 |
| 3 | Muse Spark 1.3 | 162 |
| 4 | Claude Sonnet 5.5 | 138 |
| 5 | GPT-5.6 Terra | 112 |
| 6 | GPT-6 Sol | 101 |
| 7 | Claude Opus 5.5 | 97 |
| 8 | Step 5 Preview | 84 |
| 9 | Grok 4.7 | 80 |
| 10 | GLM-5.3 | 75 |
Blended cost per 1M tokens (3:1 input/output) — lower is cheaper. Minimum intelligence score of 40.
| # | Model | $/1M |
|---|---|---|
| 1 | GLM 5.3 Flash | $0.237 |
| 2 | Ling 3.1 Flash | $0.45 |
| 3 | MiMo-V2.6-Pro | $0.544 |
| 4 | Step 5 Preview | $1.43 |
| 5 | Gemini 3.8 Flash | $1.50 |
| 6 | Muse Spark 1.3 | $2.00 |
| 7 | GLM-5.3 | $2.15 |
| 8 | Grok 4.7 | $3.00 |
| 9 | Qwen3.8 Max | $3.00 |
| 10 | Grok 4.6 | $3.00 |