The SWE-rebench leaderboard shows no movement in the top positions; AnthropicFable 5 remains at 64.5% plus or minus 1.41%, with GrokGrok 4.5 and AnthropicOpus 5 holding their second and third positions at 63.8% and 63.4% respectively. The confidence intervals are wide enough that these three models occupy a statistical plateau where ranking shifts would require changes well outside current uncertainty bounds. Below the top tier, the spread widens noticeably: Z.aiGLM-5.2 sits at 62.9% plus or minus 1.19%, OpenAIGPT-5.6 Sol at 62.3% plus or minus 1.83%, and JunieJunieAgent at 61.8% plus or minus 0.54%, indicating that agents with tighter confidence intervals (like Junie) may have undergone more stable evaluation conditions than models with larger error margins. The Artificial Analysis benchmark presents a different picture, with Claude Opus 5 leading at 63.1 rather than Fable 5, suggesting the two benchmarks measure or weight code-solving performance differently or that the models respond differently to their evaluation protocols. The SWE-rebench methodology appears to favor Anthropic's Fable variant and Grok 4.5 specifically, while Artificial Analysis ranks Opus higher, a divergence worth noting when interpreting which benchmark better predicts real coding task performance. Neither benchmark shows dramatic velocity in the upper ranks, which either reflects genuine saturation on the problem set or indicates that incremental improvements at this performance ceiling require substantially larger modeling or data investments than the recent release cycle has provided.
Cole Brennan
Daily rankings from SWE-rebench, a benchmark designed to fairly compare LLM capabilities on real-world software engineering tasks. Unlike other evaluations, it uses a standardized scaffolding for all models, continuously updates its dataset to prevent contamination, and runs each model five times to account for stochastic variance.
| # | Model | Score |
|---|---|---|
| 1 | AnthropicFable 5 [high]Model | 64.5%± 1.41% |
| 2 | GrokGrok 4.5 [high]Model | 63.8%± 0.60% |
| 3 | AnthropicOpus 5 [high]Model | 63.4%± 1.35% |
| 4 | Z.aiGLM-5.2 [high]Model | 62.9%± 1.19% |
| 5 | OpenAIGPT-5.6 Sol [medium]Model | 62.3%± 1.83% |
| 6 | JunieJunieAgent | 61.8%± 0.54% |
| 7 | AnthropicClaude CodeAgent | 60.4%± 1.03% |
| 8 | OpenAICodexAgent | 58.0%± 1.29% |
| 9 | AnthropicSonnet 5 [high]Model | 56.8%± 0.94% |
| 10 | CursorCursorAgent | 51.7%± 0.84% |
Artificial Analysis composite index across coding, math, and reasoning benchmarks.
| # | Model | Score | tok/s | $/1M |
|---|---|---|---|---|
| 1 | Claude Opus 5 | 63.1 | 52 | $10.00 |
| 2 | Claude Fable 5 | 62.1 | 63 | $20.00 |
| 3 | GPT-5.6 Sol | 60.9 | 64 | $11.25 |
| 4 | Grok 4.6 | 60.9 | 56 | $3.00 |
| 5 | Kimi K3 | 59.7 | 34 | $6.00 |
| 6 | GLM-5.3 | 59.5 | 0 | $2.15 |
| 7 | Qwen3.8 Max | 58.1 | 43 | $3.00 |
| 8 | Qwen3.8 2.4T A95B | 57.7 | 45 | $3.00 |
| 9 | Claude Opus 4.8 | 57.3 | 0 | $10.00 |
| 10 | Muse Spark 1.2 | 56.8 | 0 | $2.00 |
Output tokens per second — higher is faster. Minimum intelligence score of 40.
| # | Model | tok/s |
|---|---|---|
| 1 | Gemini 3.7 Flash | 327 |
| 2 | Gemini 3.6 Flash | 196 |
| 3 | GPT-5.6 Luna | 139 |
| 4 | Nex-N2-Pro | 136 |
| 5 | Inkling Small | 135 |
| 6 | MiniMax-M3 | 133 |
| 7 | GPT-5.3 Codex | 128 |
| 8 | Gemini 3.1 Pro Preview | 123 |
| 9 | GPT-5.6 Terra | 111 |
| 10 | DeepSeek V4 Flash 0731 | 109 |
Blended cost per 1M tokens (3:1 input/output) — lower is cheaper. Minimum intelligence score of 40.
| # | Model | $/1M |
|---|---|---|
| 1 | DeepSeek V4 Flash | $0.168 |
| 2 | Hy3 | $0.241 |
| 3 | GPT-5.6 Luna | $0.45 |
| 4 | MiniMax-M3 | $0.525 |
| 5 | Solar Pro 4 | $0.525 |
| 6 | Inkling Small | $0.525 |
| 7 | DeepSeek V4 Pro | $0.544 |
| 8 | MiMo-V2.5-Pro | $0.544 |
| 9 | DeepSeek V4 Flash 0731 | $0.66 |
| 10 | Nex-N2-Pro | $1.00 |