The SWE-rebench leaderboard shows no movement since the previous cycle, with AnthropicFable 5 holding the top position at 64.5 percent and the full ranking unchanged across all seventeen entries. On Artificial Analysis, two new entrants appear: GLM-5.3-Flash enters at position nine with a score of 57.5, pushing Claude Opus 4.8 from ninth to tenth, and Granite 4.2 8B debuts at position 170 with 19.6, followed by Granite 4.2 3B at position 217 with 14.3. The stability in SWE-rebench contrasts sharply with the Artificial Analysis benchmark, where the expanded leaderboard now includes 439 models compared to the previous 437. Within the top tier, scores remain tightly clustered and unchanged: Claude Opus 5 leads at 63.1, followed by Claude Fable 5 at 62.1 and GPT-5.6 Sol and Grok 4.6 both at 60.9. The SWE-rebench methodology, which evaluates coding agents on real software engineering tasks with controlled conditions and reproducible metrics, appears sufficiently mature that week-to-week variation falls within confidence intervals marked by standard errors ranging from 0.54 to 1.83 percentage points. The absence of movement in the top positions suggests either that the tested models have reached a performance plateau on this benchmark or that new model releases are not substantially outpacing existing leaders. Artificial Analysis, by contrast, continues to absorb new model submissions, though the criteria for inclusion and the evaluation methodology are less transparent than SWE-rebench's publicly documented task construction.
Cole Brennan
Daily rankings from SWE-rebench, a benchmark designed to fairly compare LLM capabilities on real-world software engineering tasks. Unlike other evaluations, it uses a standardized scaffolding for all models, continuously updates its dataset to prevent contamination, and runs each model five times to account for stochastic variance.
| # | Model | Score |
|---|---|---|
| 1 | AnthropicFable 5 [high]Model | 64.5%± 1.41% |
| 2 | GrokGrok 4.5 [high]Model | 63.8%± 0.60% |
| 3 | AnthropicOpus 5 [high]Model | 63.4%± 1.35% |
| 4 | Z.aiGLM-5.2 [high]Model | 62.9%± 1.19% |
| 5 | OpenAIGPT-5.6 Sol [medium]Model | 62.3%± 1.83% |
| 6 | JunieJunieAgent | 61.8%± 0.54% |
| 7 | AnthropicClaude CodeAgent | 60.4%± 1.03% |
| 8 | OpenAICodexAgent | 58.0%± 1.29% |
| 9 | AnthropicSonnet 5 [high]Model | 56.8%± 0.94% |
| 10 | CursorCursorAgent | 51.7%± 0.84% |
Artificial Analysis composite index across coding, math, and reasoning benchmarks.
| # | Model | Score | tok/s | $/1M |
|---|---|---|---|---|
| 1 | Claude Opus 5 | 63.1 | 55 | $10.00 |
| 2 | Claude Fable 5 | 62.1 | 68 | $20.00 |
| 3 | GPT-5.6 Sol | 60.9 | 70 | $8.00 |
| 4 | Grok 4.6 | 60.9 | 61 | $3.00 |
| 5 | Kimi K3 | 59.7 | 38 | $6.00 |
| 6 | GLM-5.3 | 59.5 | 81 | $2.15 |
| 7 | Qwen3.8 Max | 58.1 | 24 | $3.00 |
| 8 | Qwen3.8 2.4T A95B | 57.7 | 24 | $3.00 |
| 9 | GLM-5.3-Flash | 57.5 | 42 | $0.237 |
| 10 | Claude Opus 4.8 | 57.3 | 0 | $10.00 |
Output tokens per second — higher is faster. Minimum intelligence score of 40.
| # | Model | tok/s |
|---|---|---|
| 1 | Gemini 3.7 Flash | 366 |
| 2 | Gemini 3.6 Flash | 197 |
| 3 | Nex-N2-Pro | 142 |
| 4 | DeepSeek V4 Flash 0731 | 129 |
| 5 | Gemini 3.1 Pro Preview | 125 |
| 6 | GPT-5.6 Luna | 124 |
| 7 | GPT-5.3 Codex | 121 |
| 8 | DeepSeek V4 Flash Vision | 120 |
| 9 | Inkling Small | 115 |
| 10 | GPT-5.6 Terra | 113 |
Blended cost per 1M tokens (3:1 input/output) — lower is cheaper. Minimum intelligence score of 40.
| # | Model | $/1M |
|---|---|---|
| 1 | DeepSeek V4 Flash | $0.168 |
| 2 | GLM-5.3-Flash | $0.237 |
| 3 | Hy3 | $0.241 |
| 4 | GPT-5.6 Luna | $0.45 |
| 5 | MiniMax-M3 | $0.525 |
| 6 | Solar Pro 4 | $0.525 |
| 7 | Inkling Small | $0.525 |
| 8 | DeepSeek V4 Pro | $0.544 |
| 9 | MiMo-V2.5-Pro | $0.544 |
| 10 | DeepSeek V4 Flash 0731 | $0.66 |