The SWE-rebench rankings remain stable at the top tier, with AnthropicFable 5 holding first place at 64.5 percent and the top six entries unchanged, suggesting the coding agent benchmark has reached a plateau where further gains require more than incremental model improvements. On Artificial Analysis, the shifts are more pronounced but confined to the mid-tier: Grok 4.6 enters at number four, displacing Kimi K3 to fifth, while three new entrants appear lower in the rankings (Solar Pro 4 at 40, K-EXAONE 2.0 at 105, and Solar Open2 250B at 66, plus A.X-K2 at 77), indicating the broader evaluation is capturing a wider range of models without disrupting the established hierarchy of top performers. The stability in SWE-rebench contrasts with the churn in Artificial Analysis, which may reflect differences in evaluation scope: the coding benchmark appears to be narrower and more resistant to model shuffling, while the broader analysis benchmark accommodates new releases more readily but with limited impact on the leadership. Neither benchmark shows the kind of discontinuous jump that would indicate a methodological shift or a fundamentally new approach to code generation, only the expected incremental repositioning as variants and new models cycle through.
Cole Brennan
Daily rankings from SWE-rebench, a benchmark designed to fairly compare LLM capabilities on real-world software engineering tasks. Unlike other evaluations, it uses a standardized scaffolding for all models, continuously updates its dataset to prevent contamination, and runs each model five times to account for stochastic variance.
| # | Model | Score |
|---|---|---|
| 1 | AnthropicFable 5 [high]Model | 64.5%± 1.41% |
| 2 | GrokGrok 4.5 [high]Model | 63.8%± 0.60% |
| 3 | AnthropicOpus 5 [high]Model | 63.4%± 1.35% |
| 4 | Z.aiGLM-5.2 [high]Model | 62.9%± 1.19% |
| 5 | OpenAIGPT-5.6 Sol [medium]Model | 62.3%± 1.83% |
| 6 | JunieJunieAgent | 61.8%± 0.54% |
| 7 | AnthropicClaude CodeAgent | 60.4%± 1.03% |
| 8 | OpenAICodexAgent | 58.0%± 1.29% |
| 9 | AnthropicSonnet 5 [high]Model | 56.8%± 0.94% |
| 10 | CursorCursorAgent | 51.7%± 0.84% |
Artificial Analysis composite index across coding, math, and reasoning benchmarks.
| # | Model | Score | tok/s | $/1M |
|---|---|---|---|---|
| 1 | Claude Opus 5 | 63.1 | 53 | $10.00 |
| 2 | Claude Fable 5 | 62.1 | 68 | $20.00 |
| 3 | GPT-5.6 Sol | 60.9 | 65 | $11.25 |
| 4 | Grok 4.6 | 60.9 | 67 | $3.00 |
| 5 | Kimi K3 | 59.7 | 39 | $6.00 |
| 6 | Qwen3.8 Max | 58.1 | 49 | $3.00 |
| 7 | Claude Opus 4.8 | 57.3 | 0 | $10.00 |
| 8 | Muse Spark 1.2 | 56.8 | 0 | $2.00 |
| 9 | GPT-5.6 Terra | 56.6 | 120 | $4.50 |
| 10 | GPT-5.5 | 56.3 | 0 | $11.25 |
Output tokens per second — higher is faster. Minimum intelligence score of 40.
| # | Model | tok/s |
|---|---|---|
| 1 | Gemini 3.6 Flash | 220 |
| 2 | GPT-5.6 Luna | 158 |
| 3 | Nex-N2-Pro | 138 |
| 4 | Inkling Small | 134 |
| 5 | Gemini 3.1 Pro Preview | 127 |
| 6 | GPT-5.6 Terra | 120 |
| 7 | DeepSeek V4 Flash 0731 | 120 |
| 8 | GLM-5.2 | 119 |
| 9 | GPT-5.3 Codex | 119 |
| 10 | MiniMax-M3 | 97 |
Blended cost per 1M tokens (3:1 input/output) — lower is cheaper. Minimum intelligence score of 40.
| # | Model | $/1M |
|---|---|---|
| 1 | DeepSeek V4 Flash | $0.168 |
| 2 | DeepSeek V4 Flash 0731 | $0.175 |
| 3 | Hy3 | $0.241 |
| 4 | GPT-5.6 Luna | $0.45 |
| 5 | MiniMax-M3 | $0.525 |
| 6 | Solar Pro 4 | $0.525 |
| 7 | Inkling Small | $0.525 |
| 8 | DeepSeek V4 Pro 0813 | $0.544 |
| 9 | DeepSeek V4 Pro | $0.544 |
| 10 | MiMo-V2.5-Pro | $0.544 |