The SWE-rebench rankings remain frozen at their previous positions, with AnthropicFable 5 holding first place at 64.5% ± 1.41%, followed by GrokGrok 4.5 at 63.8% ± 0.60%, and AnthropicOpus 5 at 63.4% ± 1.35%, suggesting either a measurement plateau or a pause in the release cadence for coding agents. On the Artificial Analysis benchmark, the roster has shifted substantially: Claude Opus 5 now leads at 63.1 (up from second), while Claude Fable 5 dropped to 62.1 (down from first), and a new entry, DeepSeek V4 Flash Vision at 51.5, has inserted itself at position 26, pushing all subsequent models down one slot. The divergence between these two benchmarks is notable, SWE-rebench evaluates coding task completion in a controlled sandbox environment with reproducible test suites, while Artificial Analysis appears to measure general capability across a broader evaluation matrix, which explains why the same models rank differently across the two tests. The stability in SWE-rebench contrasts with the churn in Artificial Analysis, where the top tier remains competitive but entries below position 25 experience consistent downward pressure as newer variants appear. Neither benchmark methodology is disclosed in detail here, but the consistency of SWE-rebench's top performers and their tight clustering (all within 8.3 percentage points) suggests the task distribution may be saturating for frontier models, while Artificial Analysis's wider spread indicates greater differentiation across capability profiles.
Cole Brennan
Daily rankings from SWE-rebench, a benchmark designed to fairly compare LLM capabilities on real-world software engineering tasks. Unlike other evaluations, it uses a standardized scaffolding for all models, continuously updates its dataset to prevent contamination, and runs each model five times to account for stochastic variance.
| # | Model | Score |
|---|---|---|
| 1 | AnthropicFable 5 [high]Model | 64.5%± 1.41% |
| 2 | GrokGrok 4.5 [high]Model | 63.8%± 0.60% |
| 3 | AnthropicOpus 5 [high]Model | 63.4%± 1.35% |
| 4 | Z.aiGLM-5.2 [high]Model | 62.9%± 1.19% |
| 5 | OpenAIGPT-5.6 Sol [medium]Model | 62.3%± 1.83% |
| 6 | JunieJunieAgent | 61.8%± 0.54% |
| 7 | AnthropicClaude CodeAgent | 60.4%± 1.03% |
| 8 | OpenAICodexAgent | 58.0%± 1.29% |
| 9 | AnthropicSonnet 5 [high]Model | 56.8%± 0.94% |
| 10 | CursorCursorAgent | 51.7%± 0.84% |
Artificial Analysis composite index across coding, math, and reasoning benchmarks.
| # | Model | Score | tok/s | $/1M |
|---|---|---|---|---|
| 1 | Claude Opus 5 | 63.1 | 55 | $10.00 |
| 2 | Claude Fable 5 | 62.1 | 64 | $20.00 |
| 3 | GPT-5.6 Sol | 60.9 | 70 | $8.00 |
| 4 | Grok 4.6 | 60.9 | 57 | $3.00 |
| 5 | Kimi K3 | 59.7 | 38 | $6.00 |
| 6 | GLM-5.3 | 59.5 | 80 | $2.15 |
| 7 | Qwen3.8 Max | 58.1 | 24 | $3.00 |
| 8 | Qwen3.8 2.4T A95B | 57.7 | 24 | $3.00 |
| 9 | Claude Opus 4.8 | 57.3 | 0 | $10.00 |
| 10 | Muse Spark 1.2 | 56.8 | 0 | $2.00 |
Output tokens per second — higher is faster. Minimum intelligence score of 40.
| # | Model | tok/s |
|---|---|---|
| 1 | Gemini 3.7 Flash | 345 |
| 2 | Gemini 3.6 Flash | 195 |
| 3 | MiniMax-M3 | 139 |
| 4 | Nex-N2-Pro | 133 |
| 5 | Inkling Small | 131 |
| 6 | GPT-5.6 Luna | 130 |
| 7 | Gemini 3.1 Pro Preview | 123 |
| 8 | DeepSeek V4 Flash 0731 | 122 |
| 9 | GPT-5.3 Codex | 121 |
| 10 | DeepSeek V4 Flash Vision | 120 |
Blended cost per 1M tokens (3:1 input/output) — lower is cheaper. Minimum intelligence score of 40.
| # | Model | $/1M |
|---|---|---|
| 1 | DeepSeek V4 Flash | $0.168 |
| 2 | Hy3 | $0.241 |
| 3 | GPT-5.6 Luna | $0.45 |
| 4 | MiniMax-M3 | $0.525 |
| 5 | Solar Pro 4 | $0.525 |
| 6 | Inkling Small | $0.525 |
| 7 | DeepSeek V4 Pro | $0.544 |
| 8 | MiMo-V2.5-Pro | $0.544 |
| 9 | DeepSeek V4 Flash 0731 | $0.66 |
| 10 | DeepSeek V4 Flash Vision | $0.66 |