The SWE-rebench rankings remain static across the top tier, with AnthropicFable 5 holding 64.5% and the next four positions unchanged through GPT-5.6 Sol at 62.3%. The stability here reflects narrow confidence intervals (most under 1.8 percentage points), suggesting these measurements have settled into reliable territory. Artificial Analysis, by contrast, shows substantial churn below the top twenty entries, with models like LFM2.5-1.2B-Thinking newly appearing at rank 402 and numerous repositionings throughout the mid-tier that indicate either methodological differences between the two benchmarks or genuine performance variance on different evaluation sets. The gap between SWE-rebench's top performer and its seventeenth-ranked model spans 47.4 percentage points (64.5% to 17.1%), whereas Artificial Analysis compresses the same span from rank 1 to 17 into just 9.9 points (63.1 to 53.2), suggesting SWE-rebench may be testing a narrower or more discriminative problem space. Within Artificial Analysis, the long tail below rank 100 shows models clustering tightly at single-digit scores, making positional shifts there largely noise. The absence of movement in SWE-rebench's top ten over this cycle indicates either that the benchmark has reached saturation at its current scale or that frontier models are converging on its difficulty floor, a pattern worth examining against the benchmark's design rather than interpreting as stagnation in model capability.
Cole Brennan
Daily rankings from SWE-rebench, a benchmark designed to fairly compare LLM capabilities on real-world software engineering tasks. Unlike other evaluations, it uses a standardized scaffolding for all models, continuously updates its dataset to prevent contamination, and runs each model five times to account for stochastic variance.
| # | Model | Score |
|---|---|---|
| 1 | AnthropicFable 5 [high]Model | 64.5%± 1.41% |
| 2 | GrokGrok 4.5 [high]Model | 63.8%± 0.60% |
| 3 | AnthropicOpus 5 [high]Model | 63.4%± 1.35% |
| 4 | Z.aiGLM-5.2 [high]Model | 62.9%± 1.19% |
| 5 | OpenAIGPT-5.6 Sol [medium]Model | 62.3%± 1.83% |
| 6 | JunieJunieAgent | 61.8%± 0.54% |
| 7 | AnthropicClaude CodeAgent | 60.4%± 1.03% |
| 8 | OpenAICodexAgent | 58.0%± 1.29% |
| 9 | AnthropicSonnet 5 [high]Model | 56.8%± 0.94% |
| 10 | CursorCursorAgent | 51.7%± 0.84% |
Artificial Analysis composite index across coding, math, and reasoning benchmarks.
| # | Model | Score | tok/s | $/1M |
|---|---|---|---|---|
| 1 | Claude Opus 5 | 63.1 | 53 | $10.00 |
| 2 | Claude Fable 5 | 62.1 | 67 | $20.00 |
| 3 | GPT-5.6 Sol | 60.9 | 78 | $11.25 |
| 4 | Grok 4.6 | 60.9 | 58 | $3.00 |
| 5 | Kimi K3 | 59.7 | 35 | $6.00 |
| 6 | GLM-5.3 | 59.5 | 0 | $2.15 |
| 7 | Qwen3.8 Max | 58.1 | 45 | $3.00 |
| 8 | Qwen3.8 2.4T A95B | 57.7 | 45 | $3.00 |
| 9 | Claude Opus 4.8 | 57.3 | 0 | $10.00 |
| 10 | Muse Spark 1.2 | 56.8 | 0 | $2.00 |
Output tokens per second — higher is faster. Minimum intelligence score of 40.
| # | Model | tok/s |
|---|---|---|
| 1 | Gemini 3.7 Flash | 328 |
| 2 | Gemini 3.6 Flash | 201 |
| 3 | Nex-N2-Pro | 139 |
| 4 | GPT-5.6 Luna | 132 |
| 5 | GPT-5.3 Codex | 132 |
| 6 | Inkling Small | 129 |
| 7 | Gemini 3.1 Pro Preview | 120 |
| 8 | GPT-5.6 Terra | 119 |
| 9 | MiniMax-M3 | 116 |
| 10 | DeepSeek V4 Flash 0731 | 112 |
Blended cost per 1M tokens (3:1 input/output) — lower is cheaper. Minimum intelligence score of 40.
| # | Model | $/1M |
|---|---|---|
| 1 | DeepSeek V4 Flash | $0.168 |
| 2 | Hy3 | $0.241 |
| 3 | GPT-5.6 Luna | $0.45 |
| 4 | MiniMax-M3 | $0.525 |
| 5 | Solar Pro 4 | $0.525 |
| 6 | Inkling Small | $0.525 |
| 7 | DeepSeek V4 Pro | $0.544 |
| 8 | MiMo-V2.5-Pro | $0.544 |
| 9 | DeepSeek V4 Flash 0731 | $0.66 |
| 10 | Nex-N2-Pro | $1.00 |