The SWE-rebench results remain static across the top tier, with Anthropic Fable 5 holding 64.5% and the next four positions unchanged through Grok 4.5, Opus 5, GLM-5.2, and GPT-5.6 Sol, all within their confidence intervals from the previous cycle. The Artificial Analysis benchmark shows more churn in the middle ranks but reveals a pattern worth noting: Claude Opus 4.8 gained 0.4 points to reach 47.8, Claude Opus 4.7 held steady at 44.3, and Claude Sonnet 5 remained at 45.1, suggesting Anthropic's line is consolidating performance across different model sizes. DeepSeek V3 jumped 2.1 points from 8.3 to 10.4, placing it at #212 and marking the largest single-step gain in the visible range. Qwen3 32B moved from 5.7 to 8.0 (entry #232), a 2.3-point improvement. Lower down the list, gpt-oss-20b gained 1.7 points to 10.8, and Trinity Large Thinking moved from #190 to #181 with a 1.1-point gain to 13.1. The methodology here differs between benchmarks: SWE-rebench measures code generation on real software engineering tasks with confidence intervals, while Artificial Analysis reports single-point scores without error bounds, making the two incomparable directly. Neither benchmark shows the kind of concentrated improvement at the top that would suggest a fundamental breakthrough. The gains are distributed, incremental, and within the noise of typical model iteration.
Cole Brennan
Daily rankings from SWE-rebench, a benchmark designed to fairly compare LLM capabilities on real-world software engineering tasks. Unlike other evaluations, it uses a standardized scaffolding for all models, continuously updates its dataset to prevent contamination, and runs each model five times to account for stochastic variance.
| # | Model | Score |
|---|---|---|
| 1 | AnthropicFable 5 [high]Model | 64.5%± 1.41% |
| 2 | GrokGrok 4.5 [high]Model | 63.8%± 0.60% |
| 3 | AnthropicOpus 5 [high]Model | 63.4%± 1.35% |
| 4 | Z.aiGLM-5.2 [high]Model | 62.9%± 1.19% |
| 5 | OpenAIGPT-5.6 Sol [medium]Model | 62.3%± 1.83% |
| 6 | JunieJunieAgent | 61.8%± 0.54% |
| 7 | AnthropicClaude CodeAgent | 60.4%± 1.03% |
| 8 | OpenAICodexAgent | 58.0%± 1.29% |
| 9 | AnthropicSonnet 5 [high]Model | 56.8%± 0.94% |
| 10 | CursorCursorAgent | 51.7%± 0.84% |
Artificial Analysis composite index across coding, math, and reasoning benchmarks.
| # | Model | Score | tok/s | $/1M |
|---|---|---|---|---|
| 1 | Claude Fable 5.1 | 56.8 | 68 | $20.00 |
| 2 | GPT-6 Astra | 54.7 | 63 | $20.00 |
| 3 | Claude Opus 5 | 54.1 | 52 | $10.00 |
| 4 | Claude Fable 5 | 53.2 | 62 | $20.00 |
| 5 | Muse Spark 1.3 | 53 | 221 | $2.00 |
| 6 | GPT-5.6 Sol | 51.3 | 73 | $8.00 |
| 7 | Grok 4.6 | 50.6 | 57 | $3.00 |
| 8 | Kimi K3 | 50.2 | 42 | $6.00 |
| 9 | GLM-5.3 | 48.6 | 75 | $2.15 |
| 10 | Claude Opus 4.8 | 47.8 | 0 | $10.00 |
Output tokens per second — higher is faster. Minimum intelligence score of 40.
| # | Model | tok/s |
|---|---|---|
| 1 | Gemini 3.8 Flash | 356 |
| 2 | Gemini 3.7 Flash | 325 |
| 3 | Muse Spark 1.2 | 262 |
| 4 | Muse Spark 1.3 | 221 |
| 5 | Gemini 3.6 Flash | 188 |
| 6 | GPT-5.6 Terra | 121 |
| 7 | DeepSeek V4 Flash Vision | 120 |
| 8 | DeepSeek V4 Flash 0731 | 118 |
| 9 | GPT-5.6 Luna | 117 |
| 10 | Claude Sonnet 5 | 80 |
Blended cost per 1M tokens (3:1 input/output) — lower is cheaper. Minimum intelligence score of 40.
| # | Model | $/1M |
|---|---|---|
| 1 | Qwen3.8-Flash-Next | $0.23 |
| 2 | GLM-5.3-Flash | $0.237 |
| 3 | GPT-5.6 Luna | $0.45 |
| 4 | DeepSeek V4 Flash 0731 | $0.66 |
| 5 | DeepSeek V4 Flash Vision | $0.66 |
| 6 | Qwen3.8 27B | $1.13 |
| 7 | Gemini 3.8 Flash | $1.50 |
| 8 | Gemini 3.7 Flash | $1.50 |
| 9 | Gemini 3.6 Flash | $1.50 |
| 10 | DeepSeek V4 Pro 0813 | $1.98 |