The SWE-rebench rankings remain static across both surveys, with AnthropicFable 5 holding the top position at 64.5 percent and the same seventeen models occupying positions one through seventeen in identical order. The Artificial Analysis benchmark shows marginal movement in its broader field: Muse Spark 1.3 gained 0.3 points to reach 53.0, climbing from position five to maintain that slot; Mistral Large 3 moved from 9.7 to 11.1 points and rose from position 211 to 202; Gemini 3.5 Flash fell from 41.9 to 39.7 points, dropping from position 28 to 32; and Gemini 3.6 Flash shifted from position 32 to 31 while holding steady at 40.3 points. The SWE-rebench evaluation uses controlled conditions with reported confidence intervals, though the methodology for the Artificial Analysis benchmark remains opaque regarding test set composition, evaluation protocol, and statistical rigor, making it difficult to assess whether these fractional gains represent meaningful capability differences or measurement noise. The coding agent benchmark's stability at the top tier suggests the frontier models have reached a performance plateau on this task, while the Artificial Analysis results show typical variance patterns consistent with evaluation noise rather than systematic improvement across the broader model landscape.
Cole Brennan
Daily rankings from SWE-rebench, a benchmark designed to fairly compare LLM capabilities on real-world software engineering tasks. Unlike other evaluations, it uses a standardized scaffolding for all models, continuously updates its dataset to prevent contamination, and runs each model five times to account for stochastic variance.
| # | Model | Score |
|---|---|---|
| 1 | AnthropicFable 5 [high]Model | 64.5%± 1.41% |
| 2 | GrokGrok 4.5 [high]Model | 63.8%± 0.60% |
| 3 | AnthropicOpus 5 [high]Model | 63.4%± 1.35% |
| 4 | Z.aiGLM-5.2 [high]Model | 62.9%± 1.19% |
| 5 | OpenAIGPT-5.6 Sol [medium]Model | 62.3%± 1.83% |
| 6 | JunieJunieAgent | 61.8%± 0.54% |
| 7 | AnthropicClaude CodeAgent | 60.4%± 1.03% |
| 8 | OpenAICodexAgent | 58.0%± 1.29% |
| 9 | AnthropicSonnet 5 [high]Model | 56.8%± 0.94% |
| 10 | CursorCursorAgent | 51.7%± 0.84% |
Artificial Analysis composite index across coding, math, and reasoning benchmarks.
| # | Model | Score | tok/s | $/1M |
|---|---|---|---|---|
| 1 | Claude Fable 5.1 | 56.8 | 70 | $20.00 |
| 2 | GPT-6 Astra | 54.7 | 61 | $20.00 |
| 3 | Claude Opus 5 | 54.1 | 49 | $10.00 |
| 4 | Claude Fable 5 | 53.2 | 59 | $20.00 |
| 5 | Muse Spark 1.3 | 53 | 177 | $2.00 |
| 6 | GPT-5.6 Sol | 51.3 | 79 | $8.00 |
| 7 | Grok 4.6 | 50.6 | 57 | $3.00 |
| 8 | Kimi K3 | 50.2 | 39 | $6.00 |
| 9 | GLM-5.3 | 48.6 | 77 | $2.15 |
| 10 | Gemini 3.8 Flash | 47.1 | 340 | $1.50 |
Output tokens per second — higher is faster. Minimum intelligence score of 40.
| # | Model | tok/s |
|---|---|---|
| 1 | Gemini 3.8 Flash | 340 |
| 2 | Gemini 3.7 Flash | 284 |
| 3 | Muse Spark 1.2 | 224 |
| 4 | Gemini 3.6 Flash | 186 |
| 5 | Muse Spark 1.3 | 177 |
| 6 | DeepSeek V4 Flash Vision | 121 |
| 7 | DeepSeek V4 Flash 0731 | 118 |
| 8 | GPT-5.6 Luna | 110 |
| 9 | GPT-5.6 Terra | 104 |
| 10 | GPT-5.6 Sol | 79 |
Blended cost per 1M tokens (3:1 input/output) — lower is cheaper. Minimum intelligence score of 40.
| # | Model | $/1M |
|---|---|---|
| 1 | Qwen3.8-Flash-Next | $0.23 |
| 2 | GLM-5.3-Flash | $0.237 |
| 3 | GPT-5.6 Luna | $0.45 |
| 4 | DeepSeek V4 Flash Vision | $0.66 |
| 5 | DeepSeek V4 Flash 0731 | $0.66 |
| 6 | Qwen3.8 27B | $1.13 |
| 7 | Gemini 3.8 Flash | $1.50 |
| 8 | Gemini 3.7 Flash | $1.50 |
| 9 | Gemini 3.6 Flash | $1.50 |
| 10 | DeepSeek V4 Pro 0813 | $1.98 |