The SWE-rebench results show no movement since the previous cycle: AnthropicFable 5 holds position one at 64.5% with a confidence interval of 1.41%, followed by GrokGrok 4.5 at 63.8% and AnthropicOpus 5 at 63.4%, identical to prior rankings. The Artificial Analysis leaderboard exhibits modest churn across its 442-entry roster, with two new entrants, Granite 4.2 30B at position 142 (23.7) and Agnes 2.5 Pro Beta at position 29 (49.1), displacing models that previously held those slots. The coding benchmark's stability across top performers suggests either convergence toward a ceiling or insufficient temporal resolution to detect meaningful gains; the tight error bars (0.54 to 1.83 percentage points) indicate the measurements themselves are reliable, but the lack of differentiation between cycles raises a question about whether SWE-rebench remains sensitive to model improvements at this performance tier. Conversely, the Artificial Analysis benchmark's broader movement pattern, particularly the entry of newer variants like Granite 4.2 and Agnes 2.5 Pro Beta, implies that general-capability benchmarks continue to register iterative progress across the field, though the magnitude of individual shifts remains modest and the methodology underlying Artificial Analysis scores is opaque, making it difficult to assess whether the observed reordering reflects genuine capability changes or variance in evaluation conditions.
Cole Brennan
Daily rankings from SWE-rebench, a benchmark designed to fairly compare LLM capabilities on real-world software engineering tasks. Unlike other evaluations, it uses a standardized scaffolding for all models, continuously updates its dataset to prevent contamination, and runs each model five times to account for stochastic variance.
| # | Model | Score |
|---|---|---|
| 1 | AnthropicFable 5 [high]Model | 64.5%± 1.41% |
| 2 | GrokGrok 4.5 [high]Model | 63.8%± 0.60% |
| 3 | AnthropicOpus 5 [high]Model | 63.4%± 1.35% |
| 4 | Z.aiGLM-5.2 [high]Model | 62.9%± 1.19% |
| 5 | OpenAIGPT-5.6 Sol [medium]Model | 62.3%± 1.83% |
| 6 | JunieJunieAgent | 61.8%± 0.54% |
| 7 | AnthropicClaude CodeAgent | 60.4%± 1.03% |
| 8 | OpenAICodexAgent | 58.0%± 1.29% |
| 9 | AnthropicSonnet 5 [high]Model | 56.8%± 0.94% |
| 10 | CursorCursorAgent | 51.7%± 0.84% |
Artificial Analysis composite index across coding, math, and reasoning benchmarks.
| # | Model | Score | tok/s | $/1M |
|---|---|---|---|---|
| 1 | Claude Opus 5 | 63.1 | 57 | $10.00 |
| 2 | Claude Fable 5 | 62.1 | 71 | $20.00 |
| 3 | GPT-5.6 Sol | 60.9 | 78 | $8.00 |
| 4 | Grok 4.6 | 60.9 | 60 | $3.00 |
| 5 | Kimi K3 | 59.7 | 38 | $6.00 |
| 6 | GLM-5.3 | 59.5 | 66 | $2.15 |
| 7 | Qwen3.8 Max | 58.1 | 26 | $3.00 |
| 8 | Qwen3.8 2.4T A95B | 57.7 | 24 | $3.00 |
| 9 | GLM-5.3-Flash | 57.5 | 49 | $0.237 |
| 10 | Claude Opus 4.8 | 57.3 | 0 | $10.00 |
Output tokens per second — higher is faster. Minimum intelligence score of 40.
| # | Model | tok/s |
|---|---|---|
| 1 | Gemini 3.7 Flash | 399 |
| 2 | Gemini 3.6 Flash | 186 |
| 3 | Agnes 2.5 Pro Beta | 141 |
| 4 | Nex-N2-Pro | 139 |
| 5 | DeepSeek V4 Flash 0731 | 137 |
| 6 | GPT-5.3 Codex | 126 |
| 7 | GPT-5.6 Luna | 123 |
| 8 | Gemini 3.1 Pro Preview | 123 |
| 9 | DeepSeek V4 Flash Vision | 117 |
| 10 | GPT-5.6 Terra | 107 |
Blended cost per 1M tokens (3:1 input/output) — lower is cheaper. Minimum intelligence score of 40.
| # | Model | $/1M |
|---|---|---|
| 1 | Agnes 2.5 Pro Beta | $0.15 |
| 2 | DeepSeek V4 Flash | $0.168 |
| 3 | Qwen3.8-Flash-Next | $0.23 |
| 4 | GLM-5.3-Flash | $0.237 |
| 5 | Hy3 | $0.241 |
| 6 | GPT-5.6 Luna | $0.45 |
| 7 | MiniMax-M3 | $0.525 |
| 8 | Solar Pro 4 | $0.525 |
| 9 | Inkling Small | $0.525 |
| 10 | DeepSeek V4 Pro | $0.544 |