The SWE-rebench results show no movement in the top tier: AnthropicFable 5 holds position one at 64.5% with confidence intervals tight enough to establish clear separation from Grok 4.5 at 63.8% and Opus 5 at 63.4%. All three models maintain their prior rankings, and the spread between first and fifth place remains 2.2 percentage points, suggesting the frontier on coding tasks has plateaued or these measurements have reached the precision limits of the benchmark itself. The Artificial Analysis benchmark presents a different picture: Gemini 3.8 Flash enters at position eight with 58.7, displacing Qwen3.8 Max to ninth, while the rest of the top twenty shifts down one position. This insertion is methodologically notable because Artificial Analysis covers a broader evaluation surface than SWE-rebench alone, and the entry of a Google model in the upper ranks hints at different task distributions or evaluation protocols between the two datasets. However, the discrepancy between rankings raises a practical question: SWE-rebench measures direct code generation on real repository issues with verifiable outcomes, whereas Artificial Analysis appears to aggregate multiple capability dimensions. The coding-specific benchmark shows mature performance clustering at the top, with Anthropic and OpenAI variants dominating, while the broader benchmark permits more competition from models like Gemini that may excel on non-coding tasks. Without knowing whether Artificial Analysis includes or weights SWE-rebench results, the gap between these two leaderboards suggests they are measuring related but distinct problems, and neither alone captures the full picture of model capability on engineering tasks.
Cole Brennan
Daily rankings from SWE-rebench, a benchmark designed to fairly compare LLM capabilities on real-world software engineering tasks. Unlike other evaluations, it uses a standardized scaffolding for all models, continuously updates its dataset to prevent contamination, and runs each model five times to account for stochastic variance.
| # | Model | Score |
|---|---|---|
| 1 | AnthropicFable 5 [high]Model | 64.5%± 1.41% |
| 2 | GrokGrok 4.5 [high]Model | 63.8%± 0.60% |
| 3 | AnthropicOpus 5 [high]Model | 63.4%± 1.35% |
| 4 | Z.aiGLM-5.2 [high]Model | 62.9%± 1.19% |
| 5 | OpenAIGPT-5.6 Sol [medium]Model | 62.3%± 1.83% |
| 6 | JunieJunieAgent | 61.8%± 0.54% |
| 7 | AnthropicClaude CodeAgent | 60.4%± 1.03% |
| 8 | OpenAICodexAgent | 58.0%± 1.29% |
| 9 | AnthropicSonnet 5 [high]Model | 56.8%± 0.94% |
| 10 | CursorCursorAgent | 51.7%± 0.84% |
Artificial Analysis composite index across coding, math, and reasoning benchmarks.
| # | Model | Score | tok/s | $/1M |
|---|---|---|---|---|
| 1 | Claude Fable 5.1 | 65.7 | 69 | $20.00 |
| 2 | Claude Opus 5 | 63.1 | 48 | $10.00 |
| 3 | Claude Fable 5 | 62.1 | 59 | $20.00 |
| 4 | GPT-5.6 Sol | 60.9 | 76 | $8.00 |
| 5 | Grok 4.6 | 60.9 | 56 | $3.00 |
| 6 | Kimi K3 | 59.7 | 39 | $6.00 |
| 7 | GLM-5.3 | 59.5 | 72 | $2.15 |
| 8 | Gemini 3.8 Flash | 58.7 | 297 | $1.50 |
| 9 | Qwen3.8 Max | 58.1 | 40 | $3.00 |
| 10 | Qwen3.8 2.4T A95B | 57.7 | 39 | $3.00 |
Output tokens per second — higher is faster. Minimum intelligence score of 40.
| # | Model | tok/s |
|---|---|---|
| 1 | Gemini 3.8 Flash | 297 |
| 2 | Gemini 3.7 Flash | 292 |
| 3 | Quasar 438B | 185 |
| 4 | Muse Spark 1.2 | 163 |
| 5 | Gemini 3.6 Flash | 161 |
| 6 | Agnes 2.5 Pro Beta | 154 |
| 7 | Nex-N2-Pro | 131 |
| 8 | GPT-5.6 Luna | 116 |
| 9 | GPT-5.3 Codex | 116 |
| 10 | Gemini 3.1 Pro Preview | 114 |
Blended cost per 1M tokens (3:1 input/output) — lower is cheaper. Minimum intelligence score of 40.
| # | Model | $/1M |
|---|---|---|
| 1 | Agnes 2.5 Pro Beta | $0.15 |
| 2 | DeepSeek V4 Flash | $0.168 |
| 3 | Qwen3.8-Flash-Next | $0.23 |
| 4 | GLM-5.3-Flash | $0.237 |
| 5 | Hy3 | $0.241 |
| 6 | GPT-5.6 Luna | $0.45 |
| 7 | MiniMax-M3 | $0.525 |
| 8 | Solar Pro 4 | $0.525 |
| 9 | Inkling Small | $0.525 |
| 10 | DeepSeek V4 Pro | $0.544 |