The SWE-rebench rankings remain unchanged at the top tier, with AnthropicFable 5 holding at 64.5 percent and the next eight positions identical to the previous cycle, but the Artificial Analysis leaderboard has undergone significant restructuring below the top 20. Claude Opus 5 and Claude Fable 5 have swapped positions, now ranking first and second respectively at 60.7 and 59.9, displacing GPT-5.6 Sol to third, while Kimi K3 enters the top five at 57.1. The most substantial movement occurs in the 80 to 420 entry range, where G9v3-39A5B debuts at position 88 on Artificial Analysis, pushing prior entries down by one rank through the tail of the list. Within the SWE-rebench benchmark, the consistency of scores and confidence intervals across the top performers suggests either a plateau in incremental gains or stable evaluation conditions, though the methodology of both benchmarks warrants scrutiny: SWE-rebench appears to measure code generation on pull request resolution tasks with relatively tight confidence bounds, while Artificial Analysis aggregates across a broader set of capabilities, making direct comparison between the two frameworks problematic. The data shows no evidence of breakthrough performance on either metric; the top SWE-rebench model at 64.5 percent leaves substantial room for improvement on a task designed to reflect real-world software engineering, and the Artificial Analysis rankings reflect primarily reshuffling rather than systematic score inflation across the board.
Cole Brennan
Daily rankings from SWE-rebench, a benchmark designed to fairly compare LLM capabilities on real-world software engineering tasks. Unlike other evaluations, it uses a standardized scaffolding for all models, continuously updates its dataset to prevent contamination, and runs each model five times to account for stochastic variance.
| # | Model | Score |
|---|---|---|
| 1 | AnthropicFable 5 [high]Model | 64.5%± 1.41% |
| 2 | GrokGrok 4.5 [high]Model | 63.8%± 0.60% |
| 3 | AnthropicOpus 5 [high]Model | 63.4%± 1.35% |
| 4 | Z.aiGLM-5.2 [high]Model | 62.9%± 1.19% |
| 5 | OpenAIGPT-5.6 Sol [medium]Model | 62.3%± 1.83% |
| 6 | JunieJunieAgent | 61.8%± 0.54% |
| 7 | AnthropicClaude CodeAgent | 60.4%± 1.03% |
| 8 | OpenAICodexAgent | 58.0%± 1.29% |
| 9 | AnthropicSonnet 5 [high]Model | 56.8%± 0.94% |
| 10 | CursorCursorAgent | 51.7%± 0.84% |
Artificial Analysis composite index across coding, math, and reasoning benchmarks.
| # | Model | Score | tok/s | $/1M |
|---|---|---|---|---|
| 1 | Claude Opus 5 | 60.7 | 60 | $10.00 |
| 2 | Claude Fable 5 | 59.9 | 67 | $20.00 |
| 3 | GPT-5.6 Sol | 58.9 | 80 | $11.25 |
| 4 | Kimi K3 | 57.1 | 34 | $6.00 |
| 5 | Claude Opus 4.8 | 55.7 | 0 | $10.00 |
| 6 | GPT-5.6 Terra | 55 | 131 | $4.50 |
| 7 | GPT-5.5 | 54.8 | 0 | $11.25 |
| 8 | Grok 4.5 | 53.8 | 63 | $3.00 |
| 9 | Claude Opus 4.7 | 53.5 | 0 | $10.00 |
| 10 | Claude Sonnet 5 | 53.4 | 88 | $4.00 |
Output tokens per second — higher is faster. Minimum intelligence score of 40.
| # | Model | tok/s |
|---|---|---|
| 1 | Gemini 3.5 Flash | 281 |
| 2 | Gemini 3.6 Flash | 231 |
| 3 | GLM-5.2 | 212 |
| 4 | Qwen3.7 Max | 208 |
| 5 | Muse Spark 1.1 | 206 |
| 6 | GPT-5.6 Luna | 172 |
| 7 | Nex-N2-Pro | 136 |
| 8 | Gemini 3.1 Pro Preview | 134 |
| 9 | GPT-5.6 Terra | 131 |
| 10 | Inkling Small | 131 |
Blended cost per 1M tokens (3:1 input/output) — lower is cheaper. Minimum intelligence score of 40.
| # | Model | $/1M |
|---|---|---|
| 1 | DeepSeek V4 Flash | $0.171 |
| 2 | DeepSeek V4 Flash 0731 | $0.175 |
| 3 | Hy3 | $0.241 |
| 4 | GPT-5.6 Luna | $0.45 |
| 5 | MiniMax-M3 | $0.525 |
| 6 | Inkling Small | $0.525 |
| 7 | DeepSeek V4 Pro | $0.544 |
| 8 | MiMo-V2.5-Pro | $0.544 |
| 9 | Nex-N2-Pro | $1.00 |
| 10 | GPT-5.4 mini | $1.69 |