The SWE-rebench standings show no movement at the top tier, with AnthropicFable 5 holding 64.5% and the next four positions (Grok 4.5, Opus 5, GLM-5.2, GPT-5.6 Sol) maintaining their rankings between 62.3% and 63.8%, all within margins that overlap their confidence intervals. The Artificial Analysis benchmark, by contrast, reveals substantial reordering across its full 422-entry leaderboard: Claude Opus 5 climbed to first place with 63.1 (up 2.4 points from 60.7), displacing Fable 5 to second at 62.1, while GPT-5.6 Sol moved from third to third with a 2.0-point gain to 60.9, and deeper in the list models like Muse Spark 1.2 jumped nine positions and 2.7 points to number 7. The divergence between these two benchmarks reflects their different methodologies: SWE-rebench measures coding agent performance on concrete repository tasks with tight confidence intervals that constrain apparent movement, while Artificial Analysis appears to evaluate broader model capabilities with larger score spreads. The top-tier SWE-rebench stability suggests either that the models tested are genuinely separated by meaningful gaps (the 1.1-point spread between first and fifth is real), or that the benchmark's variance is high enough to mask smaller changes, whereas Artificial Analysis's volatility in the 30-50 rank range and consistent 1-3 point gains across the board hints at either a recalibrated evaluation, a different test set, or systematic score inflation. Without knowing whether Artificial Analysis changed its methodology or simply re-ran the same test, the pattern is ambiguous: the gains could reflect genuine model improvement, measurement drift, or both.
Cole Brennan
Daily rankings from SWE-rebench, a benchmark designed to fairly compare LLM capabilities on real-world software engineering tasks. Unlike other evaluations, it uses a standardized scaffolding for all models, continuously updates its dataset to prevent contamination, and runs each model five times to account for stochastic variance.
| # | Model | Score |
|---|---|---|
| 1 | AnthropicFable 5 [high]Model | 64.5%± 1.41% |
| 2 | GrokGrok 4.5 [high]Model | 63.8%± 0.60% |
| 3 | AnthropicOpus 5 [high]Model | 63.4%± 1.35% |
| 4 | Z.aiGLM-5.2 [high]Model | 62.9%± 1.19% |
| 5 | OpenAIGPT-5.6 Sol [medium]Model | 62.3%± 1.83% |
| 6 | JunieJunieAgent | 61.8%± 0.54% |
| 7 | AnthropicClaude CodeAgent | 60.4%± 1.03% |
| 8 | OpenAICodexAgent | 58.0%± 1.29% |
| 9 | AnthropicSonnet 5 [high]Model | 56.8%± 0.94% |
| 10 | CursorCursorAgent | 51.7%± 0.84% |
Artificial Analysis composite index across coding, math, and reasoning benchmarks.
| # | Model | Score | tok/s | $/1M |
|---|---|---|---|---|
| 1 | Claude Opus 5 | 63.1 | 53 | $10.00 |
| 2 | Claude Fable 5 | 62.1 | 61 | $20.00 |
| 3 | GPT-5.6 Sol | 60.9 | 63 | $11.25 |
| 4 | Kimi K3 | 59.7 | 37 | $6.00 |
| 5 | Qwen3.8 Max | 58.1 | 69 | $3.00 |
| 6 | Claude Opus 4.8 | 57.3 | 0 | $10.00 |
| 7 | Muse Spark 1.2 | 56.8 | 0 | $2.00 |
| 8 | GPT-5.6 Terra | 56.6 | 125 | $4.50 |
| 9 | GPT-5.5 | 56.3 | 0 | $11.25 |
| 10 | Grok 4.5 | 55.8 | 53 | $3.00 |
Output tokens per second — higher is faster. Minimum intelligence score of 40.
| # | Model | tok/s |
|---|---|---|
| 1 | Gemini 3.5 Flash | 226 |
| 2 | Qwen3.7 Max | 196 |
| 3 | Muse Spark 1.1 | 192 |
| 4 | Gemini 3.6 Flash | 189 |
| 5 | GPT-5.6 Luna | 169 |
| 6 | Gemini 3.1 Pro Preview | 128 |
| 7 | GPT-5.6 Terra | 125 |
| 8 | Nex-N2-Pro | 119 |
| 9 | Inkling Small | 119 |
| 10 | GLM-5.2 | 115 |
Blended cost per 1M tokens (3:1 input/output) — lower is cheaper. Minimum intelligence score of 40.
| # | Model | $/1M |
|---|---|---|
| 1 | DeepSeek V4 Flash | $0.168 |
| 2 | DeepSeek V4 Flash 0731 | $0.175 |
| 3 | Hy3 | $0.241 |
| 4 | GPT-5.6 Luna | $0.45 |
| 5 | MiniMax-M3 | $0.525 |
| 6 | Inkling Small | $0.525 |
| 7 | DeepSeek V4 Pro | $0.544 |
| 8 | MiMo-V2.5-Pro | $0.544 |
| 9 | Nex-N2-Pro | $1.00 |
| 10 | Qwen3.6 Plus | $1.13 |