On the SWE-rebench coding benchmark, the rankings remain stable from the previous cycle: AnthropicFable 5 holds position one at 64.5% (±1.41%), followed by GrokGrok 4.5 at 63.8% (±0.60%), and AnthropicOpus 5 at 63.4% (±1.35%). The top seven positions cluster tightly between 60.4% and 64.5%, with confidence intervals that overlap substantially, making any claim of meaningful separation within this tier premature given the measurement uncertainty. The Artificial Analysis benchmark tells a different story: Claude Opus 5.5 enters at position one with a score of 57.6, displacing Claude Fable 5.1 to second place at 53.4, while GPT-6 Astra moves to third at 52.7. The new entries Claude Opus 5.5 and GPT-6 Sol (at 47.5, position seven) suggest recent model releases, though the Artificial Analysis methodology differs fundamentally from SWE-rebench and measures different problem classes, making direct cross-benchmark comparison invalid. Within each benchmark independently, the movements are modest: on SWE-rebench, the top performers show no ranking changes from the previous cycle, indicating evaluation consistency or saturation at the frontier; on Artificial Analysis, the displacement of Fable 5.1 by Opus 5.5 represents a 4.2-point gain that, while directional, lacks context about whether this reflects genuine capability advancement or variation in task difficulty. Without details on the Artificial Analysis evaluation methodology, confidence intervals, or the specific problems tested, whether these shifts constitute meaningful progress or normal measurement variance remains unclear.
Cole Brennan
Daily rankings from SWE-rebench, a benchmark designed to fairly compare LLM capabilities on real-world software engineering tasks. Unlike other evaluations, it uses a standardized scaffolding for all models, continuously updates its dataset to prevent contamination, and runs each model five times to account for stochastic variance.
| # | Model | Score |
|---|---|---|
| 1 | AnthropicFable 5 [high]Model | 64.5%± 1.41% |
| 2 | GrokGrok 4.5 [high]Model | 63.8%± 0.60% |
| 3 | AnthropicOpus 5 [high]Model | 63.4%± 1.35% |
| 4 | Z.aiGLM-5.2 [high]Model | 62.9%± 1.19% |
| 5 | OpenAIGPT-5.6 Sol [medium]Model | 62.3%± 1.83% |
| 6 | JunieJunieAgent | 61.8%± 0.54% |
| 7 | AnthropicClaude CodeAgent | 60.4%± 1.03% |
| 8 | OpenAICodexAgent | 58.0%± 1.29% |
| 9 | AnthropicSonnet 5 [high]Model | 56.8%± 0.94% |
| 10 | CursorCursorAgent | 51.7%± 0.84% |
Artificial Analysis composite index across coding, math, and reasoning benchmarks.
| # | Model | Score | tok/s | $/1M |
|---|---|---|---|---|
| 1 | Claude Opus 5.5 | 57.6 | 0 | $8.00 |
| 2 | Claude Fable 5.1 | 53.4 | 64 | $20.00 |
| 3 | GPT-6 Astra | 52.7 | 54 | $20.00 |
| 4 | Claude Opus 5 | 50.8 | 52 | $10.00 |
| 5 | Claude Fable 5 | 49.6 | 0 | $20.00 |
| 6 | Muse Spark 1.3 | 48.1 | 220 | $2.00 |
| 7 | GPT-6 Sol | 47.5 | 116 | $4.00 |
| 8 | GPT-5.6 Sol | 47 | 62 | $8.00 |
| 9 | Grok 4.7 | 46.4 | 50 | $3.00 |
| 10 | MiMo-V2.6-Pro | 46.3 | 54 | $0.544 |
Output tokens per second — higher is faster. Minimum intelligence score of 40.
| # | Model | tok/s |
|---|---|---|
| 1 | Gemini 3.8 Flash | 301 |
| 2 | Muse Spark 1.3 | 220 |
| 3 | GPT-6 Sol | 116 |
| 4 | GPT-5.6 Terra | 89 |
| 5 | Step 5 Preview | 75 |
| 6 | Grok 4.6 | 70 |
| 7 | Claude Fable 5.1 | 64 |
| 8 | GLM-5.3 | 63 |
| 9 | GPT-5.6 Sol | 62 |
| 10 | GPT-6 Astra | 54 |
Blended cost per 1M tokens (3:1 input/output) — lower is cheaper. Minimum intelligence score of 40.
| # | Model | $/1M |
|---|---|---|
| 1 | GLM 5.3 Flash | $0.237 |
| 2 | MiMo-V2.6-Pro | $0.544 |
| 3 | Step 5 Preview | $1.43 |
| 4 | Gemini 3.8 Flash | $1.50 |
| 5 | Muse Spark 1.3 | $2.00 |
| 6 | GLM-5.3 | $2.15 |
| 7 | Grok 4.7 | $3.00 |
| 8 | Qwen3.8 Max | $3.00 |
| 9 | Grok 4.6 | $3.00 |
| 10 | GPT-6 Sol | $4.00 |