On the SWE-rebench coding benchmark, the top tier remains frozen: AnthropicFable 5 holds 64.5% with Grok 4.5 and Anthropic Opus 5 trailing at 63.8% and 63.4% respectively, all within their established confidence intervals and showing no meaningful movement from the previous cycle. The Artificial Analysis general-purpose benchmark tells a different story, with Claude Opus 5.5 now leading at 57.6 (up from an unranked position), displacing Claude Opus 5 to seventh place and shuffling the top ten substantially: Claude Sonnet 5.5 enters at 56.0, Claude Fable 5.1 at 53.4, and the GPT-6 variants (Astra at 52.7, Sol at 51.8) consolidate ground previously held by older models. This reshuffling reflects the release cadence of frontier models rather than methodological discovery. The coding benchmark's stability is notable given the variance margins, most top performers fall within 0.6 to 1.83 percentage points of uncertainty, yet no reordering occurs, suggesting either that the task has reached saturation for leading systems or that confidence intervals genuinely constrain detectability of small gains. The Artificial Analysis list shows more flux: Claude Haiku 5.5 and GLM-5.3-Flash appear as new entries at positions 19 and 21, while GLM 5.3 Flash dropped entirely, indicating version-specific performance variation rather than sustained architectural progress. Taken together, the benchmarks suggest the frontier has plateaued on structured code tasks while general capabilities continue to shift with model releases, a pattern consistent with specialization rather than across-the-board improvement.
Cole Brennan
Daily rankings from SWE-rebench, a benchmark designed to fairly compare LLM capabilities on real-world software engineering tasks. Unlike other evaluations, it uses a standardized scaffolding for all models, continuously updates its dataset to prevent contamination, and runs each model five times to account for stochastic variance.
| # | Model | Score |
|---|---|---|
| 1 | AnthropicFable 5 [high]Model | 64.5%± 1.41% |
| 2 | GrokGrok 4.5 [high]Model | 63.8%± 0.60% |
| 3 | AnthropicOpus 5 [high]Model | 63.4%± 1.35% |
| 4 | Z.aiGLM-5.2 [high]Model | 62.9%± 1.19% |
| 5 | OpenAIGPT-5.6 Sol [medium]Model | 62.3%± 1.83% |
| 6 | JunieJunieAgent | 61.8%± 0.54% |
| 7 | AnthropicClaude CodeAgent | 60.4%± 1.03% |
| 8 | OpenAICodexAgent | 58.0%± 1.29% |
| 9 | AnthropicSonnet 5 [high]Model | 56.8%± 0.94% |
| 10 | CursorCursorAgent | 51.7%± 0.84% |
Artificial Analysis composite index across coding, math, and reasoning benchmarks.
| # | Model | Score | tok/s | $/1M |
|---|---|---|---|---|
| 1 | Claude Opus 5.5 | 57.6 | 98 | $8.00 |
| 2 | Claude Sonnet 5.5 | 56 | 133 | $4.00 |
| 3 | Claude Fable 5.1 | 53.4 | 67 | $20.00 |
| 4 | GPT-6 Astra | 52.7 | 48 | $20.00 |
| 5 | Gemini 4 Argon | 52.6 | 0 | $4.00 |
| 6 | GPT-6.1 Sol | 51.8 | 57 | $4.00 |
| 7 | Claude Opus 5 | 50.8 | 0 | $10.00 |
| 8 | Claude Fable 5 | 49.6 | 0 | $20.00 |
| 9 | Muse Spark 1.3 | 48.1 | 103 | $2.00 |
| 10 | GPT-6 Sol | 47.6 | 0 | $4.00 |
Output tokens per second — higher is faster. Minimum intelligence score of 40.
| # | Model | tok/s |
|---|---|---|
| 1 | Claude Haiku 5.5 | 240 |
| 2 | Ling 3.1 Flash | 215 |
| 3 | Gemini 3.8 Flash | 145 |
| 4 | Claude Sonnet 5.5 | 133 |
| 5 | Muse Spark 1.3 | 103 |
| 6 | GPT-5.6 Terra | 99 |
| 7 | Claude Opus 5.5 | 98 |
| 8 | Step 5 Preview | 89 |
| 9 | GLM-5.3 | 70 |
| 10 | Claude Fable 5.1 | 67 |
Blended cost per 1M tokens (3:1 input/output) — lower is cheaper. Minimum intelligence score of 40.
| # | Model | $/1M |
|---|---|---|
| 1 | Claude Haiku 5.5 | $0.20 |
| 2 | GLM-5.3-Flash | $0.237 |
| 3 | Ling 3.1 Flash | $0.45 |
| 4 | MiMo-V2.6-Pro | $0.544 |
| 5 | Step 5 Preview | $1.43 |
| 6 | Gemini 3.8 Flash | $1.50 |
| 7 | Muse Spark 1.3 | $2.00 |
| 8 | GLM-5.3 | $2.15 |
| 9 | Grok 4.7 | $3.00 |
| 10 | Qwen3.8 Max | $3.00 |