The SWE-rebench rankings remain static across the top tier, with AnthropicFable 5 holding 64.5% ± 1.41% at first place and the next six positions unchanged through OpenAICodexAgent at 58.0% ± 1.29%. This stability in the coding benchmark suggests either that the test set has reached saturation at the frontier or that recent model updates have not substantially altered performance on this particular evaluation. The Artificial Analysis benchmark, by contrast, shows considerable churn: Claude Sonnet 5.5 enters at rank 2 with 56.0, displacing Claude Fable 5.1 down one position to 3, while the rest of the top 20 shifts accordingly with no score changes, only ranking adjustments. The gap between the two benchmarks is worth noting. SWE-rebench scores cluster tightly in the 60s and 50s for leading models, with error bars ranging from ±0.54% to ±1.83%, reflecting controlled experimental conditions. Artificial Analysis scores begin at 57.6 for Claude Opus 5.5 and decline more steeply down the list, reaching single digits by rank 250. This divergence hints at different task difficulty profiles: SWE-rebench may be measuring a narrower, more saturated capability space (software engineering problem solving with defined test cases), while Artificial Analysis appears to cover a broader skill distribution. Neither benchmark reveals methodological details in the provided data, making it difficult to assess whether the stability in SWE-rebench reflects genuine parity or simply coarse-grained measurement. The entry of Claude Sonnet 5.5 into Artificial Analysis's top two is the only concrete movement worth tracking, but without prior scores for this model on that benchmark, its significance cannot be determined.
Cole Brennan
Daily rankings from SWE-rebench, a benchmark designed to fairly compare LLM capabilities on real-world software engineering tasks. Unlike other evaluations, it uses a standardized scaffolding for all models, continuously updates its dataset to prevent contamination, and runs each model five times to account for stochastic variance.
| # | Model | Score |
|---|---|---|
| 1 | AnthropicFable 5 [high]Model | 64.5%± 1.41% |
| 2 | GrokGrok 4.5 [high]Model | 63.8%± 0.60% |
| 3 | AnthropicOpus 5 [high]Model | 63.4%± 1.35% |
| 4 | Z.aiGLM-5.2 [high]Model | 62.9%± 1.19% |
| 5 | OpenAIGPT-5.6 Sol [medium]Model | 62.3%± 1.83% |
| 6 | JunieJunieAgent | 61.8%± 0.54% |
| 7 | AnthropicClaude CodeAgent | 60.4%± 1.03% |
| 8 | OpenAICodexAgent | 58.0%± 1.29% |
| 9 | AnthropicSonnet 5 [high]Model | 56.8%± 0.94% |
| 10 | CursorCursorAgent | 51.7%± 0.84% |
Artificial Analysis composite index across coding, math, and reasoning benchmarks.
| # | Model | Score | tok/s | $/1M |
|---|---|---|---|---|
| 1 | Claude Opus 5.5 | 57.6 | 96 | $8.00 |
| 2 | Claude Sonnet 5.5 | 56 | 145 | $4.00 |
| 3 | Claude Fable 5.1 | 53.4 | 69 | $20.00 |
| 4 | GPT-6 Astra | 52.7 | 62 | $20.00 |
| 5 | Claude Opus 5 | 50.8 | 0 | $10.00 |
| 6 | Claude Fable 5 | 49.6 | 0 | $20.00 |
| 7 | Muse Spark 1.3 | 48.1 | 189 | $2.00 |
| 8 | GPT-6 Sol | 47.5 | 84 | $4.00 |
| 9 | GPT-5.6 Sol | 47 | 0 | $8.00 |
| 10 | Grok 4.7 | 46.4 | 72 | $3.00 |
Output tokens per second — higher is faster. Minimum intelligence score of 40.
| # | Model | tok/s |
|---|---|---|
| 1 | Gemini 3.8 Flash | 238 |
| 2 | Muse Spark 1.3 | 189 |
| 3 | Claude Sonnet 5.5 | 145 |
| 4 | GPT-5.6 Terra | 109 |
| 5 | Claude Opus 5.5 | 96 |
| 6 | GPT-6 Sol | 84 |
| 7 | Step 5 Preview | 84 |
| 8 | GLM-5.3 | 75 |
| 9 | Grok 4.6 | 75 |
| 10 | Grok 4.7 | 72 |
Blended cost per 1M tokens (3:1 input/output) — lower is cheaper. Minimum intelligence score of 40.
| # | Model | $/1M |
|---|---|---|
| 1 | GLM 5.3 Flash | $0.237 |
| 2 | MiMo-V2.6-Pro | $0.544 |
| 3 | Step 5 Preview | $1.43 |
| 4 | Gemini 3.8 Flash | $1.50 |
| 5 | Muse Spark 1.3 | $2.00 |
| 6 | GLM-5.3 | $2.15 |
| 7 | Grok 4.7 | $3.00 |
| 8 | Qwen3.8 Max | $3.00 |
| 9 | Grok 4.6 | $3.00 |
| 10 | Claude Sonnet 5.5 | $4.00 |