The SWE-rebench standings remain frozen at the top, with AnthropicFable 5 holding 64.5% and the next four positions unchanged, but the Artificial Analysis rankings show material churn across the full leaderboard. GPT-6.1 Sol enters at #5 (51.8), pushing Claude Opus 5 down one slot to #6, while lower-ranked models experience more dramatic shifts: JT-4.1 Flash 236B A21B surges from #62 (27.3) to #39 (33.9), a gain of 6.6 points that stands out as the most substantial movement in the visible range. Ling 3.0 Tiny drops sharply from #155 (15.3) to #216 (11.1), losing 4.2 points, suggesting the benchmark may be reweighting reasoning or coding tasks where that model underperforms. Most other movements cluster between one and three positions, typical churn in a crowded field. The Artificial Analysis benchmark itself lacks transparency about evaluation methodology, making it difficult to assess whether these shifts reflect genuine capability changes or volatility in how tasks are selected and scored. SWE-rebench's stability at the top suggests either convergence among leading models or that its test set has reached saturation; the absence of new entries in the top 17 over this cycle is notable. Without visibility into what changed in Artificial Analysis's test construction or weighting, the practical significance of these rankings remains unclear.
Cole Brennan
Daily rankings from SWE-rebench, a benchmark designed to fairly compare LLM capabilities on real-world software engineering tasks. Unlike other evaluations, it uses a standardized scaffolding for all models, continuously updates its dataset to prevent contamination, and runs each model five times to account for stochastic variance.
| # | Model | Score |
|---|---|---|
| 1 | AnthropicFable 5 [high]Model | 64.5%± 1.41% |
| 2 | GrokGrok 4.5 [high]Model | 63.8%± 0.60% |
| 3 | AnthropicOpus 5 [high]Model | 63.4%± 1.35% |
| 4 | Z.aiGLM-5.2 [high]Model | 62.9%± 1.19% |
| 5 | OpenAIGPT-5.6 Sol [medium]Model | 62.3%± 1.83% |
| 6 | JunieJunieAgent | 61.8%± 0.54% |
| 7 | AnthropicClaude CodeAgent | 60.4%± 1.03% |
| 8 | OpenAICodexAgent | 58.0%± 1.29% |
| 9 | AnthropicSonnet 5 [high]Model | 56.8%± 0.94% |
| 10 | CursorCursorAgent | 51.7%± 0.84% |
Artificial Analysis composite index across coding, math, and reasoning benchmarks.
| # | Model | Score | tok/s | $/1M |
|---|---|---|---|---|
| 1 | Claude Opus 5.5 | 57.6 | 96 | $8.00 |
| 2 | Claude Sonnet 5.5 | 56 | 145 | $4.00 |
| 3 | Claude Fable 5.1 | 53.4 | 69 | $20.00 |
| 4 | GPT-6 Astra | 52.7 | 55 | $20.00 |
| 5 | GPT-6.1 Sol | 51.8 | 80 | $4.00 |
| 6 | Claude Opus 5 | 50.8 | 0 | $10.00 |
| 7 | Claude Fable 5 | 49.6 | 0 | $20.00 |
| 8 | Muse Spark 1.3 | 48.1 | 189 | $2.00 |
| 9 | GPT-6 Sol | 47.5 | 80 | $4.00 |
| 10 | GPT-5.6 Sol | 47 | 0 | $8.00 |
Output tokens per second — higher is faster. Minimum intelligence score of 40.
| # | Model | tok/s |
|---|---|---|
| 1 | Gemini 3.8 Flash | 222 |
| 2 | Muse Spark 1.3 | 189 |
| 3 | Claude Sonnet 5.5 | 145 |
| 4 | GPT-5.6 Terra | 108 |
| 5 | Claude Opus 5.5 | 96 |
| 6 | Step 5 Preview | 86 |
| 7 | Grok 4.7 | 83 |
| 8 | GPT-6.1 Sol | 80 |
| 9 | GPT-6 Sol | 80 |
| 10 | GLM-5.3 | 70 |
Blended cost per 1M tokens (3:1 input/output) — lower is cheaper. Minimum intelligence score of 40.
| # | Model | $/1M |
|---|---|---|
| 1 | GLM 5.3 Flash | $0.237 |
| 2 | MiMo-V2.6-Pro | $0.544 |
| 3 | Step 5 Preview | $1.43 |
| 4 | Gemini 3.8 Flash | $1.50 |
| 5 | Muse Spark 1.3 | $2.00 |
| 6 | GLM-5.3 | $2.15 |
| 7 | Grok 4.7 | $3.00 |
| 8 | Qwen3.8 Max | $3.00 |
| 9 | Grok 4.6 | $3.00 |
| 10 | Claude Sonnet 5.5 | $4.00 |