The SWE-rebench rankings remain frozen while the Artificial Analysis leaderboard shuffled 38 entries, introducing Mistral Large 4 Preview at position 32 and cascading everything below it downward. On SWE-rebench, AnthropicFable 5 holds 64.5% with tight confidence intervals across the top tier, and the spread between first and tenth place spans 12.8 percentage points, suggesting real separation in coding task performance. The Artificial Analysis benchmark, by contrast, shows a different stratification: Claude Opus 5.5 leads at 57.6, but the top ten compress within a 5-point band, and models ranked 250 onward cluster between 9 and 5 percent with minimal differentiation. The methodology concern here is acute. SWE-rebench appears to test concrete problem-solving on real repository issues with controlled evaluation, whereas Artificial Analysis likely aggregates diverse metrics across broader capability domains where coding may be one signal among many. The two benchmarks rank models differently enough to suggest they measure distinct things: SWE-rebench's top performer (Fable 5 at 64.5%) ranks third on Artificial Analysis (53.4), while Artificial Analysis's leader (Opus 5.5 at 57.6) sits ninth on SWE-rebench (56.8%). This divergence is not noise. It reflects whether a benchmark prioritizes narrow coding competence or distributed capability. Neither ranking change signals meaningful progress without knowing whether the Mistral insertion reflects actual improvement or simply expanded roster coverage, and the SWE-rebench stability suggests either the evaluation is mature or updates are infrequent.
Cole Brennan
Daily rankings from SWE-rebench, a benchmark designed to fairly compare LLM capabilities on real-world software engineering tasks. Unlike other evaluations, it uses a standardized scaffolding for all models, continuously updates its dataset to prevent contamination, and runs each model five times to account for stochastic variance.
| # | Model | Score |
|---|---|---|
| 1 | AnthropicFable 5 [high]Model | 64.5%± 1.41% |
| 2 | GrokGrok 4.5 [high]Model | 63.8%± 0.60% |
| 3 | AnthropicOpus 5 [high]Model | 63.4%± 1.35% |
| 4 | Z.aiGLM-5.2 [high]Model | 62.9%± 1.19% |
| 5 | OpenAIGPT-5.6 Sol [medium]Model | 62.3%± 1.83% |
| 6 | JunieJunieAgent | 61.8%± 0.54% |
| 7 | AnthropicClaude CodeAgent | 60.4%± 1.03% |
| 8 | OpenAICodexAgent | 58.0%± 1.29% |
| 9 | AnthropicSonnet 5 [high]Model | 56.8%± 0.94% |
| 10 | CursorCursorAgent | 51.7%± 0.84% |
Artificial Analysis composite index across coding, math, and reasoning benchmarks.
| # | Model | Score | tok/s | $/1M |
|---|---|---|---|---|
| 1 | Claude Opus 5.5 | 57.6 | 97 | $8.00 |
| 2 | Claude Sonnet 5.5 | 56 | 137 | $4.00 |
| 3 | Claude Fable 5.1 | 53.4 | 71 | $20.00 |
| 4 | GPT-6 Astra | 52.7 | 50 | $20.00 |
| 5 | Gemini 4 Argon | 52.6 | 0 | $4.00 |
| 6 | GPT-6.1 Sol | 51.8 | 57 | $4.00 |
| 7 | Claude Opus 5 | 50.8 | 0 | $10.00 |
| 8 | Claude Fable 5 | 49.6 | 0 | $20.00 |
| 9 | Muse Spark 1.3 | 48.1 | 123 | $2.00 |
| 10 | GPT-6 Sol | 47.6 | 0 | $4.00 |
Output tokens per second — higher is faster. Minimum intelligence score of 40.
| # | Model | tok/s |
|---|---|---|
| 1 | Ling 3.1 Flash | 218 |
| 2 | Gemini 3.8 Flash | 187 |
| 3 | Claude Sonnet 5.5 | 137 |
| 4 | Muse Spark 1.3 | 123 |
| 5 | GPT-5.6 Terra | 104 |
| 6 | Claude Opus 5.5 | 97 |
| 7 | Step 5 Preview | 90 |
| 8 | GLM-5.3 | 75 |
| 9 | Claude Fable 5.1 | 71 |
| 10 | Grok 4.7 | 68 |
Blended cost per 1M tokens (3:1 input/output) — lower is cheaper. Minimum intelligence score of 40.
| # | Model | $/1M |
|---|---|---|
| 1 | GLM 5.3 Flash | $0.237 |
| 2 | Ling 3.1 Flash | $0.45 |
| 3 | MiMo-V2.6-Pro | $0.544 |
| 4 | Step 5 Preview | $1.43 |
| 5 | Gemini 3.8 Flash | $1.50 |
| 6 | Muse Spark 1.3 | $2.00 |
| 7 | GLM-5.3 | $2.15 |
| 8 | Grok 4.7 | $3.00 |
| 9 | Qwen3.8 Max | $3.00 |
| 10 | Grok 4.6 | $3.00 |