Daily rankings from SWE-rebench, a benchmark designed to fairly compare LLM capabilities on real-world software engineering tasks. Unlike other evaluations, it uses a standardized scaffolding for all models, continuously updates its dataset to prevent contamination, and runs each model five times to account for stochastic variance.
| # | Model | Score |
|---|---|---|
| 1 | AnthropicFable 5 [high]Model | 64.5%± 1.41% |
| 2 | GrokGrok 4.5 [high]Model | 63.8%± 0.60% |
| 3 | AnthropicOpus 5 [high]Model | 63.4%± 1.35% |
| 4 | Z.aiGLM-5.2 [high]Model | 62.9%± 1.19% |
| 5 | OpenAIGPT-5.6 Sol [medium]Model | 62.3%± 1.83% |
| 6 | JunieJunieAgent | 61.8%± 0.54% |
| 7 | AnthropicClaude CodeAgent | 60.4%± 1.03% |
| 8 | OpenAICodexAgent | 58.0%± 1.29% |
| 9 | AnthropicSonnet 5 [high]Model | 56.8%± 0.94% |
| 10 | CursorCursorAgent | 51.7%± 0.84% |
Artificial Analysis composite index across coding, math, and reasoning benchmarks.
| # | Model | Score | tok/s | $/1M |
|---|---|---|---|---|
| 1 | Claude Fable 5.1 | 53.4 | 72 | $20.00 |
| 2 | GPT-6 Astra | 52.7 | 69 | $20.00 |
| 3 | Claude Opus 5 | 50.8 | 54 | $10.00 |
| 4 | Claude Fable 5 | 49.6 | 0 | $20.00 |
| 5 | Muse Spark 1.3 | 48.1 | 279 | $2.00 |
| 6 | GPT-5.6 Sol | 47 | 72 | $8.00 |
| 7 | Qwen3.8 Max | 45.4 | 41 | $3.00 |
| 8 | GLM-5.3 | 44.8 | 71 | $2.15 |
| 9 | Grok 4.6 | 44.3 | 63 | $3.00 |
| 10 | Step 5 Preview | 43.7 | 93 | $1.43 |
Output tokens per second — higher is faster. Minimum intelligence score of 40.
| # | Model | tok/s |
|---|---|---|
| 1 | Gemini 3.8 Flash | 342 |
| 2 | Muse Spark 1.3 | 279 |
| 3 | GPT-5.6 Terra | 106 |
| 4 | Step 5 Preview | 93 |
| 5 | GLM 5.3 Flash | 92 |
| 6 | Claude Fable 5.1 | 72 |
| 7 | GPT-5.6 Sol | 72 |
| 8 | GLM-5.3 | 71 |
| 9 | GPT-6 Astra | 69 |
| 10 | Grok 4.6 | 63 |
Blended cost per 1M tokens (3:1 input/output) — lower is cheaper. Minimum intelligence score of 40.
| # | Model | $/1M |
|---|---|---|
| 1 | GLM 5.3 Flash | $0.237 |
| 2 | Step 5 Preview | $1.43 |
| 3 | Gemini 3.8 Flash | $1.50 |
| 4 | Muse Spark 1.3 | $2.00 |
| 5 | GLM-5.3 | $2.15 |
| 6 | Qwen3.8 Max | $3.00 |
| 7 | Grok 4.6 | $3.00 |
| 8 | GPT-5.6 Terra | $4.50 |
| 9 | Kimi K3 | $6.00 |
| 10 | GPT-5.6 Sol | $8.00 |