Daily rankings from SWE-rebench, a benchmark designed to fairly compare LLM capabilities on real-world software engineering tasks. Unlike other evaluations, it uses a standardized scaffolding for all models, continuously updates its dataset to prevent contamination, and runs each model five times to account for stochastic variance.
| # | Model | Score |
|---|---|---|
| 1 | AnthropicFable 5 [high]Model | 64.5%± 1.41% |
| 2 | GrokGrok 4.5 [high]Model | 63.8%± 0.60% |
| 3 | AnthropicOpus 5 [high]Model | 63.4%± 1.35% |
| 4 | Z.aiGLM-5.2 [high]Model | 62.9%± 1.19% |
| 5 | OpenAIGPT-5.6 Sol [medium]Model | 62.3%± 1.83% |
| 6 | JunieJunieAgent | 61.8%± 0.54% |
| 7 | AnthropicClaude CodeAgent | 60.4%± 1.03% |
| 8 | OpenAICodexAgent | 58.0%± 1.29% |
| 9 | AnthropicSonnet 5 [high]Model | 56.8%± 0.94% |
| 10 | CursorCursorAgent | 51.7%± 0.84% |
Artificial Analysis composite index across coding, math, and reasoning benchmarks.
| # | Model | Score | tok/s | $/1M |
|---|---|---|---|---|
| 1 | Claude Opus 5 | 63.1 | 48 | $10.00 |
| 2 | Claude Fable 5 | 62.1 | 58 | $20.00 |
| 3 | GPT-5.6 Sol | 60.9 | 67 | $11.25 |
| 4 | Kimi K3 | 59.7 | 37 | $6.00 |
| 5 | Qwen3.8 Max | 58.1 | 80 | $3.00 |
| 6 | Claude Opus 4.8 | 57.3 | 0 | $10.00 |
| 7 | Muse Spark 1.2 | 56.8 | 0 | $2.00 |
| 8 | GPT-5.6 Terra | 56.6 | 116 | $4.50 |
| 9 | GPT-5.5 | 56.3 | 0 | $11.25 |
| 10 | Grok 4.5 | 55.8 | 48 | $3.00 |
Output tokens per second — higher is faster. Minimum intelligence score of 40.
| # | Model | tok/s |
|---|---|---|
| 1 | Gemini 3.5 Flash | 226 |
| 2 | Qwen3.7 Max | 197 |
| 3 | Muse Spark 1.1 | 192 |
| 4 | Gemini 3.6 Flash | 184 |
| 5 | GPT-5.6 Luna | 161 |
| 6 | Inkling Small | 144 |
| 7 | Gemini 3.1 Pro Preview | 126 |
| 8 | Nex-N2-Pro | 124 |
| 9 | GPT-5.6 Terra | 116 |
| 10 | GPT-5.3 Codex | 112 |
Blended cost per 1M tokens (3:1 input/output) — lower is cheaper. Minimum intelligence score of 40.
| # | Model | $/1M |
|---|---|---|
| 1 | DeepSeek V4 Flash | $0.168 |
| 2 | DeepSeek V4 Flash 0731 | $0.175 |
| 3 | Hy3 | $0.241 |
| 4 | GPT-5.6 Luna | $0.45 |
| 5 | MiniMax-M3 | $0.525 |
| 6 | Inkling Small | $0.525 |
| 7 | DeepSeek V4 Pro | $0.544 |
| 8 | MiMo-V2.5-Pro | $0.544 |
| 9 | Nex-N2-Pro | $1.00 |
| 10 | Qwen3.6 Plus | $1.13 |