On SWE-rebench, the top tier remains stable: OpenAI's gpt-5.5-2026-04-23-xhighModel holds 62.7% (±0.91%), followed by JunieJunieAgent at 61.6% (±0.64%) and OpenAICodexAgent at 60.4% (±1.37%), with confidence intervals that do not overlap meaningfully enough to suggest ranking instability. The spread between positions 1 and 24 spans from 62.7% to 16.5%, a 46-point gap that reflects genuine capability stratification rather than noise. What merits scrutiny is the benchmark's design: SWE-rebench evaluates agent-based code completion on real software engineering tasks, a more applied setting than many token-prediction benchmarks, yet the evaluation methodology, particularly how task success is defined and whether partial credit is awarded, remains underspecified in the data provided. On Artificial Analysis, the roster has undergone substantial motion: Gemini 3.6 Flash enters at rank 15 (50.1 points), displacing prior entries downward, while Trinity Large Thinking drops from rank 112 (24.5 points) to rank 157 (18.2 points), a 6.3-point decline suggesting either a recalibration of the benchmark or a shift in evaluation conditions. The Artificial Analysis leaderboard exhibits denser clustering in the 30-50 point range than SWE-rebench, with many models separated by tenths of a point; this density raises questions about whether differences of 0.3-0.5 points reflect reproducible capability gaps or measurement variance. Neither benchmark's methodology clarifies whether scores are averaged across multiple runs, how task selection bias is controlled, or whether confidence intervals reflect statistical significance or merely reported uncertainty bounds. The two benchmarks show weak rank correlation at the extremes, Claude Fable 5 ranks first on Artificial Analysis (59.9) but does not appear on the SWE-rebench top 24, while gpt-5.5-xhigh dominates SWE-rebench but ranks sixth on Artificial Analysis (54.8), suggesting they measure distinct problem classes or that one benchmark's evaluation is more sensitive to architectural choices that do not generalize.
Cole Brennan
Daily rankings from SWE-rebench, a benchmark designed to fairly compare LLM capabilities on real-world software engineering tasks. Unlike other evaluations, it uses a standardized scaffolding for all models, continuously updates its dataset to prevent contamination, and runs each model five times to account for stochastic variance.
| # | Model | Score |
|---|---|---|
| 1 | OpenAIgpt-5.5-2026-04-23-xhighModel | 62.7%± 0.91% |
| 2 | JunieJunieAgent | 61.6%± 0.64% |
| 3 | OpenAICodexAgent | 60.4%± 1.37% |
| 4 | AnthropicClaude CodeAgent | 59.6%± 1.98% |
| 5 | OpenAIgpt-5.5-2026-04-23-mediumModel | 58.9%± 0.78% |
| 6 | AnthropicClaude Opus 4.8-xhighModel | 56.5%± 1.20% |
| 7 | OpenAIgpt-5.4-2026-03-05-mediumModel | 54.9%± 1.02% |
| 8 | AnthropicClaude Opus 4.7-highModel | 53.1%± 1.45% |
| 9 | CursorCursorAgent | 53.0%± 0.53% |
| 10 | AnthropicClaude Sonnet 4.6Model | 51.3%± 0.55% |
Artificial Analysis composite index across coding, math, and reasoning benchmarks.
| # | Model | Score | tok/s | $/1M |
|---|---|---|---|---|
| 1 | Claude Fable 5 | 59.9 | 70 | $20.00 |
| 2 | GPT-5.6 Sol | 58.9 | 67 | $11.25 |
| 3 | Kimi K3 | 57.1 | 38 | $6.00 |
| 4 | Claude Opus 4.8 | 55.7 | 64 | $10.00 |
| 5 | GPT-5.6 Terra | 55 | 149 | $5.63 |
| 6 | GPT-5.5 | 54.8 | 92 | $11.25 |
| 7 | Grok 4.5 | 53.8 | 73 | $3.00 |
| 8 | Claude Opus 4.7 | 53.5 | 61 | $10.00 |
| 9 | Claude Sonnet 5 | 53.4 | 86 | $4.00 |
| 10 | GPT-5.4 | 51.4 | 149 | $5.63 |
Output tokens per second — higher is faster. Minimum intelligence score of 40.
| # | Model | tok/s |
|---|---|---|
| 1 | Gemini 3.6 Flash | 311 |
| 2 | Gemini 3.5 Flash | 283 |
| 3 | Qwen3.7 Max | 207 |
| 4 | GPT-5.6 Luna | 203 |
| 5 | GLM-5.2 | 191 |
| 6 | GPT-5.4 mini | 179 |
| 7 | GPT-5.2 Codex | 162 |
| 8 | GPT-5.6 Terra | 149 |
| 9 | GPT-5.4 | 149 |
| 10 | Gemini 3.1 Pro Preview | 135 |
Blended cost per 1M tokens (3:1 input/output) — lower is cheaper. Minimum intelligence score of 40.
| # | Model | $/1M |
|---|---|---|
| 1 | DeepSeek V4 Flash | $0.175 |
| 2 | Hy3 | $0.25 |
| 3 | MiniMax-M3 | $0.525 |
| 4 | DeepSeek V4 Pro | $0.544 |
| 5 | MiMo-V2.5-Pro | $0.544 |
| 6 | Nex-N2-Pro | $1.00 |
| 7 | GPT-5.4 mini | $1.69 |
| 8 | Kimi K2.6 | $1.71 |
| 9 | Kimi K2.7 Code | $1.71 |
| 10 | Muse Spark 1.1 | $2.00 |