The Inference Report

October 1, 2026

The SWE-rebench leaderboard holds steady at the top with Anthropic Fable 5 maintaining 64.5% ± 1.41%, followed by Grok 4.5 at 63.8% ± 0.60% and Anthropic Opus 5 at 63.4% ± 1.35%, showing no movement in the top tier. Notably, the confidence intervals on SWE-rebench remain tight across the tested models, suggesting controlled evaluation conditions, though the gap between the leader and tenth-place Cursor Agent (51.7% ± 0.84%) spans 12.8 percentage points, indicating clear stratification in code-solving capability. On the Artificial Analysis benchmark, Gemini 4 Argon enters at number five with 52.6 points as a new entrant, displacing prior entries by one position, while GPT-6 Luna climbed from #34 to #33 with a score increase from 37.3 to 38.1. Solar Mini 4 also appears as a new entry at #91 with 24.1 points. The broader Artificial Analysis leaderboard shows tight clustering in the 5 to 10 point range across positions 380 to 450, where parameter count and model size variations produce minimal score differentiation, raising questions about whether that benchmark's resolution distinguishes meaningfully between smaller models. The SWE-rebench results align with the intuition that coding tasks reward architectural sophistication and training data quality more than raw parameters, whereas Artificial Analysis's crowded lower tier suggests saturation or floor effects in its evaluation methodology.

Cole Brennan

Daily rankings from SWE-rebench, a benchmark designed to fairly compare LLM capabilities on real-world software engineering tasks. Unlike other evaluations, it uses a standardized scaffolding for all models, continuously updates its dataset to prevent contamination, and runs each model five times to account for stochastic variance.

#ModelScore
1AnthropicFable 5 [high]Model64.5%± 1.41%
2GrokGrok 4.5 [high]Model63.8%± 0.60%
3AnthropicOpus 5 [high]Model63.4%± 1.35%
4Z.aiGLM-5.2 [high]Model62.9%± 1.19%
5OpenAIGPT-5.6 Sol [medium]Model62.3%± 1.83%
6JunieJunieAgent61.8%± 0.54%
7AnthropicClaude CodeAgent60.4%± 1.03%
8OpenAICodexAgent58.0%± 1.29%
9AnthropicSonnet 5 [high]Model56.8%± 0.94%
10CursorCursorAgent51.7%± 0.84%

Artificial Analysis composite index across coding, math, and reasoning benchmarks.

#ModelScoretok/s$/1M
1Claude Opus 5.557.696$8.00
2Claude Sonnet 5.556145$4.00
3Claude Fable 5.153.470$20.00
4GPT-6 Astra52.754$20.00
5Gemini 4 Argon52.60$4.00
6GPT-6.1 Sol51.873$4.00
7Claude Opus 550.80$10.00
8Claude Fable 549.60$20.00
9Muse Spark 1.348.1177$2.00
10GPT-6 Sol47.683$4.00

Output tokens per second — higher is faster. Minimum intelligence score of 40.

#Modeltok/s
1Gemini 3.8 Flash238
2Muse Spark 1.3177
3Claude Sonnet 5.5145
4GPT-5.6 Terra106
5Claude Opus 5.596
6Step 5 Preview88
7GPT-6 Sol83
8Grok 4.783
9GPT-6.1 Sol73
10Claude Fable 5.170

Blended cost per 1M tokens (3:1 input/output) — lower is cheaper. Minimum intelligence score of 40.

#Model$/1M
1GLM 5.3 Flash$0.237
2MiMo-V2.6-Pro$0.544
3Step 5 Preview$1.43
4Gemini 3.8 Flash$1.50
5Muse Spark 1.3$2.00
6GLM-5.3$2.15
7Grok 4.7$3.00
8Qwen3.8 Max$3.00
9Grok 4.6$3.00
10Claude Sonnet 5.5$4.00