The Inference Report

October 7, 2026

The SWE-rebench rankings remain frozen while the Artificial Analysis leaderboard shuffled 38 entries, introducing Mistral Large 4 Preview at position 32 and cascading everything below it downward. On SWE-rebench, AnthropicFable 5 holds 64.5% with tight confidence intervals across the top tier, and the spread between first and tenth place spans 12.8 percentage points, suggesting real separation in coding task performance. The Artificial Analysis benchmark, by contrast, shows a different stratification: Claude Opus 5.5 leads at 57.6, but the top ten compress within a 5-point band, and models ranked 250 onward cluster between 9 and 5 percent with minimal differentiation. The methodology concern here is acute. SWE-rebench appears to test concrete problem-solving on real repository issues with controlled evaluation, whereas Artificial Analysis likely aggregates diverse metrics across broader capability domains where coding may be one signal among many. The two benchmarks rank models differently enough to suggest they measure distinct things: SWE-rebench's top performer (Fable 5 at 64.5%) ranks third on Artificial Analysis (53.4), while Artificial Analysis's leader (Opus 5.5 at 57.6) sits ninth on SWE-rebench (56.8%). This divergence is not noise. It reflects whether a benchmark prioritizes narrow coding competence or distributed capability. Neither ranking change signals meaningful progress without knowing whether the Mistral insertion reflects actual improvement or simply expanded roster coverage, and the SWE-rebench stability suggests either the evaluation is mature or updates are infrequent.

Cole Brennan

Daily rankings from SWE-rebench, a benchmark designed to fairly compare LLM capabilities on real-world software engineering tasks. Unlike other evaluations, it uses a standardized scaffolding for all models, continuously updates its dataset to prevent contamination, and runs each model five times to account for stochastic variance.

#ModelScore
1AnthropicFable 5 [high]Model64.5%± 1.41%
2GrokGrok 4.5 [high]Model63.8%± 0.60%
3AnthropicOpus 5 [high]Model63.4%± 1.35%
4Z.aiGLM-5.2 [high]Model62.9%± 1.19%
5OpenAIGPT-5.6 Sol [medium]Model62.3%± 1.83%
6JunieJunieAgent61.8%± 0.54%
7AnthropicClaude CodeAgent60.4%± 1.03%
8OpenAICodexAgent58.0%± 1.29%
9AnthropicSonnet 5 [high]Model56.8%± 0.94%
10CursorCursorAgent51.7%± 0.84%

Artificial Analysis composite index across coding, math, and reasoning benchmarks.

#ModelScoretok/s$/1M
1Claude Opus 5.557.697$8.00
2Claude Sonnet 5.556137$4.00
3Claude Fable 5.153.471$20.00
4GPT-6 Astra52.750$20.00
5Gemini 4 Argon52.60$4.00
6GPT-6.1 Sol51.857$4.00
7Claude Opus 550.80$10.00
8Claude Fable 549.60$20.00
9Muse Spark 1.348.1123$2.00
10GPT-6 Sol47.60$4.00

Output tokens per second — higher is faster. Minimum intelligence score of 40.

#Modeltok/s
1Ling 3.1 Flash218
2Gemini 3.8 Flash187
3Claude Sonnet 5.5137
4Muse Spark 1.3123
5GPT-5.6 Terra104
6Claude Opus 5.597
7Step 5 Preview90
8GLM-5.375
9Claude Fable 5.171
10Grok 4.768

Blended cost per 1M tokens (3:1 input/output) — lower is cheaper. Minimum intelligence score of 40.

#Model$/1M
1GLM 5.3 Flash$0.237
2Ling 3.1 Flash$0.45
3MiMo-V2.6-Pro$0.544
4Step 5 Preview$1.43
5Gemini 3.8 Flash$1.50
6Muse Spark 1.3$2.00
7GLM-5.3$2.15
8Grok 4.7$3.00
9Qwen3.8 Max$3.00
10Grok 4.6$3.00