The Inference Report

October 4, 2026

The SWE-rebench leaderboard shows complete stability across the top tier, with AnthropicFable 5 holding 64.5% ± 1.41%, GrokGrok 4.5 at 63.8% ± 0.60%, and AnthropicOpus 5 at 63.4% ± 1.35%, identical to the previous cycle. The confidence intervals are tight enough that these positions reflect genuine performance differences rather than noise, yet the lack of movement suggests either that the test set has reached saturation among frontier models or that incremental improvements are genuinely stalled. The gap between rank 5 (OpenAIGPT-5.6 Sol at 62.3% ± 1.83%) and rank 14 (DeepSeekDeepSeek-V4 Pro at 40.2% ± 1.29%) is substantial and persistent, indicating a clear separation between high-capability and mid-tier systems. On the Artificial Analysis benchmark, the data shows comprehensive churn throughout the ranking, with Ling 3.1 Flash entering at rank 22 as a new entry, pushing all subsequent models down by one position. This suggests Artificial Analysis tracks a broader ecosystem with more frequent model releases, whereas SWE-rebench's frozen leaderboard may reflect either fewer new submissions or the difficulty of breaking into a plateau dominated by three Anthropic variants. Neither benchmark shows evidence of methodological weakness in the raw numbers, though SWE-rebench's lack of movement warrants scrutiny into whether the test remains sensitive to real capability differences or has become a ceiling effect masking progress in specialized domains.

Cole Brennan

Daily rankings from SWE-rebench, a benchmark designed to fairly compare LLM capabilities on real-world software engineering tasks. Unlike other evaluations, it uses a standardized scaffolding for all models, continuously updates its dataset to prevent contamination, and runs each model five times to account for stochastic variance.

#ModelScore
1AnthropicFable 5 [high]Model64.5%± 1.41%
2GrokGrok 4.5 [high]Model63.8%± 0.60%
3AnthropicOpus 5 [high]Model63.4%± 1.35%
4Z.aiGLM-5.2 [high]Model62.9%± 1.19%
5OpenAIGPT-5.6 Sol [medium]Model62.3%± 1.83%
6JunieJunieAgent61.8%± 0.54%
7AnthropicClaude CodeAgent60.4%± 1.03%
8OpenAICodexAgent58.0%± 1.29%
9AnthropicSonnet 5 [high]Model56.8%± 0.94%
10CursorCursorAgent51.7%± 0.84%

Artificial Analysis composite index across coding, math, and reasoning benchmarks.

#ModelScoretok/s$/1M
1Claude Opus 5.557.697$8.00
2Claude Sonnet 5.556138$4.00
3Claude Fable 5.153.469$20.00
4GPT-6 Astra52.763$20.00
5Gemini 4 Argon52.60$4.00
6GPT-6.1 Sol51.859$4.00
7Claude Opus 550.80$10.00
8Claude Fable 549.60$20.00
9Muse Spark 1.348.1162$2.00
10GPT-6 Sol47.6101$4.00

Output tokens per second — higher is faster. Minimum intelligence score of 40.

#Modeltok/s
1Gemini 3.8 Flash247
2Ling 3.1 Flash216
3Muse Spark 1.3162
4Claude Sonnet 5.5138
5GPT-5.6 Terra112
6GPT-6 Sol101
7Claude Opus 5.597
8Step 5 Preview84
9Grok 4.780
10GLM-5.375

Blended cost per 1M tokens (3:1 input/output) — lower is cheaper. Minimum intelligence score of 40.

#Model$/1M
1GLM 5.3 Flash$0.237
2Ling 3.1 Flash$0.45
3MiMo-V2.6-Pro$0.544
4Step 5 Preview$1.43
5Gemini 3.8 Flash$1.50
6Muse Spark 1.3$2.00
7GLM-5.3$2.15
8Grok 4.7$3.00
9Qwen3.8 Max$3.00
10Grok 4.6$3.00