The Inference Report

October 9, 2026

The SWE-rebench rankings remain static at the top tier, with AnthropicFable 5 holding 64.5% and the next five models clustered within 2.2 percentage points, all showing confidence intervals that overlap substantially with their neighbors. Artificial Analysis data, by contrast, reveals significant churn across its 470-entry leaderboard, particularly in the middle ranks where Granite 4.2 30B has shifted from position 163 to 193, and several DeepSeek variants have moved multiple positions, though the magnitude of these shifts often falls within typical measurement noise for scores in the 12-15 range. The two benchmarks measure different problem spaces: SWE-rebench evaluates real-world software engineering tasks with controlled conditions and reported error margins, while Artificial Analysis appears to sample a broader cross-section of models with no disclosed methodology or uncertainty quantification. At the SWE-rebench top, the stability reflects either genuine convergence in capability or plateau in discrimination power; the absence of new entries in the top ten positions suggests the frontier has settled. Artificial Analysis's fluidity at ranks 160-200 warrants skepticism about whether those shifts represent meaningful performance changes or noise from unmeasured experimental variance. Neither benchmark clarifies whether the apparent separation between Fable 5 (64.5%) and the field reflects a genuine algorithmic advance or saturation effects on the task distribution itself.

Cole Brennan

Daily rankings from SWE-rebench, a benchmark designed to fairly compare LLM capabilities on real-world software engineering tasks. Unlike other evaluations, it uses a standardized scaffolding for all models, continuously updates its dataset to prevent contamination, and runs each model five times to account for stochastic variance.

#ModelScore
1AnthropicFable 5 [high]Model64.5%± 1.41%
2GrokGrok 4.5 [high]Model63.8%± 0.60%
3AnthropicOpus 5 [high]Model63.4%± 1.35%
4Z.aiGLM-5.2 [high]Model62.9%± 1.19%
5OpenAIGPT-5.6 Sol [medium]Model62.3%± 1.83%
6JunieJunieAgent61.8%± 0.54%
7AnthropicClaude CodeAgent60.4%± 1.03%
8OpenAICodexAgent58.0%± 1.29%
9AnthropicSonnet 5 [high]Model56.8%± 0.94%
10CursorCursorAgent51.7%± 0.84%

Artificial Analysis composite index across coding, math, and reasoning benchmarks.

#ModelScoretok/s$/1M
1Claude Opus 5.557.697$8.00
2Claude Sonnet 5.556137$4.00
3Claude Fable 5.153.472$20.00
4GPT-6 Astra52.747$20.00
5Gemini 4 Argon52.60$4.00
6GPT-6.1 Sol51.858$4.00
7Claude Opus 550.80$10.00
8Claude Fable 549.60$20.00
9Muse Spark 1.348.1118$2.00
10GPT-6 Sol47.60$4.00

Output tokens per second — higher is faster. Minimum intelligence score of 40.

#Modeltok/s
1Claude Haiku 5.5235
2Ling 3.1 Flash215
3Claude Sonnet 5.5137
4Gemini 3.8 Flash131
5Muse Spark 1.3118
6GPT-5.6 Terra111
7Claude Opus 5.597
8Step 5 Preview90
9Claude Fable 5.172
10Grok 4.770

Blended cost per 1M tokens (3:1 input/output) — lower is cheaper. Minimum intelligence score of 40.

#Model$/1M
1Claude Haiku 5.5$0.20
2GLM-5.3-Flash$0.237
3Ling 3.1 Flash$0.45
4MiMo-V2.6-Pro$0.544
5Step 5 Preview$1.43
6Gemini 3.8 Flash$1.50
7Muse Spark 1.3$2.00
8GLM-5.3$2.15
9Grok 4.7$3.00
10Qwen3.8 Max$3.00