The Inference Report

October 8, 2026

On the SWE-rebench coding benchmark, the top tier remains frozen: AnthropicFable 5 holds 64.5% with Grok 4.5 and Anthropic Opus 5 trailing at 63.8% and 63.4% respectively, all within their established confidence intervals and showing no meaningful movement from the previous cycle. The Artificial Analysis general-purpose benchmark tells a different story, with Claude Opus 5.5 now leading at 57.6 (up from an unranked position), displacing Claude Opus 5 to seventh place and shuffling the top ten substantially: Claude Sonnet 5.5 enters at 56.0, Claude Fable 5.1 at 53.4, and the GPT-6 variants (Astra at 52.7, Sol at 51.8) consolidate ground previously held by older models. This reshuffling reflects the release cadence of frontier models rather than methodological discovery. The coding benchmark's stability is notable given the variance margins, most top performers fall within 0.6 to 1.83 percentage points of uncertainty, yet no reordering occurs, suggesting either that the task has reached saturation for leading systems or that confidence intervals genuinely constrain detectability of small gains. The Artificial Analysis list shows more flux: Claude Haiku 5.5 and GLM-5.3-Flash appear as new entries at positions 19 and 21, while GLM 5.3 Flash dropped entirely, indicating version-specific performance variation rather than sustained architectural progress. Taken together, the benchmarks suggest the frontier has plateaued on structured code tasks while general capabilities continue to shift with model releases, a pattern consistent with specialization rather than across-the-board improvement.

Cole Brennan

Daily rankings from SWE-rebench, a benchmark designed to fairly compare LLM capabilities on real-world software engineering tasks. Unlike other evaluations, it uses a standardized scaffolding for all models, continuously updates its dataset to prevent contamination, and runs each model five times to account for stochastic variance.

#ModelScore
1AnthropicFable 5 [high]Model64.5%± 1.41%
2GrokGrok 4.5 [high]Model63.8%± 0.60%
3AnthropicOpus 5 [high]Model63.4%± 1.35%
4Z.aiGLM-5.2 [high]Model62.9%± 1.19%
5OpenAIGPT-5.6 Sol [medium]Model62.3%± 1.83%
6JunieJunieAgent61.8%± 0.54%
7AnthropicClaude CodeAgent60.4%± 1.03%
8OpenAICodexAgent58.0%± 1.29%
9AnthropicSonnet 5 [high]Model56.8%± 0.94%
10CursorCursorAgent51.7%± 0.84%

Artificial Analysis composite index across coding, math, and reasoning benchmarks.

#ModelScoretok/s$/1M
1Claude Opus 5.557.698$8.00
2Claude Sonnet 5.556133$4.00
3Claude Fable 5.153.467$20.00
4GPT-6 Astra52.748$20.00
5Gemini 4 Argon52.60$4.00
6GPT-6.1 Sol51.857$4.00
7Claude Opus 550.80$10.00
8Claude Fable 549.60$20.00
9Muse Spark 1.348.1103$2.00
10GPT-6 Sol47.60$4.00

Output tokens per second — higher is faster. Minimum intelligence score of 40.

#Modeltok/s
1Claude Haiku 5.5240
2Ling 3.1 Flash215
3Gemini 3.8 Flash145
4Claude Sonnet 5.5133
5Muse Spark 1.3103
6GPT-5.6 Terra99
7Claude Opus 5.598
8Step 5 Preview89
9GLM-5.370
10Claude Fable 5.167

Blended cost per 1M tokens (3:1 input/output) — lower is cheaper. Minimum intelligence score of 40.

#Model$/1M
1Claude Haiku 5.5$0.20
2GLM-5.3-Flash$0.237
3Ling 3.1 Flash$0.45
4MiMo-V2.6-Pro$0.544
5Step 5 Preview$1.43
6Gemini 3.8 Flash$1.50
7Muse Spark 1.3$2.00
8GLM-5.3$2.15
9Grok 4.7$3.00
10Qwen3.8 Max$3.00