The Inference Report

September 30, 2026

The SWE-rebench standings remain frozen at the top, with AnthropicFable 5 holding 64.5% and the next four positions unchanged, but the Artificial Analysis rankings show material churn across the full leaderboard. GPT-6.1 Sol enters at #5 (51.8), pushing Claude Opus 5 down one slot to #6, while lower-ranked models experience more dramatic shifts: JT-4.1 Flash 236B A21B surges from #62 (27.3) to #39 (33.9), a gain of 6.6 points that stands out as the most substantial movement in the visible range. Ling 3.0 Tiny drops sharply from #155 (15.3) to #216 (11.1), losing 4.2 points, suggesting the benchmark may be reweighting reasoning or coding tasks where that model underperforms. Most other movements cluster between one and three positions, typical churn in a crowded field. The Artificial Analysis benchmark itself lacks transparency about evaluation methodology, making it difficult to assess whether these shifts reflect genuine capability changes or volatility in how tasks are selected and scored. SWE-rebench's stability at the top suggests either convergence among leading models or that its test set has reached saturation; the absence of new entries in the top 17 over this cycle is notable. Without visibility into what changed in Artificial Analysis's test construction or weighting, the practical significance of these rankings remains unclear.

Cole Brennan

Daily rankings from SWE-rebench, a benchmark designed to fairly compare LLM capabilities on real-world software engineering tasks. Unlike other evaluations, it uses a standardized scaffolding for all models, continuously updates its dataset to prevent contamination, and runs each model five times to account for stochastic variance.

#ModelScore
1AnthropicFable 5 [high]Model64.5%± 1.41%
2GrokGrok 4.5 [high]Model63.8%± 0.60%
3AnthropicOpus 5 [high]Model63.4%± 1.35%
4Z.aiGLM-5.2 [high]Model62.9%± 1.19%
5OpenAIGPT-5.6 Sol [medium]Model62.3%± 1.83%
6JunieJunieAgent61.8%± 0.54%
7AnthropicClaude CodeAgent60.4%± 1.03%
8OpenAICodexAgent58.0%± 1.29%
9AnthropicSonnet 5 [high]Model56.8%± 0.94%
10CursorCursorAgent51.7%± 0.84%

Artificial Analysis composite index across coding, math, and reasoning benchmarks.

#ModelScoretok/s$/1M
1Claude Opus 5.557.696$8.00
2Claude Sonnet 5.556145$4.00
3Claude Fable 5.153.469$20.00
4GPT-6 Astra52.755$20.00
5GPT-6.1 Sol51.880$4.00
6Claude Opus 550.80$10.00
7Claude Fable 549.60$20.00
8Muse Spark 1.348.1189$2.00
9GPT-6 Sol47.580$4.00
10GPT-5.6 Sol470$8.00

Output tokens per second — higher is faster. Minimum intelligence score of 40.

#Modeltok/s
1Gemini 3.8 Flash222
2Muse Spark 1.3189
3Claude Sonnet 5.5145
4GPT-5.6 Terra108
5Claude Opus 5.596
6Step 5 Preview86
7Grok 4.783
8GPT-6.1 Sol80
9GPT-6 Sol80
10GLM-5.370

Blended cost per 1M tokens (3:1 input/output) — lower is cheaper. Minimum intelligence score of 40.

#Model$/1M
1GLM 5.3 Flash$0.237
2MiMo-V2.6-Pro$0.544
3Step 5 Preview$1.43
4Gemini 3.8 Flash$1.50
5Muse Spark 1.3$2.00
6GLM-5.3$2.15
7Grok 4.7$3.00
8Qwen3.8 Max$3.00
9Grok 4.6$3.00
10Claude Sonnet 5.5$4.00