The Inference Report

September 25, 2026

The SWE-rebench leaderboard shows no movement at the top nine positions, with AnthropicFable 5 maintaining 64.5% and the tier below it unchanged through OpenAICodexAgent at 58.0%. The consistency across these scores, identical to the previous snapshot, reflects either a stabilization in performance or a lack of new evaluations on this benchmark. The confidence intervals remain tight enough that the ranking would require meaningful score shifts to alter, particularly in the 0.5-1.4 percentage point range where several models cluster. On Artificial Analysis, which measures a broader set of coding tasks, the ordering shows minor churn in the 45-67 position band, with Apodex 1.1 dropping from #47 to #66 on a score decline from 30.4 to 26.4, while models like Claude Sonnet 4.6 and Gemini 3.1 Pro Preview each gained one position. These shifts are modest and do not reflect fundamental reordering of capability tiers. The divergence between benchmarks is notable: models ranked highest on SWE-rebench (Fable 5, Grok 4.5, Opus 5) occupy positions 5, 27, and 1 respectively on Artificial Analysis, suggesting the benchmarks measure different aspects of coding performance or that SWE-rebench may emphasize repository-level software engineering tasks where agentic orchestration matters more than raw instruction-following ability. Without methodological transparency on how SWE-rebench constructs its evaluation, particularly whether it tests tool use, multi-step reasoning, or purely code generation, the stability of its rankings is difficult to interpret as either meaningful stasis or measurement artifact.

Cole Brennan

Daily rankings from SWE-rebench, a benchmark designed to fairly compare LLM capabilities on real-world software engineering tasks. Unlike other evaluations, it uses a standardized scaffolding for all models, continuously updates its dataset to prevent contamination, and runs each model five times to account for stochastic variance.

#ModelScore
1AnthropicFable 5 [high]Model64.5%± 1.41%
2GrokGrok 4.5 [high]Model63.8%± 0.60%
3AnthropicOpus 5 [high]Model63.4%± 1.35%
4Z.aiGLM-5.2 [high]Model62.9%± 1.19%
5OpenAIGPT-5.6 Sol [medium]Model62.3%± 1.83%
6JunieJunieAgent61.8%± 0.54%
7AnthropicClaude CodeAgent60.4%± 1.03%
8OpenAICodexAgent58.0%± 1.29%
9AnthropicSonnet 5 [high]Model56.8%± 0.94%
10CursorCursorAgent51.7%± 0.84%

Artificial Analysis composite index across coding, math, and reasoning benchmarks.

#ModelScoretok/s$/1M
1Claude Opus 5.557.6102$8.00
2Claude Fable 5.153.470$20.00
3GPT-6 Astra52.754$20.00
4Claude Opus 550.851$10.00
5Claude Fable 549.60$20.00
6Muse Spark 1.348.1191$2.00
7GPT-6 Sol47.5107$4.00
8GPT-5.6 Sol4758$8.00
9Grok 4.746.451$3.00
10MiMo-V2.6-Pro46.337$0.544

Output tokens per second — higher is faster. Minimum intelligence score of 40.

#Modeltok/s
1Gemini 3.8 Flash267
2Muse Spark 1.3191
3GPT-6 Sol107
4Claude Opus 5.5102
5GPT-5.6 Terra87
6Step 5 Preview73
7Claude Fable 5.170
8Grok 4.665
9GLM-5.364
10GPT-5.6 Sol58

Blended cost per 1M tokens (3:1 input/output) — lower is cheaper. Minimum intelligence score of 40.

#Model$/1M
1GLM 5.3 Flash$0.237
2MiMo-V2.6-Pro$0.544
3Step 5 Preview$1.43
4Gemini 3.8 Flash$1.50
5Muse Spark 1.3$2.00
6GLM-5.3$2.15
7Grok 4.7$3.00
8Qwen3.8 Max$3.00
9Grok 4.6$3.00
10GPT-6 Sol$4.00