The Inference Report

September 27, 2026

The SWE-rebench leaderboard shows no movement since the previous update, with the top tier locked in place: AnthropicFable 5 holds 64.5% plus or minus 1.41%, GrokGrok 4.5 remains at 63.8% plus or minus 0.60%, and AnthropicOpus 5 stays at 63.4% plus or minus 1.35%. The confidence intervals are tight enough that the ranking reflects real separation, particularly between the first and second positions where the gap exceeds the combined uncertainty bands. On the Artificial Analysis benchmark, movement is cosmetic rather than structural. MiMo-V2.6-Pro enters at position 30 with a score of 46.3, displacing the previous entry downward, but this represents catalog expansion rather than performance shifts among established models. The top performers remain unchanged: Claude Opus 5.5 at 57.6, Claude Fable 5.1 at 53.4, GPT-6 Astra at 52.7. The SWE-rebench methodology tests code generation against real software engineering problems with defined pass rates, making its stability meaningful, while the Artificial Analysis benchmark's evaluation criteria are less transparent from the data provided. The absence of score changes across both benchmarks suggests either genuine plateau at the frontier or evaluation cycles that have not yet captured recent model iterations. Without information about when these measurements were taken or how frequently they update, it remains unclear whether this stasis reflects convergence or lag in assessment.

Cole Brennan

Daily rankings from SWE-rebench, a benchmark designed to fairly compare LLM capabilities on real-world software engineering tasks. Unlike other evaluations, it uses a standardized scaffolding for all models, continuously updates its dataset to prevent contamination, and runs each model five times to account for stochastic variance.

#ModelScore
1AnthropicFable 5 [high]Model64.5%± 1.41%
2GrokGrok 4.5 [high]Model63.8%± 0.60%
3AnthropicOpus 5 [high]Model63.4%± 1.35%
4Z.aiGLM-5.2 [high]Model62.9%± 1.19%
5OpenAIGPT-5.6 Sol [medium]Model62.3%± 1.83%
6JunieJunieAgent61.8%± 0.54%
7AnthropicClaude CodeAgent60.4%± 1.03%
8OpenAICodexAgent58.0%± 1.29%
9AnthropicSonnet 5 [high]Model56.8%± 0.94%
10CursorCursorAgent51.7%± 0.84%

Artificial Analysis composite index across coding, math, and reasoning benchmarks.

#ModelScoretok/s$/1M
1Claude Opus 5.557.699$8.00
2Claude Fable 5.153.471$20.00
3GPT-6 Astra52.760$20.00
4Claude Opus 550.80$10.00
5Claude Fable 549.60$20.00
6Muse Spark 1.348.1161$2.00
7GPT-6 Sol47.587$4.00
8GPT-5.6 Sol470$8.00
9Grok 4.746.471$3.00
10MiMo-V2.6-Pro46.339$0.544

Output tokens per second — higher is faster. Minimum intelligence score of 40.

#Modeltok/s
1Gemini 3.8 Flash298
2Muse Spark 1.3161
3Claude Opus 5.599
4GPT-5.6 Terra90
5GPT-6 Sol87
6Grok 4.679
7GLM-5.378
8Step 5 Preview74
9Claude Fable 5.171
10Grok 4.771

Blended cost per 1M tokens (3:1 input/output) — lower is cheaper. Minimum intelligence score of 40.

#Model$/1M
1GLM 5.3 Flash$0.237
2MiMo-V2.6-Pro$0.544
3Step 5 Preview$1.43
4Gemini 3.8 Flash$1.50
5Muse Spark 1.3$2.00
6GLM-5.3$2.15
7Grok 4.7$3.00
8Qwen3.8 Max$3.00
9Grok 4.6$3.00
10GPT-6 Sol$4.00