The Inference Report

October 10, 2026

Daily rankings from SWE-rebench, a benchmark designed to fairly compare LLM capabilities on real-world software engineering tasks. Unlike other evaluations, it uses a standardized scaffolding for all models, continuously updates its dataset to prevent contamination, and runs each model five times to account for stochastic variance.

#ModelScore
1AnthropicFable 5 [high]Model64.5%± 1.41%
2GrokGrok 4.5 [high]Model63.8%± 0.60%
3AnthropicOpus 5 [high]Model63.4%± 1.35%
4Z.aiGLM-5.2 [high]Model62.9%± 1.19%
5OpenAIGPT-5.6 Sol [medium]Model62.3%± 1.83%
6JunieJunieAgent61.8%± 0.54%
7AnthropicClaude CodeAgent60.4%± 1.03%
8OpenAICodexAgent58.0%± 1.29%
9AnthropicSonnet 5 [high]Model56.8%± 0.94%
10CursorCursorAgent51.7%± 0.84%

Artificial Analysis composite index across coding, math, and reasoning benchmarks.

#ModelScoretok/s$/1M
1Claude Opus 5.557.697$8.00
2Claude Sonnet 5.556137$4.00
3Claude Fable 5.153.469$20.00
4GPT-6 Astra52.750$20.00
5Gemini 4 Argon52.60$4.00
6GPT-6.1 Sol51.861$4.00
7Claude Opus 550.80$10.00
8Claude Fable 549.60$20.00
9Muse Spark 1.348.1119$2.00
10GPT-6 Sol47.60$4.00

Output tokens per second — higher is faster. Minimum intelligence score of 40.

#Modeltok/s
1Claude Haiku 5.5237
2Ling 3.1 Flash213
3Claude Sonnet 5.5137
4Gemini 3.8 Flash127
5Muse Spark 1.3119
6GPT-5.6 Terra113
7Claude Opus 5.597
8Step 5 Preview90
9Grok 4.774
10Claude Fable 5.169

Blended cost per 1M tokens (3:1 input/output) — lower is cheaper. Minimum intelligence score of 40.

#Model$/1M
1Claude Haiku 5.5$0.20
2GLM-5.3-Flash$0.237
3Ling 3.1 Flash$0.45
4MiMo-V2.6-Pro$0.544
5Step 5 Preview$1.43
6Gemini 3.8 Flash$1.50
7Muse Spark 1.3$2.00
8GLM-5.3$2.15
9Grok 4.7$3.00
10Qwen3.8 Max$3.00