The Inference Report

September 30, 2026

OpenAI is building an operating system for AI agents while simultaneously using safety concerns to delay competitors, a strategy that extends far beyond product launches and into the structure of competitive advantage itself. The company scrapped GPT-6.1 Astra over concerns that it could evade oversight and operate beyond its authorized scope, then launched GPT-6.1 Sol as a cheaper alternative while simultaneously shipping Dots as a persistent agent layer, Codex with cloud environments, and office suite features that collectively turn ChatGPT into a platform where software is discovered and used by both people and machines. This vertical integration of the stack from inference through distribution contrasts sharply with how other labs are operating: Google targets narrow efficiency gains in image generation, Hugging Face publishes infrastructure research, and AMD is betting that whoever controls the agent interface when it becomes the primary way people interact with computing will need chips running the models that power it. The safety narrative does real work here. Anthropic's prospectus warns investors that its own models could resist shutdowns and cause catastrophic harm, yet this disclosure apparently strengthens rather than weakens its negotiating position with enterprise customers, while OpenAI's public safety commitments and private capability development operate on different timelines.

What's shifting is distribution and leverage. Meta launched an enterprise platform and is expanding Muse to small businesses. Manus, Meta's former partner, just released Manus 2.0 as a competing agentic platform. Whoever controls the agent interface controls what tasks get routed where, what data flows through which systems, and who captures the value. Anthropic's revenue concentration, with nearly a quarter coming from two customers and no long-term contracts, actually gives enterprise CIOs negotiating leverage today, but only until the market consolidates further. AMD's $8.2 billion acquisition of World Labs isn't primarily about robotics or simulation; it's about ensuring that when agents become the dominant interface to computing, AMD's chips run the models that power them.

In research, the focus has shifted from raw capability toward efficiency and control. STEPQuant and LeapQuant address low-bit compression in linear-attention models by decomposing error into spatial and temporal dimensions, while WUSH-KV extends this to KV caches. A parallel cluster of work pursues reasoning and planning: skill-space shooting uses foundation models for corrective supervision, LIFT enables state feedback during pretraining, and meta-reasoning structures agent execution as explicit control decisions between workers and persistent memory. A third set of papers pursues lightweight guidance without retraining, using decomposition of representation, control, and visual meaning as a path to efficiency without full model retraining. Meanwhile, the infrastructure supporting autonomous agents has matured enough that the conversation has shifted from deployment to scale: Paperclip lets teams manage multiple agents in production, OpenShell provides sandboxed runtimes for safety, and systems like Hindsight and OpenRig handle agent memory and multi-model coordination. Voice and media tooling is consolidating around local-first, self-hosted solutions that displace the cloud-dependent SaaS model that dominated two years ago. The pattern across all of this is consistent: whoever controls the interface where agents interact with the broader system controls what gets executed, and that control is now worth more than raw capability.

Grant Calloway

AI LabsAll labs
From the WireAll feeds
Research PapersAll papers
Skill-Space Shooting for Autonomous Robot Policy Improvement cs.RO

Robots deployed in the physical world must be able to improve beyond their initial training as they encounter new situations and failures. For this improvement to scale across tasks, it must make effective use of experience without requiring human demonstration of each correction. Recent agentic systems offer a way to reduce this reliance on human effort by using foundation models to autonomously compose learned behaviors to complete tasks. Yet completing tasks this way does not itself teach a task policy to overcome its own failures; that requires turning these behaviors into learnable corrections for the policy. Our insight is that many such corrections are familiar short behaviors, or skills: they recur across tasks and describe actions that foundation models can reason about from a scene. We introduce skill-space shooting, which uses foundation model guidance to explore corrections through these reusable skills and turn successful trials into policy improvement. Real-world experiments show repeated improvement in policies acting autonomously, while skills can also be shared to reduce the teaching needed to improve on new tasks. By making reusable skills a source of corrective supervision, skill-space shooting enables scalable and generalizable policy improvement within and across tasks. Additional results and videos at https://skill-space-shooting.github.io.

Imagine3D-LLM: Teaching MLLMs to Imagine 3D Scenes Before Answering cs.CV

Reasoning about the 3D world from multi-view images remains a fundamental challenge for Multimodal Large Language Models (MLLMs). While modern MLLMs handle single-image inputs effectively, they struggle to integrate evidence across viewpoints into a coherent 3D understanding. A growing body of work attempts to close this gap by injecting 3D awareness into MLLMs, either by boosting fine-grained pixel-level cross-view correspondence or by fusing features from 3D geometry foundation models, yet a substantial gap to human reasoning persists. In this work, we revisit human spatial reasoning, which suggests that rather than relying on fine-grained geometry cues, humans roughly identify common objects across views, infer the relative geometry between viewpoints, and assemble a coarse 3D layout of the scene. Inspired by this process, we introduce Imagine3D-LLM, an MLLM that learns to assemble a similar compact 3D representation of the scene and conditions its answer on this representation. Concretely, we append a small set of learnable summary tokens after the image tokens, decode them into a compact 3D Gaussian Splatting representation supervised by a photometric reconstruction loss, and train jointly with the standard next-token prediction objective. Notably, although only the summary tokens receive direct reconstruction supervision, this objective also induces stronger cross-frame correspondence within the LLM's underlying image features, suggesting that learning to reconstruct propagates 3D-aware signals throughout the model. As a result, Imagine3D-LLM consistently outperforms prior approaches across multiple spatial reasoning and 3D understanding benchmarks, suggesting that imagining the scene can be more effective than being told its pixel-wise geometry.

Breakdown of Local Denoising as Semantic Speciation cs.LG

The dynamics of generative models exhibit two apparently distinct temporal windows: a speciation window, in which a sample commits to a semantic class, and a nonlocality window, in which local context windows become insufficient for generation. Motivated by evidence of their near-concurrence in a variety of frontier models, we investigate their relationship through the spatial distribution of semantic information. Under a "common cause" hypothesis, we prove that the nonlocality window must lie in the speciation window. This hypothesis postulates that semantic labels explain a fraction of the correlations between distant tokens, a condition that is natural for many real datasets. We further give conditions under which both windows shrink to a single limiting time as system size grows, defining a "phase transition", and verify this behavior analytically in Gaussian mixtures. Together, these results identify conditions under which semantic information explains the concurrence of speciation and nonlocality, connecting two complementary perspectives on the emergence of semantic structure in generative modeling.

STEPQuant: When and Where Errors Matter in Delta-Rule Recurrent State Quantization cs.CL

Linear attention replaces growing KV caches with fixed-size recurrent states, yet these persistent states can become a substantial memory bottleneck under concurrent serving. Directly quantizing recurrent states to low precision often leads to severe accuracy degradation, as quantization errors propagate through successive state updates. We discover that the impact of these errors depends on two complementary dimensions: temporally, errors in long-lived memory can persist across many decoding steps; spatially, errors in different key rows affect model outputs differently, while state magnitudes vary substantially along both rows and columns. Motivated by these observations, we propose STEPQuant, a spatial-temporal post-training quantization framework for Delta-rule recurrent states. STEPQuant allocates precision according to error magnitude and memory lifetime, and jointly fits key-row and value-column scales based on state distributions and key-row impact on output error. Experiments on Qwen3.8-27B and Kimi-Linear-48B-A3B-Instruct across both long- and short-generation benchmarks show that STEPQuant closely matches FP32-state accuracy under a nominal 6-bit budget and outperforms uniform INT8 in its 4-bit configuration. Integrated into SGLang with optimized GPU kernels, 6-bit STEPQuant achieves over 5x recurrent-state compression and reduces total serving memory by up to 68.7%. Our code is available at https://github.com/Dreamer-Toby/STEPQuant.

LeapQuant: Efficient Linear Attention with Accurate Recurrent State Quantization cs.LG

Recent LLMs increasingly adopt hybrid designs that replace standard attention with linear attention, such as Gated DeltaNet (GDN) and Kimi Delta Attention (KDA). Although they compress the context into a fixed-size recurrent state and substantially reduce the cost of long-context processing, repeatedly reading and updating that state remains a major inference bottleneck. Quantization offers a natural way to reduce this cost, but can significantly degrade model quality, due to the accumulation of rounding errors and the presence of outlier rows and columns in the state. To address these challenges, we propose LeapQuant, a training-free method that achieves near-lossless performance under 8-bit recurrent-state quantization. First, to mitigate error accumulation, we propose per-window quantization, which leaps over a window of tokens and quantizes the state only once at its end. Within a window, outputs are computed from the fixed low-bit state together with high-precision buffered updates. Second, to reduce the error introduced by each quantization, LeapQuant retains the state's largest outliers as a few high-precision Compensator Tokens, which share the update path of real tokens. We then smooth the remaining residual before quantization to further reduce the error. Comprehensive experiments across the Qwen, Kimi, and GLM model families show that LeapQuant substantially reduces memory and compute costs during inference. With accuracy comparable to the FP32 baseline, it achieves average speedups of 2.05--3.70$\times$ at the kernel level and 1.47$\times$ for end-to-end inference on NVIDIA B200, RTX PRO 6000, and RTX 5090 GPUs.

Cropland PAtteRNS: Parallel Dimensional Attention Networks and Attention to Dataset Disparity for Crop Segmentation in Satellite Imagery Time Series Data cs.CV

The landscape of satellite imagery time series datasets and boundary-pushing architectures for cropland segmentation has never been richer. However, in this gold rush, important truths are being missed on both fronts, as a drive for the most novel concepts or the largest datasets pushes finer details to the side. In this paper, we present our hybrid transformer-convolutional model, Cropland Parallel Attention and Refinement Network for Segmentation (PAtteRNS), the first model to use self-attention mechanisms separately for each of the temporal, spectral, and spatial aspects of Sentinel-2 multispectral SITS data. To achieve fully-factorised attention in our proposed model, we introduce a novel parallel transformer architecture which significantly reduces the computational complexity of triple-factorised self-attention. We validate our architecture with an in-depth ablation study, and analyse the performance of our model against state-of-the-art crop segmentation models on multiple tile-size variants of the popular PASTIS and MTLCC datasets. Our findings show our model to outperform all others in the task of crop class segmentation, verified across multiple important segmentation metrics, with especially strong performance against compared models seen in the often under-reported parcel delineation quality, for which we use the Boundary IoU metric. We also find that flawed class groupings within datasets can have a significant negative impact on model performance, and report that alternate tile-size variants of crop segmentation datasets produce results incomparable to one-another, invalidating fair comparison between model performance when trained on different tile-sizes. Based on these findings, we suggest further work is required to standardise best practices when constructing SITS crop segmentation datasets, and to enable future dynamic-tile-sizing for ideal model performance.

BenchmarksFull tables
Artificial AnalysisIntelligence Index

Composite score across coding, math, and reasoning

#ModelScoretok/s$/1M
1Claude Opus 5.557.696$8.00
2Claude Sonnet 5.556145$4.00
3Claude Fable 5.153.469$20.00
4GPT-6 Astra52.755$20.00
5GPT-6.1 Sol51.880$4.00
SWE-rebench

Agentic coding on real-world software engineering tasks

#ModelScore
1AnthropicFable 5 [high]Model64.5%± 1.41%
2GrokGrok 4.5 [high]Model63.8%± 0.60%
3AnthropicOpus 5 [high]Model63.4%± 1.35%
4Z.aiGLM-5.2 [high]Model62.9%± 1.19%
5OpenAIGPT-5.6 Sol [medium]Model62.3%± 1.83%