The Inference Report

August 23, 2026

The AI industry is splitting into two operating modes: one focused on appearing responsible at scale, the other focused on shipping products that users actually pay for. OpenAI lobbies California for stronger safety rules it once opposed, yet no leading lab has publicly documented containment plans for rogue models. Meanwhile, Anthropic's Claude 3.5 Sonnet struggles to convert corporate users despite technical sophistication, while narrower tools like Legora in legal tech and Inherent's Faraday in scientific replication capture real demand. When Faraday outperforms Claude and GPT-4 at replicating research, the conversation about foundation model superiority becomes secondary to the conversation about who ships products users will pay for.

The market's preference for task-specific deployment over general capability extends across every layer of the stack. In signal processing, the field has shifted decisively toward interpretability and physical grounding: Polar MKANs enforce disentanglement in RF fingerprinting, LuminaECG embeds cardiologist-grounded primitives into vision-language reasoning, and self-supervised pretraining strategies now dominate data-scarce regimes. On GitHub, agentic coding tools are consolidating around skill harnesses and memory layers rather than competing on agent architecture itself, while infrastructure quietly consolidates around workflow automation, observability, and open-source alternatives to rented tooling. The SWE-rebench benchmark shows saturation at the top tier, with the gap between first and seventeenth place spanning 47.4 percentage points, suggesting frontier models may be converging on the benchmark's difficulty floor.

The pattern repeats in infrastructure: automation in air traffic control gets faster and more complex while regulatory and operational frameworks lag, creating new failure modes no one planned for. The labs are racing to appear responsible while building containment theater. The market is racing past them toward applications that work, even if they're narrower. The real economic value isn't in foundation models themselves but in task-specific deployment, human-AI collaboration at scale, and the infrastructure that makes composition reliable enough to ship to production.

Grant Calloway

AI LabsAll labs

No lab headlines.

From the WireAll feeds
Research Papers — FocusedAll papers
Combining Improvements in Uplink AI-RAN eess.SP

One of the major transformative factors in 6G will be the integration of Artificial Intelligence (AI) to become a native part of Radio Access Network (RAN). While most physical-layer AI features have so far been evaluated in isolation using link-level simulations, their combined behavior in a realistic multi-cell, multi-UE deployment has remained largely unexplored. In this paper, we present system-level performance results when multiple uplink AI features are enabled together, achieved by integrating accurate link-level and system-level simulators. To infer state-of-the-art deep-learning-aided Multiple Input Multiple Output (MIMO) receivers under the dynamic allocations produced by a realistic uplink scheduler, we propose a mirrored data augmentation method that decouples receiver performance from scheduled allocation size. In addition to these Physical Layer (PHY) receiver features, we combine several recent advances in deep reinforcement learning to train uplink power control and link adaptation that outperform a heuristic baseline and further boost the gains obtainable from the AI receiver alone. The system-level results show that the combined AI features improve the mean uplink user throughput by roughly 27% compared to a non-AI baseline, confirming that the individual PHY and Medium Access Control (MAC) AI features provide complementary gains when deployed jointly.

PI-AMFM: Permutation-Invariant Learning for Variable-Cardinality AM-FM Mode Decomposition in Biomedical Signal Analysis eess.SP

Physiological recordings often contain nonstationary oscillatory components whose number and dynamics vary across signals. Amplitude- and frequency-modulated (AM-FM) representations are well suited to characterizing such dynamics and have shown broad utility in biomedical signal analysis. Recent approaches have incorporated neural networks to learn mode decomposition patterns from data, but component cardinality is often predefined or determined through separate stopping or selection mechanisms. We propose a permutation-invariant neural framework for variable-cardinality AM-FM mode decomposition (PI-AMFM). PI-AMFM combines a multiscale temporal encoder, Mamba backbone, and component-presence estimation, with permutation-invariant Hungarian matching during training. On synthetic AM-FM signals, PI-AMFM achieved lower decomposition, instantaneous-frequency, reconstruction, and mode-count errors than the compared methods while preserving the overall trajectory pattern in a crossing-chirp example. On photoplethysmographic recordings, recovered modes captured cardiac and respiratory dynamics despite training only on synthetic signals. These results support the feasibility of PI-AMFM for variable-cardinality decomposition of nonstationary biomedical signals.

Token Communication-Assisted Collaborative Embodied Artificial Intelligence: Concepts, Framework, and Opportunities eess.SP

Collaborative embodied artificial intelligence (CEAI) enables multiple physical agents to perceive, reason, and act cooperatively in dynamic environments. Effective communication is essential for CEAI, yet CEAI agents must exchange not only large multimodal observations but also task-relevant insights, intents, and interactive information over long horizons. This article investigates token communication (TokCom) as a native intelligence interface for CEAI, in which tokens serve jointly as compact semantic carriers for communication and fundamental inference units for generative foundation models (GFMs). We first discuss how TokCom supports insight sharing, intent alignment, and interactive control among embodied agents. We then propose a TokCom-assisted CEAI framework driven by a task-adaptive communication protocol. Comprising a compact codebook, syntax rules, and contextual examples, this protocol guides GFM-based transceivers to distill messages into compact tokens and reconstruct them after wireless transmission. A case study on collaborative object transport demonstrates that the proposed TokCom framework substantially reduces the source payload bit consumption while preserving task efficiency and showing robustness under noisy channels. Finally, we outline future research directions.

Grounding Time-Series Foundation Models in Digital Twin Topology for Predictive Maintenance eess.SP

Digital twins increasingly support downstream analytical tasks that depend on time-series data, motivating interest in time-series foundation models (TSFMs) as scalable backbones. However, TSFMs are primarily pretrained for temporal continuation and often underperform on unseen tasks such as regression, and systematic empirical comparisons against state-of-the-art dedicated models in digital twin contexts remain limited. This paper makes three contributions. First, we benchmark five well-known TSFMs with frozen backbones on remaining useful life (RUL) prediction using the C-MAPSS dataset, finding that multivariate architectures substantially outperform univariate ones, particularly under varying operating conditions. This raises a deeper question: when cross-channel dependencies can be modeled through pretrained weights, target-task adaptation, and digital twin-derived representations, how much does each contribute, and are they complementary? Second, we propose a topology-informed fusion approach in which topological constraints, derived from the asset structure the digital twin stores among its information models, explicitly shape cross-attention, so that fused representations respect the physical system's local connectivity rather than relying on unconstrained all-to-all interactions. Third, we conduct an ablation study across C-MAPSS subsets of varying operational complexity that isolates the three sources and their interactions. The sources prove complementary rather than redundant, and topology-constrained attention outperforms unconstrained fusion, though by a small margin, enabling a frozen TSFM informed by digital twin representations to remain competitive or in some cases exceed state-of-the-art performance on this regression task.

Geometry-Aided Channel Deduction with Partial Channel Estimates and Uncalibrated Digital Twin eess.SP

The acquisition of high-dimensional channel state information (CSI) in wireless MIMO-OFDM communications usually requires high pilot overhead, or relies on accurate and complete positional or environmental information. In this paper, we propose a geometry-aided channel deduction (GCD) approach, which utilizes an uncalibrated digital twin (DT) with only approximate environmental geometry and positions to assist the channel acquisition. The key rationale behind is that, even imprecise geometric information, which can be easily obtained in advance through radio sensing technologies or existing geographic databases, provides certain structural features about the current channel; meanwhile, the coarse instantaneous channel estimates using only a small amount of pilots provide dedicated information that aligns with the channel structure and further compensates for the geometry inaccuracy and other channel unknowns. To this end, we first extract geometric features from the DT, which contain only simple structural information of the channel. Then we propose random prompt augmentation, a novel method to generate an appropriate prompt that converts geometric multi-path structure into a CSI-like representation while suppressing the disturbance of other unknown channel parameters. The prompt is then fused with the pilot-based instantaneous channel estimate via a channel deduction network. To further enhance the network's versatility, we incorporate pilot configurations into the existing learning architecture to support variable pilot patterns. Comprehensive experiments validate the superiority of the proposed method, which demonstrates high channel acquisition quality, low pilot overhead, and strong robustness. Furthermore, the structural prompt also serves as scenario-related context, enabling our approach to generalize well in new scenarios.

End-to-End Optical Semantic Communication over a Nonlinear WDM Fiber Link eess.SP

Emerging optical-network applications increasingly use received data for inference and control rather than exact source reproduction, creating an opportunity to trade bit-level fidelity for greater transmission reach and efficiency. We propose an end-to-end optical semantic communication system for joint image classification and reconstruction over a nonlinear wavelength-division multiplexed (WDM) fiber channel. The system maps each image directly into a fixed-length sequence of channel symbols that preserves task-relevant information, without explicit source compression or channel coding. Experiments on the MNIST dataset cover launch powers from -9 to +3 dBm, fiber lengths up to 800 km, and 16-, 64-, and 256-Quadrature Amplitude Modulation (QAM) formats. At 0 dBm, classification accuracy remains between 98.92% and 99.31% across all tested link lengths and modulation orders, while requiring fewer transmitted symbols than a Low-Density Parity-Check (LDPC)-coded JPEG baseline at every tested modulation order. These results show that semantic communication can simultaneously extend optical reach and reduce transmission resources by conveying only task-relevant information.

BenchmarksFull tables
Artificial AnalysisIntelligence Index

Composite score across coding, math, and reasoning

#ModelScoretok/s$/1M
1Claude Opus 563.153$10.00
2Claude Fable 562.167$20.00
3GPT-5.6 Sol60.978$11.25
4Grok 4.660.958$3.00
5Kimi K359.735$6.00
SWE-rebench

Agentic coding on real-world software engineering tasks

#ModelScore
1AnthropicFable 5 [high]Model64.5%± 1.41%
2GrokGrok 4.5 [high]Model63.8%± 0.60%
3AnthropicOpus 5 [high]Model63.4%± 1.35%
4Z.aiGLM-5.2 [high]Model62.9%± 1.19%
5OpenAIGPT-5.6 Sol [medium]Model62.3%± 1.83%