The Inference Report

May 15, 2026

"The constraint is no longer capability. It's capital, energy, and talent retention."

The infrastructure of AI is revealing itself as fundamentally material rather than neutral. NV Energy's decision to cut water to 49,000 Lake Tahoe residents in favor of Nevada data centers represents the most literal expression of this reality: when energy is scarce, it flows to the highest-value customer, not the most critical one. This same logic is reshaping compensation across the sector. Anthropic is moving Claude from unlimited subscriptions to per-call billing starting June 15, a metering mechanism that transforms how agent work gets valued. McKinsey is shifting partner compensation toward equity. Cisco cut nearly 4,000 jobs while posting record revenue. These aren't isolated business moves. They're synchronized signals of capital and labor being redistributed as companies move from research to production at scale. When partnerships fail, OpenAI exploring legal action against Apple over a ChatGPT integration that failed to deliver subscribers, SpaceXAI losing 50 employees since its February merger, the Musk v. Altman litigation, the failure is rarely about capability. It's about returns not matching expectations and executives fighting to protect their positions when they don't.

The market is consolidating around operational deployment rather than model advancement. OpenAI is pushing Codex into production workflows and real-time steering at Sea Limited. AWS is moving beyond model selection into prompt optimization tooling that compares outputs across five models simultaneously. Anthropic is signing enterprise deals with PwC and the Gates Foundation. Microsoft and IBM are both positioning themselves as implementation engines for existing organizations rather than research frontiers. The infrastructure layer, AMD optimizing inference, Hugging Face publishing embedding improvements, NVIDIA shipping games on GeForce NOW, is competing for the same outcome: making the layer between model and customer's problem invisible. The labs that win will be the ones that own that layer.

GitHub trending repositories show developers solving concrete friction points: persistent memory and agentic workflows (AgentMemory), computer vision infrastructure (Roboflow's Supervision, NVIDIA's video analytics), and offline inference (Supertone's on-device TTS, LocalAI's hardware-agnostic serving). What's absent is any major push on reasoning or long-horizon planning. The momentum is instead in making existing approaches reliable, composable, and deployable at scale. The gap between established tools and emerging specialized stacks suggests the market for agentic frameworks is fragmenting rather than converging. Across signal processing research, the same pattern holds: self-supervised contrastive methods paired with structured neural architectures are winning over generic approaches because they encode domain-specific priors. Uncertainty quantification is moving from post-hoc calibration to a training objective. Lightweight adaptation techniques are replacing full retraining. The field is optimizing for practical deployment constraints, latency, bandwidth, label scarcity, not theoretical capability.

Grant Calloway

AI LabsAll labs
From the WireAll feeds
Research Papers — FocusedAll papers
Structured Pose-Conditioned Flow Matching for Generative 5G CSI Augmentation eess.SP

With the growing demand for privacy-preserving and occlusion-resilient human pose recognition (HPR), 5G channel state information (CSI) offers a promising contactless sensing modality by integrating communication and sensing capabilities. However, collecting large-scale synchronized CSI-pose pairs remains costly in practical 5G systems. To address this limitation, we propose StructFlow-HPR, a structured pose-conditioned flow matching framework for generative CSI augmentation. StructFlow-HPR learns a continuous latent transport process from Gaussian noise to real CSI representations under pose guidance, while preserving the receiver-frequency topology of CSI through a reconstruction-preserving autoencoder. A pose-conditioned Transformer is further designed to model the latent velocity field and generate pose-aligned CSI samples via ordinary differential equation sampling. Experiments on real-world 5G sensing data show that StructFlow-HPR can produce realistic CSI-pose pairs and improve downstream HPR performance under limited-data conditions.

Untangling the Geometry and Speed for RF Sensing Spectrograms eess.SP

A fundamental challenge in RF sensing is that Doppler signatures observed by a link entangle the target's motion with the sensing geometry, resulting in limited applicability to unconstrained real-world settings. In this paper, we establish a new foundation for physically interpretable RF sensing that disentangles reflector speed from geometry, jointly recovering the speed, geometry factor, relative amplitude, and width of each dominant Doppler ridge. More specifically, we first develop a compact parametric representation of WiFi spectrograms and establish its low-dimensional structure through a systematic computer-vision analysis of a large and diverse human-activity dataset, thereby providing a tractable foundation for learning. Building on this representation, we then design a physics-informed autoencoder whose structured bottleneck and differentiable RF forward model enforce physically meaningful estimates of reflector speed and geometry. We further introduce a synthetic-to-real training framework, eliminating the need for real WiFi training data. We extensively validate the proposed framework under both known and time-varying geometries, using both independently generated synthetic test sets and 31 real WiFi experiments. The results demonstrate the superior performance in speed and geometry extraction, robustly recovering the underlying geometry, speeds, Doppler-ridge amplitudes, and ridge widths across all settings, while substantially outperforming the strongest baselines.

Risk-Aware Online Conformal State Probing eess.SP

AI-based autonomous agents, typically hosted at data centers, must acquire state information from robots or edge devices in order to issue informed control decisions. Managing uncertainty about the state is particularly consequential in safety-critical settings, in which average-case guarantees are insufficient. In this context, we study a sequential decision maker process that jointly decides which actions to take and when to probe given access to an arbitrary state prediction model. We propose online conformal state probing (OCSP), an action and probing policy that certifies worst-case reliability levels without relying on distributional assumptions. OCSP is designed to provably control the missed query error (MQE), i.e., the fraction of instances where probing would have been beneficial, while minimizing the probing rate. OCSP can be applied to existing pre-trained value-based control policies without requiring retraining or fine-tuning. We validate OCSP through numerical simulations to verify theoretical guarantees and to assess performance trade-offs as a function of the calibration of the state predictor.

Unlocking Cross-Scenario Physical Layer Security: A Mixture-of-Experts Framework with Generative Diffusion Models eess.SP

The future 6G networks are expected to incorporate a proliferation of wireless services in diverse environments, which presents a significant challenge for information security. Conventionally optimization always requires recalculation and learning strategy often suffers poor generalization, which are thus incapable for the security provisioning with wide scenario coverage. In this paper, we propose an adaptive and robust learning framework that leverages a mixture-of-experts (MoE) architecture to achieve cross-scenario physical layer security guarantee. Specifically, we first select a few representative scenarios and establish the scenario-specific generative diffusion model (GDM)-based experts for secure transmission beamforming with artificial noise. The diffusion nature of experts learns the overall probability distribution of security strategy solution landscape and the Transformer-based denoising process enhances the ability to generalize across varying network configurations. Then, a lightweight gating network is constructed to identify the scenarios by engineering the channel features and select the most relevant experts. Finally, an attention-based combiner is introduced to synthesize the security proposals from the top-rated experts to produce a high-fidelity security strategy to cover the unseen scenarios. Simulation results demonstrate that the proposed GDM-based MoE framework can accurately recognize the scenarios and properly select the experts, maintaining near-optimal secrecy rates across a continuum of wireless scenarios and outperforming traditional single-model paradigms.

Fast Graph Laplacian Estimation using Effective Resistance eess.SP

Inferring network topology from noisy node observations is a central problem in graph signal processing. In this paper, we consider Laplacian-constrained graph estimation for Gaussian Markov random fields, focusing on the underdetermined regime in which the number of samples is smaller than the number of graph nodes. Existing approaches often formulate the problem as a sparsity-regularized maximum-likelihood estimation problem. While effective, such methods typically require iterative optimization and are often computationally demanding, particularly under Laplacian constraints. Instead, we propose a non-iterative estimator of graph Laplacians that uses effective resistance for regularization, and evaluate the method using a simple sparsification procedure. Experiments show that with some trade-off in edge and weight recovery on the considered dataset, computational cost for moderately sized graphs can be substantially reduced.

Goal-oriented probabilistic forecasting for dynamic PRB allocation in 5G networks eess.SP

Efficient physical resource block (PRB) allocation in 5G networks requires accurate demand forecasting. Conventional methods minimize symmetric error metrics (MAE, RMSE), ignoring the operational cost asymmetry where under-provisioning (service degradation) is far costlier than over-provisioning (wasted capacity). We propose a goal-oriented probabilistic forecasting framework that aligns model training with the operator's decision-making objectives. Specifically, we train DeepAR and Temporal Fusion Transformer (TFT) models using the Pinball Loss function and derive the optimal allocation quantile from the operator's cost matrix. Evaluation on a real beam-level 5G traffic dataset shows that the proposed approach reduces operational cost compared to MSE-trained baselines while maintaining calibrated uncertainty estimates. The framework enables dynamic PRB allocation that explicitly balances service reliability against resource efficiency.

BenchmarksFull tables
Artificial AnalysisIntelligence Index

Composite score across coding, math, and reasoning

#ModelScoretok/s$/1M
1GPT-5.560.266$11.25
2Claude Opus 4.757.362$10.94
3Gemini 3.1 Pro Preview57.2126$4.50
4GPT-5.456.883$5.63
5Kimi K2.653.943$1.71
SWE-rebench

Agentic coding on real-world software engineering tasks

#ModelScore
1Claude Opus 4.665.3%
2gpt-5.2-2025-12-11-medium64.4%
3GLM-562.8%
4Junie62.8%
5gpt-5.4-2026-03-05-medium62.8%