The Inference Report

May 28, 2026

Nvidia's $150 billion commitment to Taiwan represents a decisive realignment of AI infrastructure power away from Washington's policy ambitions and toward the physical realities of chips, power, and proximity to manufacturing. The decision is not a vote of confidence in US incentives but a calculation that Taiwan's existing TSMC ecosystem offers lower friction than any regulatory environment. Snowflake's $6 billion AWS deal for AI CPUs and DigitalBridge's $1 billion acquisition of ArcLight energy underscore the same pattern: capital flows to whoever guarantees reliable compute and power, not to whoever promises the best policy. This shift occurs precisely as the gap between AI's claimed capabilities and its actual performance widens. Google's misspelling of its own name in AI search results and reports that agents ignore evidence and struggle to learn expose fundamental brittleness in the technology, yet companies like Remote report 50% revenue-per-employee gains from AI adoption and Cognition reached $492 million in annualized run rate. The resolution is structural rather than technical: organizations absorb failures through human review, sandboxed environments, and constrained deployment. Productivity gains are real. Safety is purchased through constraint.

The competitive battleground has shifted from foundation model capability to the infrastructure that coordinates multiple agents across fragmented enterprise systems. OpenAI's Codex positioning as agent orchestration connective tissue and Nvidia's framing of AI factories as token factories converting power into intelligence both signal that near-term value capture lies in inference economics, cost per token, and performance per watt rather than raw model capability. Yet Hugging Face's ITBench benchmark revealed frontier models scoring below 50 percent on agentic enterprise IT tasks, exposing the chasm between static benchmark performance and actual agent behavior in production workflows. Anthropic's coding agents for social science research and Hugging Face's local-first robotics deployment show the market already fragmenting by use case and deployment constraint. What remains conspicuously absent is any lab claiming general superiority in agentic reasoning; instead announcements focus on embedding agents into specific workflows where switching costs are real.

Regulation is becoming a competitive advantage for incumbents. Illinois passing America's strongest AI safety bill requiring third-party audits of models from OpenAI, Anthropic, and Google creates compliance costs that lock in dominant players while raising barriers for new entrants. Cognition's $25 billion valuation after eight months, Kirkland & Ellis committing $500 million to build proprietary AI technology, and OpenAI's foundation allocating $250 million to research AI's economic impact all point to the same conclusion: winners are determined by speed and user lock-in before rules harden, not by regulatory favorability. The GitHub trending set confirms this pattern. Infrastructure tooling like Streamlit and NocoBase gains traction because it replaces something developers were already doing. Meanwhile, repos claiming compatibility with twenty platforms and selling taste as a feature show coordinated promotion around specific AI platforms rather than organic adoption. The real signal concentrates in unglamorous spaces where actual constraints live: cost, latency, and the gap between benchmark numbers and production behavior. That is where the work is happening.

Grant Calloway

AI LabsAll labs
From the WireAll feeds
Research Papers — FocusedAll papers
Claude Code Complete User Handbook cs.NI

Claude Code is an agentic work environment: a language model operating in a loop with filesystem access, shell execution, browser control, scheduled and cloud execution, external tool connections through the Model Context Protocol, and multi-agent orchestration. Its capability envelope now exceeds what one practitioner can supervise by attention alone, and its failure modes are systemic rather than local: an unreviewed hook, an over-scoped connector, a stale completion condition, an autonomous routine inheriting every credential on an account. This book is a task-oriented reference for operating that system safely and productively, written for practitioners accountable for the result. It advances four propositions. First, capability without a defined and observable completion condition is not productivity. Second, instruction, permission enforcement, sandboxing and operating-system isolation are four distinct layers of a control stack, only two of which are enforced, and conflating them is the most common cause of loss of control. Third, third-party skills, plugins, marketplaces, channels and MCP servers are software supply-chain dependencies and must be governed as such. Fourth, the correct unit of trust in agentic work is observed evidence, not an agent's closing statement. Thirty-four chapters run from installation to a fully verified capstone, with a governance part on managed policy, data residency and retention, observability and accessibility. Every product claim carries a citation to a primary source; an evidence ledger records where a claim in circulation was found wrong, what a later re-verification changed, and what remains unverified. Controls are mapped to seventeen external frameworks in a crosswalk, and an organisational adoption maturity model is proposed. Claims not confirmable from primary sources are labelled UNVERIFIED rather than softened.

FaulT-Bench: Towards Benchmarking Network Troubleshooting LLM Agents under Unreliable User Tickets cs.NI

LLM-based agents are increasingly proposed for network fault diagnosis, but existing benchmarks evaluate them only on accurate tickets and always assume a fault is present, conditions rarely met in practice. We present FaulT-Bench, a benchmark of 200 troubleshooting scenarios across eight network topologies, five reimplemented from public practitioner labs, spanning genuine faults, false fault reports, incorrect device attribution, and incorrect root-cause claims. To isolate how ticket wording affects diagnosis, we further rewrite 72 false-premise tickets into five reporter personas that vary reporter confidence and verifiable detail one factor at a time, holding the network state fixed. Our automated harness deploys each scenario in Kathará, lets agents interact through the NIKA tool interface, and scores free-text diagnoses with an LLM judge across outcome, fix, and reasoning quality. Evaluating SADE, ReAct, and Claude Code, we find all three are near-saturated on accurate tickets and robust to misdirection, yet degrade sharply when the network is healthy and the ticket is wrong, probing until a benign condition can be promoted to a root cause rather than concluding nothing is wrong. Persona rewrites show that how a ticket is written matters more than what it claims: a confidently wrong report is handled about as well as an accurate one, while a vague, underspecified report degrades performance sharply. The three agents also fail differently, from constant over-diagnosis to unanswered runs, at very different cost. These results position FaulT-Bench as a benchmark for developing agentic systems that can reason reliably over the noisy, unreliable tickets of real-world network troubleshooting.

AERIS: Offline Policy Improvement for Multi-UAV Integrated Sensing and Communication cs.NI

Unmanned aerial vehicle (UAV)-enabled integrated sensing and communication (ISAC) is a promising 6G paradigm, but dynamic multi-UAV ISAC control must jointly balance communication quality, sensing reliability, and flight safety under stochastic mobility. Existing optimization methods often require repeated global non-convex solving, while online reinforcement learning (RL) depends on risky trial-and-error flights that may cause sensing loss or collision-risk events. This paper proposes AERIS, an offline policy improvement framework for multi-UAV ISAC. AERIS learns from fixed flight logs under centralized training and decentralized execution, so each UAV acts from local histories while training uses logged global information to assess team-level effects. We further design STAR-CRDT, an offline multi-agent RL algorithm that performs support-aware local action rectification and distills only trusted improvements into the decentralized actor. We prove an offline-support policy improvement guarantee. Experiments show that STAR-CRDT improves the main ISAC objective return by 29.3% over the strongest baseline. It further improves communication sum rate, sensing pass rate, and sensing margin by 3.4%, 4.8%, and 69.1%, while reducing collision-risk events by 54.2%. On unseen real-road maps built from OpenStreetMap data, STAR-CRDT still obtains the best return.

Place, Slice and Schedule: Hierarchical O-RAN Control of a Tethered mmWave UAV-gNB cs.NI

Unmanned aerial vehicle (UAV)-mounted 5G New Radio base stations (gNBs) can augment terrestrial networks with an on-demand, repositionable Frequency Range 2 (FR2) capacity layer. This flexibility, however, couples the physical network topology with radio-resource management: UAV movement reshapes blockage, channel quality, and the set of effectively served users, while traffic demand, queues, and service requirements evolve at a much faster timescale. Existing Open Radio Access Network (O-RAN)-enabled UAV studies optimize trajectory, deployment, association, or resource allocation, but typically in isolation, without coordinating slow aerial control with fast per-user scheduling. We instead exploit O-RAN disaggregation, Key Performance Indicator (KPI) monitoring, and multi-timescale RAN Intelligent Controller (RIC) control to address this coupling: a Non-Real-Time RIC rApp uses aggregated KPIs and radio-environment context to jointly control tethered UAV placement and the enhanced Mobile Broadband (eMBB)/Ultra-Reliable Low-Latency Communication (URLLC) slice budget, while a Near-Real-Time RIC xApp allocates per-user resources within that budget. We realize this xApp as a permutation-equivariant DeepSets Soft Actor-Critic (D-SAC) scheduler that treats the users as an unordered set, trained in a Sionna RT ray traced channel. The resulting hierarchical controller improves eMBB SLA satisfaction by up to 17% and URLLC on-time delivery by up to 42% over classical and learned schedulers; the learned rApp further raises URLLC on-time delivery by up to 20% over baselines.

Advanced LLM-Enhanced Intent-Based 5G Network Management using Dynamic Semantic Routes cs.NI

As the use of Artificial Intelligence (AI) and Large Language Models (LLMs) is becoming common in everyday applications, their ability to interpret natural language has increased significantly. An emerging application of AI is integration with network management and orchestration practices. An instance of this integration is LLM-enhanced intent-based networking, where network operators will control a network using natural language. This work presents the use of dynamic routes with a semantic router to identify an intent from a network operator's prompt and extract necessary details for intent fulfillment in intent-based 5G+ core networks. Furthermore, the performance of static route selection is assessed by evaluating multiple encoders and dynamic route detail extraction accuracy against a series of realistic operator prompts. The presented results show that static and dynamic routes are successful in detail extraction and schema formatting.

Channel-Token Attention for Reliable Dynamic Spectrum Access under Bursty Primary-User Traffic cs.NI

Dynamic spectrum access must coordinate secondary users under bursty primary-user activity while preserving packet reliability and delay. We present TACAN, a centralized policy that represents each channel as a token containing occupancy history and automatic-modulation-classification entropy; a context token supplies queue class, delay and user identity. A Transformer encoder is warm-started from an occupancy-greedy policy and refined with proximal policy optimization. The frozen policies were trained to maintain a channel assignment in every slot, including when queues were empty. We therefore replay them on held-out trajectories and distinguish standby assignment success from packet-present access and packet delivery. In a 20-channel network with 60 primary devices and 4 secondary users, TACAN achieves 92.53% +/- 0.47 packet-present access success, compared with 89.94% for Greedy and 83.53% for PPO+MLP. Its paired gain over Greedy is 2.59 points (parametric 95% CI 1.89-3.29), with wins in all five seeds; the exact two-sided sign-test value is 0.0625. The gain rises from 0.57 points at normal primary-user load to 7.67 points at extreme load. TACAN also reduces mean delivery delay from 1.208 to 1.123 slots and the conditional user-reliability gap from 9.69 to 3.15 points. Delivered packets per SU-slot remain arrival-limited (30.12% versus 30.11%), so no packet-throughput gain is claimed.

BenchmarksFull tables
Artificial AnalysisIntelligence Index

Composite score across coding, math, and reasoning

#ModelScoretok/s$/1M
1GPT-5.560.281$11.25
2Claude Opus 4.757.355$10.94
3Gemini 3.1 Pro Preview57.2132$4.50
4GPT-5.456.889$5.63
5Qwen3.7 Max56.6206$3.75
SWE-rebench

Agentic coding on real-world software engineering tasks

#ModelScore
1gpt-5.5-2026-04-23-xhigh62.7%
2Codex60.4%
3Claude Code59.6%
4gpt-5.5-2026-04-23-medium58.9%
5gpt-5.4-2026-03-05-medium54.9%