The Inference Report

September 1, 2026

Consolidation is reshaping AI's competitive terrain, and the winners won't be whoever builds the best models. OpenAI is locking down government partnerships and consumer monetization while controlling regulatory narratives. Anthropic, Google, and Nvidia are pursuing distinct infrastructure angles: security as competitive advantage, specialized foundation models for enterprise forecasting, and hardware control through the MediaTek partnership. Meanwhile, legal and contractual leverage is becoming weaponized. Anthropic faces a Sony lawsuit over staff chat messages praising piracy. OpenAI cut off Cursor's model access after SpaceX acquired its parent company. Apple presented evidence of alleged data theft by a former employee now at OpenAI. These aren't isolated disputes. They're markers of incumbents using legal pressure to constrain competitors and control training data, talent, and model access. The EU's Digital Services Act now applies to ChatGPT and Reddit, imposing compliance costs that scale with user growth and landing heaviest on companies that have already won distribution.

The market is fracturing into defensible layers. Wall Street banks are cutting law firm fees because AI made routine work faster. Insurance claims adjusters report 98 percent negative Glassdoor reviews mentioning AI. Blue Voice raised $6 million to build domain-specific models for police departments because general-purpose models can't compete on local knowledge. Clipto hit profitability at $15 million ARR by solving a specific problem at scale. The Pentagon now runs ChatGPT, Grok, and Gemini on a central portal. Nvidia's $3.5 billion MediaTek investment reveals the real anxiety: Big Tech is building its own AI chips because relying on Nvidia creates a chokepoint they can't tolerate. Neoclouds are capturing $25 billion in revenue, and Gartner expects them to claim 20 percent of the $267 billion AI cloud market. Hyperscalers are losing pricing power.

Research and infrastructure development reflect the same logic. Papers cluster around efficient inference through structural redesign, principled evaluation under resource constraints, and recovery of latent reasoning from model internals. Progress comes from architectural clarity, not parameter scaling alone. On GitHub, the shift is decisive: agents are becoming operational infrastructure, not novelty. What's trending is tooling that routes work between models based on cost and capability, skill libraries that standardize what agents can do, and visualization layers that make behavior auditable. Data handling for training and inference is concentrated. Developers are extending what exists rather than demanding migration. ParadeDB extends Postgres instead of replacing it. The operative assumption across all layers is identical: the winner controls the input, the output, or the platform the output runs on.

Grant Calloway

AI LabsAll labs
From the WireAll feeds
Research PapersAll papers
Context-Aware Interleaved Batching for WhisperX cs.CL

While WhisperX accelerates speech transcription via intra-audio batching, it isolates audio segments, losing the historical context needed for coherent punctuation and terminology transcription. Conversely, standard Whisper retains context sequentially but suffers from slow inference and hallucination loops. To achieve the best of both worlds, we propose Context-Aware Interleaved Batching. By using VAD-derived segment boundaries, our algorithm stabilizes Whisper's text conditioning, allowing us to safely maintain continuous historical context across batched audio segments. As demonstrated on long-form audio benchmarks, this approach reduces Word Error Rate (WER) and improves proper noun transcription, all while maintaining high-throughput inference speeds.

Constant Individual Regret in General Games cs.LG

Uncoupled no-regret dynamics provide a decentralized route to equilibrium, but prior guarantees for individual regret retain a polylogarithmic dependence on the horizon. We remove this dependence for every finite $N$-player normal-form game under full-information feedback. We introduce \emph{ECHO-OFTRL}: optimistic follow-the-regularized-leader (OFTRL) equipped with an EMA cascade for high-order optimism (ECHO), where EMA denotes exponential moving average. The algorithm is deterministic and fully uncoupled. If $m_{\max}$ denotes the largest action-set size, then, simultaneously for every horizon $T\geq1$, it guarantees that each of the $N$ players in the game incurs regret upper bounded by $O(\textrm{poly}(N, \log m_{\max}))$. Our algorithm leverages a new form of optimism inspired by modern filter design.

SUN: Persistent Programs For Language-Grounded Control-to-Learning-to-Real Policies cs.RO

Bridging model-based control and learned policies in long-horizon manipulation has harbored a silent disagreement: control executes specified objectives, learning amortizes that behavior into a reactive policy, yet existing protocols discard task semantics, leaving rewards hand-crafted and behavior drifting from what control verified.We introduce Semantically UNified (SUN) Programs, typed executables where geometric and contact relations are defined once and compiled into aligned Model Predictive Control (MPC) costs, satisfaction predicates, RL rewards, transition guards, and diagnostics. Our system, Kuafu, driven by large vision language systems, automatically synthesizes SUN Programs from language and scene semantics, screens feasibility via MPC, and retains semantics while training stage-conditioned policies. Across nine tasks, Kuafu achieves 82.03% macro-success, outperforming sparse-reward (35.67%) and Stage-BC (24.75%) baselines. At 8192-way scale, it generates 10.57x the successful trajectory time per hour of human teleoperation. With 500 trajectories per task, Kuafu data trains DP3 policies to 46.0% simulation success (vs. 22.4% for alternatives) and 34.7% on physical Franka and Kinova robots. These results establish that simulation-screened task semantics can effectively amortize control into robust policies, without demonstrations or manual dense rewards, unifying symbolic planning and data-driven execution.

Sharp Approximation Rates for Neural Networks with Affine Latent Parameterizations cs.LG

Many parameter-efficient methods generate the parameters of a large neural network from a low-dimensional latent representation. Given an architecture $Φ$ with $P_Φ$ parameter slots, we write $\boldsymbolθ_f=\mathcal{G}(\boldsymbolξ_f)$, where $\mathcal{G}\colon\mathbb{R}^M\to\mathbb{R}^{P_Φ}$ is a parameter generator and $\boldsymbolξ_f\in\mathbb{R}^M$ is a latent representation of the target function $f$. The architecture $Φ$ and the generator $\mathcal{G}$ are shared across the entire target class, while each target $f$ is represented by its own latent vector $\boldsymbolξ_f$, with $Φ_{\mathcal{G}(\boldsymbolξ_f)}$ approximating $f$. This framework encompasses hypernetworks, low-dimensional parameterizations, parameter-efficient adaptation, and model compression. Understanding the tradeoff between the latent dimension $M$ and the network budget $P$ is therefore fundamental to characterizing the expressive efficiency of these methods. We study this tradeoff for affine generators and fully connected ReLU architectures. More precisely, optimizing jointly over architectures $Φ$ satisfying $P_Φ\leq P$ and affine generators $\mathcal{G}:\mathbb{R}^M\to \mathbb{R}^{P_Φ}$, we prove that the optimal worst-case uniform approximation error over the unit ball of $α$-Hölder functions on $[0,1]^d$, where $0<α\leq1$, has the sharp order $ \bigl(P\min\{M,P\}\bigr)^{-α/d}. $ In particular, our result shows that even a fixed-dimensional latent space suffices to achieve vanishing approximation error as the network budget increases.

Auditing Anonymous AI Models: A Four-Stage Protocol for Black-Box Identity Verification cs.SE

The 2025--2026 AI market has seen a wave of stealth releases: frontier models launched anonymously on developer platforms under codenames. For their users, identity determines data-handling terms, supply-chain risk, and capability expectations. No validated methodology exists for black-box identity verification of anonymous models: practitioner checklists lack accuracy evidence, and self-identification is untrustworthy by design. We propose a four-stage forensic audit protocol for API-served models. Stage 0 reconstructs launch-time configuration from archived platform snapshots (Internet Archive), exposing preview--production drift. Stage 1 fingerprints configuration (context, output ceiling, reasoning, modality) against the platform catalog. Stage 2 tests tokenizer identity with a cross-length differential that rejects short-prompt collisions. Stage 3 corroborates with behavioral probes. We test declaration consistency on 10 known-identity releases (7 exact, 2 precision-differences, 1 partial, 0 counter-directional), not end-to-end identification under anonymity. Identification is validated prospectively on a flagship case whose 2026-08-23 analysis pointed to the GLM-5.3 version line and whose official reveal confirmed those family and version-line inferences (deployment variant was not pre-asserted; Flash was consistent post-reveal), and on three Stage-0-only cases where the protocol produced a graded hypothesis or declined rather than guessed. A standard-library-only implementation is provided as supplementary material.

Configurable Semantic Chunking for Biomedical Information Extraction in Retrieval-Augmented Generation cs.CL

BioMedRAG introduced retrieval-augmented generation with a learned chunk scorer for biomedical information extraction. However, it relies on fixed-size chunking which can fragment semantic evidence. We propose a configurable semantic chunking framework that addresses this limitation by combining entity-preserving windows, trigger-centered chunking, proposition-first extraction, tiered trigger prioritization, and hierarchical relation resolution. The framework integrates with BioMedRAG by replacing only the chunk construction stage while preserving the embedding model, learned chunk scorer, generator, and evaluation protocol. We evaluate the framework on biomedical relation extraction benchmarks (GM-CIHT, DDI, ChemProt) and adverse event classification (ADE). On GM-CIHT, the full hybrid configuration achieves 82.6% F1, improving over the fixed-size baseline (74.2% F1) by 8.4 points under our experimental setup. Cross-dataset analysis shows that semantic chunking improves extraction datasets with explicit relation cues, such as GM-CIHT and DDI, while fixed chunking remains competitive or stronger for dense biochemical extraction and binary classification settings such as ChemProt and ADE. By externalizing chunking logic into configuration files, the framework provides an interpretable and adaptable alternative to rigid fixed-size chunking for biomedical RAG pipelines.

BenchmarksFull tables
Artificial AnalysisIntelligence Index

Composite score across coding, math, and reasoning

#ModelScoretok/s$/1M
1Claude Opus 563.154$10.00
2Claude Fable 562.162$20.00
3GPT-5.6 Sol60.984$8.00
4Grok 4.660.958$3.00
5Kimi K359.736$6.00
SWE-rebench

Agentic coding on real-world software engineering tasks

#ModelScore
1AnthropicFable 5 [high]Model64.5%± 1.41%
2GrokGrok 4.5 [high]Model63.8%± 0.60%
3AnthropicOpus 5 [high]Model63.4%± 1.35%
4Z.aiGLM-5.2 [high]Model62.9%± 1.19%
5OpenAIGPT-5.6 Sol [medium]Model62.3%± 1.83%