The Inference Report

July 24, 2026

The AI industry has entered a bifurcated mode of operation where survival depends entirely on which side of the capability frontier you occupy. Google's negative cash flow quarter exposes the brutal mathematics of the arms race: the company is burning capital faster than it can monetize it, yet cannot afford to slow down. AMD's Helios system and Etched's $10.3 billion valuation confirm that the genuine bottleneck and profit center sits in the hardware layer, not the models themselves. Policy is moving faster than markets can absorb their own externalities, with a proposed Kill Switch Act handing the Department of Homeland Security authority to order AI shutdowns while guardrails from OpenAI and Anthropic are already blocking cybersecurity researchers from legitimate vulnerability work. The result is a two-tier system: builders with sufficient capital racing toward capability and market share, everyone else managing collateral damage.

The education sector reveals this pattern most clearly. Anthropic and OpenAI are flooding schools with free and discounted tools while universities quietly abandon AI detection systems they recognize as unreliable. The companies with the most compute are colonizing institutions least equipped to negotiate with them. This colonization extends across infrastructure: the actual moat is no longer in the models but in the middleware layer that routes work between them, filters outputs, or constrains behavior. AegisAI's $36 million raise to stop AI-driven phishing, Runway's model router for selecting between competing systems, and OpenAI's Presence service for automating support work all point to the same dynamic. On the hardware side, AMD launched ROCm Infera for distributed inference orchestration across GPU clusters with 2.6x goodput improvement on agentic workloads, while Hyperloom reduces inference optimization from weeks to hours. NVIDIA is cementing relationships with governments and research institutions before competitive intensity peaks, announcing a joint research lab with KAIST in Seoul. Whoever controls the inference orchestration layer for agentic systems controls the margin on the next wave of deployment.

China's emergence of Kimi K3 as a credible alternative to Western models, paired with sovereign wealth fund expectations of strong growth in Chinese AI companies, signals that infrastructure cost advantage is beginning to matter more than marginal capability improvements. Intel's 25 percent revenue jump and Meta's $12 billion data center financing at higher borrowing costs show the capital requirements have moved from abstract to concrete, reshaping corporate balance sheets and investor risk calculations. Research papers cluster around targeted decomposition rather than end-to-end scaling: multimodal integration work like VLM-IE3D and MIRROR, systems-level lifecycle management through OpenForgeRL and Agentic Context Management, and controlled investigation of representational failures. GitHub trending repos show developers prioritizing infrastructure friction and agent composition over flashy new capabilities, with routing abstractions like OmniRoute addressing provider lock-in and orchestration layers for generative media pipelines emerging as sustained traction. The industry has crossed into a phase where the question is no longer whether AI works, but who can afford to keep the lights on and who gets squeezed out when the meter stops running.

Grant Calloway

AI LabsAll labs
From the WireAll feeds
Research PapersAll papers
3D-Aware VLMs with Implicit and Explicit Geometries cs.CV

Despite rapid progress, most existing vision-language models (VLMs) built from 2D visual inputs often struggle when handling various 3D tasks that require fine-grained spatial understanding and reasoning. To bridge this gap, we present VLM-IE3D, a unified framework that enhances the 3D spatial awareness of VLMs by equipping them with both implicit and explicit 3D geometries learned from RGB videos. Our VLM-IE3D introduces Implicit Geometry Tokens (IGTs) that capture high-level geometric priors from input videos, as well as complementary Explicit Geometry Tokens (EGTs) that encode detailed geometric structures from reconstructed 3D attributes. On top of that, VLM-IE3D comes with a 3D-aware adapter that effectively fuses the two types of geometric representations with 2D visual cues. This RGB-only design injects strong 3D inductive biases for fine-grained spatial understanding and reasoning without requiring any additional 3D inputs. Extensive experiments show that VLM-IE3D achieves superior performance consistently across various 3D tasks including 3D video detection, 3D visual grounding, 3D dense captioning, and spatial reasoning. Code and models are available at https://github.com/Vegetebird/VLM-IE3D.

Expanding Flow Maps cs.LG

Flow-based generative models have enabled remarkable progress in fast and controllable generation across continuous and discrete state spaces, yet existing parameterizations are constrained to fixed dimensions or fixed sequence lengths. Here, we introduce Expanding Generative Flows (EFlows), which define flows between distributions of increasing dimensionality along an expanding interpolant that grows the state by augmenting it with conditional noise. Building on this construction, we propose Expanding Flow Maps (EFMs), a new class of flow maps that distill the expanding interpolant into efficient few-step generative models. Each EFM factors the map between any two timesteps into two learnable operations: an expand operator, which augments the state space with new coordinates or tokens conditioned on the current state, and a transport map, which pushes the expanded state forward along the interpolant. Composing these operators yields a single map that jointly expands and denoises the state, recovering existing fixed-canvas flows and flow maps as the special case in which the expand operator is the identity. We further extend the framework to the discrete simplex, enabling variable-size graph generation and variable-length sequence generation. Across both continuous and discrete modalities, we establish EFlows and EFMs as a principled framework for settings in which output size is itself a learned, controllable degree of freedom.

GraphVid: Interactive Graph-Controllable Video Generation cs.CV

Controllable video generation remains challenging due to the difficulty of specifying precise multi-object interactions using text prompts or motion-control inputs that primarily constrain pixel movement. In practice, trajectory-based control often requires users to draw accurate tracks for multiple objects, which scales poorly with scene complexity and becomes ambiguous under occlusion or overlap. To enable flexible yet precise multi-subject control, we introduce $\textbf{GraphVid}$, a graph-conditioned image-to-video generation model that enables interactive control through structured interaction graphs. We further curate $\textbf{GraphVid-Bench}$, a large-scale interaction-centric video dataset with structured relational annotations to enable training of interaction-aware video generation models. Despite using substantially less training data and fewer trainable parameters than prior motion-control methods, GraphVid delivers strong controllability and video quality. Compared with Motion-I2V, GraphVid reduces FID by up to 39.9% and FVD by 37.6%, while improving PSNR (9.87=>15.98) and SSIM (0.38=>0.61). Our results highlight the potential of structured semantic interfaces as a powerful paradigm for controllable video generation.

Barzilai-Borwein Fails Superlinear Convergence on an Open Set of Quadratics for Every Dimension $n\geq 4$ math.OC

Barzilai--Borwein (BB) method has shown strong practical performance in continuous optimization, yet its convergence dynamics remains poorly understood. In particular, a central unresolved question is whether BB converges superlinearly for almost every strictly convex quadratic problem and initialization. We provide a negative answer to this question. Specifically, for every finite dimension $n\geq4$, we construct a nonempty open, hence positive-Lebesgue-measure, family of strictly convex quadratic problems and initial points for which the long Barzilai--Borwein method (BB1) converges but cannot converge root-superlinearly. More precisely, with the explicit constants $ρ_{\min}=10^{-6},ρ_{\max}=0.61$, every spectral component of the gradient is bounded above and below by the corresponding geometric sequence. Consequently, the gradient norm and the energy norm of the error satisfy two-sided geometric estimates with the same rates, while the objective gap satisfies the corresponding estimates with squared rates. In particular, all three quantities are bounded below by geometric sequences, ruling out superlinear convergence. The construction is highly nontrivial, based on a computer-assisted proof of a nonresonant, attracting seven-cycle of the projectivized BB dynamics in dimension four.

Synthetic data generation framework for quality control automation in gravure printing cs.CV

Quality control in printing, particularly in rotogravure printing, still depends on slow, costly, and subjective manual inspection. Automated surface defect detection is critical for maintaining high-quality standards in rotogravure printing. Deep learning models give prospects for automation. However, training robust deep learning models, such as YOLO or Vision Transformers, is heavily hindered by the extreme scarcity of real-world industrial defects images. To overcome this limitation, this paper introduces a novel synthetic data generation framework tailored for rotogravure printing quality control. The proposed pipeline automatically generates high-fidelity images of specific printing defects (creases, streaks, misregistration, etc.) and outputs corresponding bounding boxes and annotations. To validate the framework, a synthetic dataset of 7533 images was generated and used to train the state-of-the-art object-detection model RFDETR. Experimental results demonstrate that the model trained on our synthetic data achieves a Mean Average Precision (mAP) of 80.9\% on real industrial testing samples. This framework provides a zero-cost, rapid-deployment solution for automating defect inspection in printing lines without requiring massive manual data collection.

Surprisal Theory is Tautological (without Rational Grounding) cs.CL

Surprisal theory holds that the human processing difficulty of a linguistic unit in context is an affine function of its surprisal under some language model. I argue this claim is a tautology without further constraint: for any non-negative difficulty measure over units in context, there exists a language model whose surprisal is an affine function of it under mild technical conditions. Therefore, because any pattern of difficulty is consistent with some language model, without an additional constraint on the language model, surprisal theory makes no falsifiable predictions. The tautology was long obscured by an assumption implicit in two decades of psycholinguistic work---that the relevant language model is the distribution that generated the training corpus, so that improving corpus fit improves predictions of human behavior. Recent empirical work has undermined this assumption, demonstrating that better corpus models can be worse predictors of processing difficulty. I conclude that breaking the tautology requires a rationalist intervention, i.e., the relevant language model must be derived from a non-empirically motivated model of the comprehender, which could be based on, for instance, memory constraints or processing goals, and that, thus, does not depend on the behavioral data surprisal theory is meant to explain.

BenchmarksFull tables
Artificial AnalysisIntelligence Index

Composite score across coding, math, and reasoning

#ModelScoretok/s$/1M
1Claude Fable 559.959$20.00
2GPT-5.6 Sol58.962$11.25
3Kimi K357.133$6.00
4Claude Opus 4.855.761$10.00
5GPT-5.6 Terra55122$5.63
SWE-rebench

Agentic coding on real-world software engineering tasks

#ModelScore
1OpenAIgpt-5.5-2026-04-23-xhighModel62.7%± 0.91%
2JunieJunieAgent61.6%± 0.64%
3OpenAICodexAgent60.4%± 1.37%
4AnthropicClaude CodeAgent59.6%± 1.98%
5OpenAIgpt-5.5-2026-04-23-mediumModel58.9%± 0.78%