The Inference Report

October 3, 2026

The week's most important story isn't about what AI models can do but about who controls the infrastructure where decisions happen and data lives. Apple is tightening full-disk access permissions for AI agents, Meta is embedding Muse into televisions and toasters, IBM is moving Bob on-premises for companies that refuse to let data leave their networks, and Amazon is clearing regulatory approval for a 3,000-acre data center while taking criticism for downplaying its environmental footprint. These aren't separate announcements. They're competing moves in a single fight over the data layer and customer relationship when AI agents start acting on behalf of people.

The incentive structure reveals itself in the strategy choices. Meta's decision to open-source Muse while pushing device integration isn't altruism but a bid to own the interface layer before Apple, Google, or Amazon do. The company whose AI reaches the most wearables and household devices owns the daily interaction with users and the data that flows from it. Apple's restrictions on full-disk access aren't safety theater but a deliberate block against agents becoming a backdoor for competitors to access file systems and message histories. IBM's on-premises deployment targets regulated industries where data sovereignty is non-negotiable. The real competition isn't between models. It's between platforms fighting to be the trusted intermediary between users and their own information. OpenAI is pursuing production lock-in through bundled model selection and workflow redesign that shows concrete time savings. NVIDIA is pushing local deployment as open models become more capable and compact, creating an alternative to cloud-dependent workflows. Anthropic is investing $100 million to train 10,000 engineers, a supply-side move that signals confidence in enterprise demand while attempting to shape how developers think about AI integration.

The developer infrastructure tells the same story. GitHub's trending repositories cluster around agent optimization and reliability, not proof-of-concept. Agent-Reach eliminates API costs for sensory input, Caveman and context-mode attack token waste through compression and intelligent windowing, and skills frameworks from Google, Sentry, and independent developers are standardizing a pattern of encapsulating domain knowledge as composable, reusable units rather than embedding logic in prompts. This is infrastructure thinking applied to AI. A secondary trend toward localization is equally visible. Nanobot runs self-hosted with Python and WebUI, Openmed and vllm-mlx emphasize on-device execution for privacy and latency in regulated contexts, and medical tools like StatsPAI's causal inference library indicate that agent adoption is moving into domains where correctness and auditability matter more than speed. The hype phase is over, and what's gaining traction now are the unglamorous tools that make agents actually work at scale.

The regulatory response isn't keeping pace with the actual problem. The FTC is sending formal demands to Anthropic and OpenAI following incidents where agents breached corporate systems. The White House is rebranding AI as "super intelligence" and collecting CEO signatures on a safety pledge, a loyalty test dressed as governance that costs executives nothing and carries no enforcement mechanism. Meanwhile, NVIDIA's chip-smuggling arrests and memory price spikes on Galaxy S26 phones reveal what happens when supply becomes a weapon in platform competition. The company's foundry advantage in AI chips is creating an enforcement problem: it must prevent its own products from reaching competitors' markets while simultaneously raising the cost of every device that uses memory. The methodological shift in research and benchmarking toward diagnostic rigor and fine-grained evaluation reflects the same underlying reality. Papers on efficiency through structural simplification, learning from intermediate signals, and evaluation protocols that probe mechanisms rather than aggregate scores all point to a field moving past capability claims toward the unglamorous work of making systems reliable enough to deploy. The SWE-rebench rankings show no movement in the top tier, suggesting either a performance plateau or evaluation constraints that prevent differentiation. The broader pattern is clear: the infrastructure wars have begun, and the companies that own the data layer and control where computation happens will determine who profits from AI deployment, not the companies that built the most capable models.

Grant Calloway

AI LabsAll labs
From the WireAll feeds
Research PapersAll papers
One Basis to Animate Them All: Gaussian Blendshape Distillation for Real-Time Avatars cs.CV

3D Gaussian avatars support fast rendering, however, their real-time animation is often challenged by the costly neural inference. We address this bottleneck and show that the animation of pretrained avatar models can be closely approximated by a linear combination of identity-independent blendshapes. Building on this finding, we introduce GALA (Gaussian Animation via Linear Approximation), a distillation method that replaces per-frame heavy neural decoding with a shallow coefficient predictor and a linear blend. To improve fidelity and reduce memory requirements, we propose to construct the basis using block-local PCA under a rendering-aware metric and a memory budget. Our method learns a shallow MLP network to predict blendshape coefficients and applies to various animation architectures without retraining original models. We validate GALA by accelerating the inference of three distinct avatar models for 3D animation of facial expressions and full-bodies with clothing dynamics. Across these models, our distillation generalizes to held-out identities and reduces CPU animation cost by up to three orders of magnitude while preserving most of the rendering quality. Excellent results of our method confirm the shared linear structure of learned avatar representations and enable highly efficient and accurate animation at frame rates reaching up to 60fps on mobile devices. Project page: https://ramazan793.github.io/gala/

KaliBench: A Fine-Grained Benchmark for Cybersecurity Tool Use on Kali Linux with Runtime-Free Verifiable Rewards cs.CL

LLMs are increasingly applied to cybersecurity workflows, where they are expected to translate analysts' intent into tool invocations. However, existing evaluations focus on knowledge-based assessments or end-to-end agentic tasks, and do not directly measure LLMs' ability to generate executable commands for real-world cybersecurity tools. This gap is critical because cybersecurity operations rely on strict command-line interfaces (CLIs), where minor syntax errors, incorrect flag--value bindings, or argument misordering can invalidate execution. We introduce KaliBench, a fine-grained benchmark and dataset for natural-language--to--CLI translation on Kali Linux, comprising 8,504 query--command pairs spanning 1,642 tools across 23 capability dimensions and 5 security phases. KaliBench is constructed via a manuscript-grounded pipeline with deterministic canonicalization and alias-aware evaluation, enabling precise and reproducible assessment of tool selection and argument construction. To ensure both semantic correctness and practical executability, we develop a multi-stage verification pipeline that combines LLM-based validation, sandboxed terminal execution, and human-in-the-loop refinement. Building on these fine-grained, deterministic signals, KaliBench further enables runtime-free verifiable rewards for training. Across three evaluation modes and 24 configurations of general-purpose and security-focused open-weight models, no open-weight model exceeds 42% exact-command accuracy in the unrestricted setting, highlighting the difficulty of accurate CLI-based cybersecurity tool use without explicit tool hints. We further show that supervised fine-tuning and reinforcement learning with verifiable rewards derived from KaliBench significantly improve an 8B model and achieve performance comparable to a 685B MoE model.

Reconstruct, Practice, Go Real: Guided Self-Improvement for Embodied Agents cs.RO

Building reliable robot capabilities across diverse tasks requires substantial human effort to develop and maintain skills, design rewards, and integrate perception with control. We present Reconstruct, Practice, Go Real (RPG), a framework for autonomous improvement of robot execution systems without updating model weights. RPG identifies manipulation capabilities in an offline dataset and constructs related practice tasks in simulation. During practice, RPG uses execution feedback, privileged simulator state, and available dataset videos to diagnose failures. It develops new reusable symbolic skills, refines existing skills, and revises the system prompt based on these diagnoses. Cross-task evaluation tests individual candidate changes and merged revisions before they are retained for reuse. At test time, a multimodal LLM uses the resulting system prompt and skill library to coordinate perception and robot control. On held-out initializations of 22 manipulation tasks, RPG improves task success from 28.6% after the first practice round to 95.0% after 15 rounds, outperforming all evaluated baselines, including ASPIRE (75.5%) and CaP-Agent0 powered by GPT-6 Astra Pro (60.0%). After a common calibration and hardware-adaptation procedure, the frozen system succeeds in all 30 physical trials, with ten trials on each of three tasks. Project Website: https://rpg-robot.github.io/

Embedding Prediction Helps Image Generation cs.CV

In diffusion transformers, a class label or a text prompt is embedded once, and the same condition is reused at every denoising step. We ask whether predicted embeddings can serve as this condition instead. Next-Embedding Predictive Autoregression (NEPA) trains a Transformer to predict the next continuous embedding in a sequence. In generation, the clean image follows the noisy image, so its embeddings are the next embeddings after the condition and the noisy image. We train a NEPA model to predict them all at once with Multi-Embedding Prediction, and in Embedding Conditioned Generation, a DiT generator is conditioned on these predictions, recomputed at every denoising step, so the conditioning signal adapts to the current noisy state. Experiments on class-conditional ImageNet $256\times256$ study the condition of the generator, the design of Multi-Embedding Prediction, and the scaling of both models. The NEPA model adds a second network to every sampling step; with it, and combined with REPA, our final model, NEPA-DiT-XL, reaches an FID of 1.32 using about a third of the training compute of REPA.

ScholarCatalyst: A Benchmark for Retrieving Papers That Inspire New Research cs.AI

What makes great scientists great? Even as AI systems start to make progress on open problems, scientists remain far ahead of them at sensing which prior idea, buried in an ever-growing archive of research, a new problem needs. To study this skill, we draw on researchers who know firsthand which earlier work advanced their completed projects, with papers serving as pointers to the ideas within. Using our automated pipeline that makes author annotation scalable, we build ScholarCatalyst by having 184 lead authors of 207 recent computer science papers label which candidates did or could have advanced their project, each with a detailed rationale. We introduce a retrieval task with author-provided judgments: given an initial research question, retrieve these papers from only the literature available when the project began. Agentic search does no better than embedding retrieval (0.42 vs. 0.48 Recall@20) despite calling that same retriever as a tool. Even an agent built on Claude Fable 5.1, which may have seen the completed papers during training, reaches only 0.51 R@20. These results highlight the need for new training recipes that equip models with expert intuition for searching broad corpora. We envision ScholarCatalyst as a step toward scientific agents that can take a half-formed idea and point to the prior research it needs.

SILSA: Sliding-Window Slice Latents for Topology-Preserving High-Resolution 3D Generation cs.CV

High-resolution 3D generation increasingly relies on voxel latents and multi-stage pipelines that first predict active structure and then synthesize local geometry. While effective, this design fragments continuous surfaces into many local tokens, inflates generation cost, and often weakens topological consistency for thin or highly connected shapes. We introduce SILSA, a topology-aware 3D generation framework that represents shapes with compact sliding-window slice latents. Instead of generating expensive voxel tokens, SILSA uses a fixed set of overlapping slices along the three canonical axes, where each token summarizes a local depth window to preserve cross-sectional continuity and support single-stage rectified-flow generation. A Slice VAE encodes oriented surface samples into multi-axis slice latents and reconstructs them with a sparse volumetric decoder, while a Volumetric Anchor Lattice coordinates directional slice streams through a shared 3D workspace. To preserve structural correctness, we introduce slice-level topology supervision that matches persistence diagrams and aligns Betti transitions across neighboring slices. Experiments show that SILSA improves structural fidelity while substantially reducing generation cost. SILSA improves PSNR by $8.7\%$, coverage by $5.96$ absolute points, and Betti error by $9.2\%$ over the strongest baseline, while using $70.0\%$ fewer tokens than the next-most compact baseline and over $98\%$ fewer tokens than sparse or hierarchical tokenizers, effectively reducing training memory by $40.4\%$ and inference time by $58.5\%$. Qualitative results further show improved preservation of thin structures, repeated components, and long-range connectivity.

BenchmarksFull tables
Artificial AnalysisIntelligence Index

Composite score across coding, math, and reasoning

#ModelScoretok/s$/1M
1Claude Opus 5.557.698$8.00
2Claude Sonnet 5.556138$4.00
3Claude Fable 5.153.471$20.00
4GPT-6 Astra52.756$20.00
5Gemini 4 Argon52.60$4.00
SWE-rebench

Agentic coding on real-world software engineering tasks

#ModelScore
1AnthropicFable 5 [high]Model64.5%± 1.41%
2GrokGrok 4.5 [high]Model63.8%± 0.60%
3AnthropicOpus 5 [high]Model63.4%± 1.35%
4Z.aiGLM-5.2 [high]Model62.9%± 1.19%
5OpenAIGPT-5.6 Sol [medium]Model62.3%± 1.83%