The Inference Report

August 16, 2026

The infrastructure supporting AI development is consolidating around compute, data, and distribution channels while safety and misuse problems remain largely unmonitored and unaddressed. SpaceX's acquisition of Cursor signals that coding AI is no longer a standalone business but a feature bolted onto existing power structures with capital and scale. Anthropic's watermarking announcement addresses a symptom rather than the disease: Claude can now stamp its outputs, but enforcement remains unclear and watermarks may not survive editing or copy-paste. Amazon is harvesting Twitch streams for training data on an opt-out basis, inverting the default assumption about consent and treating data as a resource to be extracted rather than a choice to be made. The CSAM abuse case using Grok and the financial advice study showing blind spots in ChatGPT, Claude, and Perplexity both point to the same friction: these tools are deployed at scale without adequate safeguards for misuse or error, yet the industry response defaults to technical markers and voluntary opt-outs rather than architecture that prevents harm. Malaysia's emergence as a data center hub completes the picture.

Computer vision research meanwhile is shifting from static, post-hoc interventions to trainable, evidence-aware mechanisms that integrate uncertainty measurement into the forward pass. Token pruning and weight-sharing strategies address efficiency bottlenecks in vision transformers by exploiting redundancy without retraining, while pseudo-label refinement acknowledges that self-supervised encoders saturate confidence distributions, requiring adaptive rather than fixed filtering thresholds. Efforts to ground language generation in visual evidence through structural constraints and region-level perception tokens yield measurable precision gains on long-form outputs, though external benchmarks expose domain-conditional failure modes. Long-horizon understanding consistently reveals that current multimodal models function as lossy summarizers, motivating hierarchical indexing and temporal reasoning rather than end-to-end generation alone.

On coding benchmarks, AnthropicFable 5 holds the top position on SWE-rebench at 64.5 percent with tight confidence intervals that rule out meaningful movement at the top tier, preserving a three-model hierarchy with GrokGrok 4.5 at 63.8 percent and AnthropicOpus 5 at 63.4 percent. Developer infrastructure is fracturing into specialized layers: Unsloth dominates with 72k stars by solving local inference on consumer hardware, Needle compresses models to 14MB for phones and wearables, and Soup enables fine-tuning on 4GB laptop GPUs. Browser automation and agent infrastructure are consolidating around practical tooling like Ego-lite and CLI-Anything that let agents coexist with existing workflows without wholesale replacement. The real competition is not about which model is smarter or safer but about who owns compute, controls the data pipeline, and can move fastest to put tools in users' hands before accountability catches up.

Grant Calloway

AI LabsAll labs

No lab headlines.

From the WireAll feeds
Research Papers — FocusedAll papers
On the Relaxation of Conditional Independence Assumption for Image Segmentation cs.CV

In semantic segmentation, a recent line of RankSEG methods directly optimizes Dice/IoU scores at inference time, improving alignment with evaluation metrics without modifying model training. Despite its theoretical and empirical success, RankSEG relies on the restrictive Conditional Independence Assumption (CIA), which ignores crucial label correlations and therefore degrades performance in ambiguous or low-contrast scenarios. However, accounting for full label dependence is computationally prohibitive, requiring $\mathcal{O}(d^3)$ time. To address this, we replace the CIA with a Spatially Localized Dependence (SLD) structure that captures local label correlations while keeping the dependence model tractable. We further overcome the remaining computational bottleneck via a Reciprocal Moment Approximation coupled with a novel fixed-point optimization strategy that eliminates exhaustive search. The proposed algorithm achieves a highly practical $\mathcal{O}(d \log d)$ complexity and consistently outperforms conventional argmax and CIA-based RankSEG across diverse segmentation benchmarks. Improvements are significant in low-contrast or small-object scenarios, where label dependence offers valuable signals complementary to image information for accurate segmentation. The code of experiments is available at https://github.com/ZixunWang/RankSEG-DEP.

TED:Text-Axis Evidence Decomposition for Prompted Anomaly Localization cs.CV

CLIP is a powerful vision-language model, but it was not designed for fine-grained defect localization; CLIP-based anomaly detectors therefore adapt it with prompts or lightweight modules to increase defect sensitivity. We show that stronger sensitivity does not necessarily make local evidence reliable: under domain shift, adapted CLIP-AD models often assign high anomaly scores to both true defects and visually complex normal regions. The issue is not simply missing defect information, but a local scoring rule that decodes defect and hard-normal evidence, having the same anomaly evidence. We propose TED (Text-Axis Evidence Decomposition), a post-hoc scoring method that asks whether each ambiguous response is better supported by source defect patches or by source normal patches mistaken as anomalous. TED compares these supports under the host's normal-versus-anomaly text response, leaves the backbone and prompts unchanged, and requires no target-domain training. It works as a train-free score for raw VLM backbones or as a source-calibrated residual correction for adapted CLIP-AD hosts. Across frozen VLM backbones, TED substantially improves pixel-level localization over raw prompt similarity; across adapted hosts, it improves most pixel-level settings over P-AUROC, P-PRO, and P-AP. Gains are largest under stronger hard-FP competition, with mean localization gain increasing from +5.0 in low-competition regimes to about +10.9 in mid/high-competition regimes. These results suggest that recoverable defect evidence can already exist in pretrained multimodal representations, but reliable localization requires decoding it against hard-normal competitors. Code will be released at TED GitHub repository.

BadAction: Backdoor Attacks on Interactive Video Generation via Action-Guided Triggers cs.CV

Interactive video generation (IVG) models have achieved remarkable progress in producing controllable visual content guided by user-defined actions, yet their security vulnerabilities remain largely unexplored. In this paper, we present the first systematic study of backdoor attacks against the interactivity of IVG models. Based on this attack surface, we propose BadAction, which leverages action-guided triggers to achieve the attack. Specifically, BadAction implants predefined motion patterns into the action sequences of backdoor samples and associates them with a static target video. Once triggered, the backdoored model generates frozen future frames that no longer respond to subsequent user actions, while preserving normal behavior on benign action sequences. In addition, we explore a stealthier attack in which multimodal triggers jointly poison action, text, and image inputs. Experiments show that BadAction achieves average attack success rates of 91.0% with action-only triggers and 80.4% with multimodal triggers. Moreover, extensive defense evaluations show that BadAction successfully bypasses existing backdoor detection methods, revealing a critical security gap in the interactive video generation pipeline. Project page: https://wsad55.github.io/badaction01/.

MEND: Label-Free Detection, Localisation, and Correction of Latent Hallucination in World Models cs.CV

World Models are appearing as the next major frontier in computer vision. However, their robustness is currently largely unexplored. We identify the phenomenon of hallucination in latent World Models: given a state and an action, the predicted next latent can decode to a scene that never occurs. Because the prediction is statistically ordinary and is fed back autoregressively by the model, the error is both silent and compounding. We study whether such latent hallucination can be detected, localised, and corrected at inference time, on a frozen self-supervised world model in the absence of ground-truth error labels. We introduce Masked Empirical-Bayes Neural Denoising (MEND), a single conditional score network trained by denoising score matching on real transitions, whose score field serves three roles: its magnitude detects hallucination, its per-token field localises it to specific image patches, and it defines an inference-time correction direction. On two navigation environments MEND detects hallucination with an AUROC of up to 0.80 without using actions, exceeding a single-Gaussian density baseline while also localising the error (per-token AUPRC up to 0.87) and correcting it, all from one score field. Our correction reliably reduces single-step latent error and improves predictions. We identify that a part of the error is tangent to the data manifold, hence, we focus on detection and localisation while highlighting promises of the correction.

Fiber-Resolved Microstructure Quantification from Multi-Shell Diffusion MRI using Detection Transformers cs.CV

Fiber orientation and compartmental microstructure are central to the characterization of white matter tissue in diffusion MRI, yet existing methods either resolve fiber orientations without quantifying microstructure, or quantify microstructure while assuming a fixed number of compartments and a single fiber direction. Nonparametric approaches that recover both require tensor-valued diffusion encoding and computationally expensive Monte-Carlo inversion of an ill-posed inverse Laplace transform. We propose to reframe this problem as an object detection-like task, adopting the Detection Transformer (DETR) architecture to jointly predict mean diffusivity (MD), fractional anisotropy (FA), main fiber direction, and signal fraction for a variable number of compartments per voxel from standard multi-shell diffusion MRI with linear encoding. Hungarian matching during training resolves permutation invariance across compartments. We introduce mean Average Precision as a reproducible benchmark metric. Evaluated on synthetic test data with up to five compartments per voxel, our model achieves $R^2=0.95$ for MD, $R^2=0.88$ for FA, and a median angular error of 4.2°, with performance scaling naturally with compartmental signal fraction.

Emergent Multi-View Geometry Through Self-Distillation cs.CV

Over a century ago, Henri Poincaré argued that a motionless observer cannot acquire the notion of space. Yet, most visual representation learning methods operate on individual images, while those that leverage multiple views rely on RGB reconstruction, entangling geometry with appearance. We propose Poincar3, a self-supervised method that learns representations from multiple views through self-distillation instead of RGB reconstruction. We combine masked patch and image-level distillation with a teacher that observes additional views, enabling training from scratch without explicit 3D supervision. Poincar3 outperforms both previous single and multi-view self-supervised approaches such as DINOv3, MuM, and Muskie on correspondence estimation, camera pose estimation, and 3D reconstruction. Using a lightweight Poincaré adapter, we also find that our learned features encode camera motion more accurately than existing self-supervised representations.

BenchmarksFull tables
Artificial AnalysisIntelligence Index

Composite score across coding, math, and reasoning

#ModelScoretok/s$/1M
1Claude Opus 563.149$10.00
2Claude Fable 562.160$20.00
3GPT-5.6 Sol60.961$11.25
4Grok 4.660.958$3.00
5Kimi K359.737$6.00
SWE-rebench

Agentic coding on real-world software engineering tasks

#ModelScore
1AnthropicFable 5 [high]Model64.5%± 1.41%
2GrokGrok 4.5 [high]Model63.8%± 0.60%
3AnthropicOpus 5 [high]Model63.4%± 1.35%
4Z.aiGLM-5.2 [high]Model62.9%± 1.19%
5OpenAIGPT-5.6 Sol [medium]Model62.3%± 1.83%