Enterprise AI is moving faster than governance can follow, but the fractures aren't random. They're opening precisely where control matters most: between agents, across borders, inside closed platforms, and in the spaces where liability still lives. The protocol layer has become the real concentration point for risk. MCP, the agent-to-agent communication standard spreading through the industry, carries trust gaps that let malicious prompts propagate sideways through entire fleets. Security researchers spotted what appears to be a coordinated agent swarm running on Tencent infrastructure targeting Alibaba's map service, and South Korea's president is warning that AI models are now active tools in bank cyber attacks. These aren't separate incidents. They're the same failure materializing at different scales: once agents can talk to each other, the perimeter collapses. The industry built the plumbing for autonomous systems to operate at machine speed, then discovered the plumbing has no locks.
Regulation is arriving in fragments, each piece revealing where power actually sits. The EU gets watermarks on ChatGPT text while OpenAI simultaneously launches visual ads inside image generation results, embedding commerce deeper into the inference loop. Reflection AI released Beam, a 501B sparse Mixture-of-Experts model explicitly pitched at enterprises and sovereign nations that want to build local AI factories on proprietary data. This is the real competition: not model quality but infrastructure sovereignty. Chinese models and open-weight alternatives are eroding the moat that closed APIs once provided. Meanwhile, liability questions remain unsettled. Insurers and lawyers are weighing potential massive lawsuits against OpenAI and Anthropic leadership over rogue AI behavior, but the legal theory doesn't exist yet to make those claims stick.
The benchmarks tell a revealing story about where the field actually stands. SWE-rebench shows no movement from the previous cycle, with AnthropicFable 5 maintaining its lead at 64.5 percent and the top seven models holding identical positions and scores. The Artificial Analysis benchmark, by contrast, exhibits substantial churn throughout its 468-entry list, with models reordering across the full range. The divergence suggests the two benchmarks are capturing different aspects of model behavior or operating under different measurement precision rather than converging on a shared picture. Developer work is splitting into two parallel tracks. One cohort is building infrastructure for AI agents, persistent memory systems, web scraping capabilities, video production pipelines, treating agents as a new class of software that needs its own tooling. The second cohort is solving older problems with renewed urgency: vector databases, web servers, and specialized inference engines gaining traction because they do one thing well. Sustained technical investment is flowing toward the databases, search engines, and inference layers that make agents actually work, while viral momentum goes to agent orchestration and agentic video production.
The capital markets are starting to price in the uncertainty. Global venture funding hit 159 billion dollars in Q3 2026, a record for billion-dollar rounds, but pension funds are cutting US equities over AI concentration risk and bond markets are selling off corporate debt. Executives are asking CFOs the question that matters: where is the ROI? The answer, for most organizations, is still missing. There's a lag between deployment and value capture, and that gap is where the risk lives. Not in the models themselves, but in the fact that no one has built the measurement systems yet to know whether any of this is actually working.
Grant Calloway
Pipeline figures in ML papers must be repurposed across many canvases, including paper columns, 16:9 slides, portrait posters, 1:1 social teasers, 9:16 phone previews. Each format imposes a different aspect ratio on the same computational graph, where any silently broken connection misrepresents the method. We formulate aspect-ratio-adaptive flowchart relayout as a distinct task: given a raster flowchart and a target ratio, produce a structurally faithful, hallucination-free, editable layout. Existing methods fail characteristically: image-to-image models stretch blocks and reject extreme ratios, text-to-image agentic systems hallucinate content, and parse-then-render systems mis-route edges. We propose an agentic pipeline factored into Parse, Style, and Layout stages, each pairing a main agent with a critic that combines deterministic constraint checks with VLM visual feedback so connectivity is explicitly checked and prevented from being silently broken. Outputs are draw.io-editable mxGraph XML. On a curated benchmark of 100 flowcharts at five aspect ratios, evaluated by Gemini 3.1 Pro and validated against human judgments, our method reaches 68.6% Content Fidelity versus 11.2-41.4% for prior work. Project page: https://onefigureeverycanvas.vercel.app/
In this paper, we study how training data creates associations between the tokens at the start of a base model's response and the reasoning behavior that follows. First, we demonstrate that fixing particular starting token cues makes a base model's performance competitive with that of its reinforcement learning (RL)-trained counterparts on math and coding. For instance, the cue ".\n\nOkay" raises Olmo-3-7B's MATH-500 pass@1 accuracy from 42% to 78%, while "Alright," raises Qwen3-14B's from 72% to 87%. Second, RL makes these cues more likely, while fixing them recovers much of its performance gain over the base model. Third, we trace the reasoning effects of token cues to the training data. We perform causal data interventions to turn an arbitrary word, such as "chicken", into an effective reasoning cue, or remove an existing cue's effect. A similar edit makes the prompt instruction "Think duck duck goose" as effective as "Think step by step" at eliciting reasoning. We also find that the hidden state representations induced by different cues correlate with different document types from the training set. Finally, we extend our study of token cues with a case study in language model safety, finding that different cues elicit distinct refusal and compliance behaviors that correspond to different types of training data.
Worst-group accuracy (WGA) evaluates a trained predictor but does not characterize how its frozen backbone behaves when a new head is learned. We introduce BiasFlow, a hook-based toolkit for monitoring class-attribute centroid alignment (IBMI), within-class centroid separation (W-IBMI), and feature-projection sensitivity. IBMI is confounded by class-attribute correlation and is not a measure of causal feature reliance. We pair these diagnostics with BiasFlow Regularization (BFR), a supervised, composable class-conditional centroid-alignment penalty. W-IBMI verifies the quantity BFR optimizes; it is scale dependent and does not independently establish attribute removal. Across the reported small-scale benchmarks, adding BFR improves or preserves mean WGA, with gains up to +26.0 pp on UrbanCars. The principal independent stress test freezes CelebA-Std backbones and trains fresh heads on biased data: BFR+GroupDRO improves WGA from 40.7% to 64.1%, while Male probe accuracy decreases from 92.5% to 72.2%. Attribute information remains recoverable, and cross-task results are mixed. A controlled synthetic-watermark ImageNet experiment additionally improves watermark-shift accuracy by +23.0 pp under matched training. These results support evaluating centroid geometry and resistance to biased head retraining alongside WGA, within the tested protocols.
Multimodal Diffusion Transformers (MM-DiTs) jointly process visual and textual representations throughout generation. These models repeatedly update the text tokens through multimodal attention, forming dynamic contextual tokens whose function is not well understood. In this work, we introduce a framework for reading this contextual space through natural-language interrogation. We train a lightweight bottleneck network that maps intermediate contextual tokens into the input space of a frozen Large Language Model (LLM), allowing the LLM to answer questions about the emerging image directly from these hidden representations. Our reader reveals that contextual tokens encode a rich, global representation of the emerging scene: generation-specific semantics, including attributes left underspecified by the prompt, are accessible surprisingly early in denoising, while increasingly fine-grained details become readable over time. Remarkably, this information remains decodable even when the MM-DiT receives an empty prompt, showing that contextual tokens accumulate substantial image-specific information from the evolving visual representation itself. We further find that generations with more readable contextual representations tend to receive higher human-preference scores. Building on these observations, we introduce Contextual Alignment, a training technique that explicitly reinforces the visual-semantic information encoded in the contextual tokens, improving generation quality and distributional coverage. Together, our results establish contextual tokens as both an interpretable view into the internal dynamics of MM-DiTs and an effective target for improving generative models.
LLM agents that orchestrate frozen vision-language-action (VLA) policies improve across episodes through text memory, which records what the agent did but not how the task is done. A demonstration video shows it, but fits poorly into an agent's context. The full video slows every turn, fixed keyframes lose the contact detail that decides whether a grasp holds, and what the agent needs shifts from the task's structure while planning to the frames around each contact. We introduce Recursive Video In-Context Learning (RV-ICL), a training-free method that turns a demonstration into a hierarchy the agent navigates rather than a prompt it receives. The hierarchy is built from the sub-events of the demonstration, such as grasps and releases. Its levels grow finer, from keyframes of the whole task to phases, moments and short clips, and are exposed through read-only tools. The agent reads the coarse levels before planning. During execution it re-enters the hierarchy whenever a step needs more detail and loads only the clip of its current sub-goal. One demonstration per task is enough. Built on RPent, RV-ICL raises success from 92.6% to 96.5% on LIBERO-PRO and from 86.7% to 95.8% on LIBERO-Plus.
Some diffusion posterior samplers construct Gaussian-tilted intermediate distributions along the reverse process. We observe that these targets can be pulled back to clean-space posteriors with weaker conditioning, with samples transported analytically to the corresponding noisy-space target through a Gaussian bridge. For the sequential Monte Carlo (SMC) sampler MCGDiff, the effective observation variance of this pulled-back problem is up to twice the diffusion-noise variance. We exploit this structure to initialize MCGDiff directly at an intermediate time: an approximate solver samples the softened clean-space posterior, the Gaussian bridge maps these samples to the tilted target, and only the remaining SMC suffix is run. This trades asymptotic consistency for finite-particle performance. With moment-matching posterior sampling (MMPS) as the solver, the hybrid improves sliced Wasserstein distance by roughly $2\times$ at matched particle count on a structured Gaussian-mixture inverse problem, and by more than an order of magnitude when the posterior-relevant mode is rare under the prior. A prior-initialization control, which retains the bridge but drops the clean-space conditioning, shows that on MCGDiff's standard Gaussian-mixture benchmark most of the improvement is insensitive to the conditioning. Conditioning the initialization gives a further consistent gain on the structured problem, and becomes decisive on a rare-mode problem, where resampling cannot repopulate a mode absent from the initial population.
Composite score across coding, math, and reasoning
| # | Model | Score | tok/s | $/1M |
|---|---|---|---|---|
| 1 | Claude Opus 5.5 | 57.6 | 97 | $8.00 |
| 2 | Claude Sonnet 5.5 | 56 | 137 | $4.00 |
| 3 | Claude Fable 5.1 | 53.4 | 67 | $20.00 |
| 4 | GPT-6 Astra | 52.7 | 64 | $20.00 |
| 5 | Gemini 4 Argon | 52.6 | 0 | $4.00 |
Agentic coding on real-world software engineering tasks
| # | Model | Score |
|---|---|---|
| 1 | AnthropicFable 5 [high]Model | 64.5%± 1.41% |
| 2 | GrokGrok 4.5 [high]Model | 63.8%± 0.60% |
| 3 | AnthropicOpus 5 [high]Model | 63.4%± 1.35% |
| 4 | Z.aiGLM-5.2 [high]Model | 62.9%± 1.19% |
| 5 | OpenAIGPT-5.6 Sol [medium]Model | 62.3%± 1.83% |
Next generation e2e testing framework for web and mobile apps.
A Claude Code plugin that automatically captures everything Claude does during your coding sessions, compresses it with AI (using Claude's agent-sdk), and injects relevant context back into future sessions.
A collection of agent skills for CAD, robotics and hardware design
Tool for automatic PS5 executables porting to Linux and Windows
An all-in-one, pure C++ inference engine for audio models, powered by ggml. Supports TTS, STT, VAD, voice conversion, music generation, and more, with highly optimized performance. No Python dependency.
Next-generation Albumentations: dual-licensed for open-source and commercial use
A framework for high-performance medical image processing, neural network inference and visualization
🚀 World's largest GPT Image 2 prompt library, updated daily — 2000+ curated prompts with preview images, 16 languages. OpenAI's next-gen image model with pixel-perfect text rendering, cross-image consistency, and commercial-grade illustration. Free & open source.
A persistent workspace for development work that self-improves and continues beyond one session.