OpenAI is winning market share through integration, not model superiority. The company is threading AI into the friction points where users already work, embedding virtual try-ons and shopping integrations directly into ChatGPT while Shopify builds Canvas for conversational store construction and ServiceNow launches Flow as a standalone service desk across Teams, Slack, and email. This pattern holds across the field: the companies capturing value aren't necessarily those with the largest models but those closest to where work actually happens. GitHub's trending repos confirm this shift. Frameworks like Ponytail and Skills let agents leverage existing code and conventions rather than reasoning from first principles. LocalAI runs models without GPU infrastructure. ComfyUI-MCP and video toolkits push agents toward media output. The unglamorous work of token management and tool sandboxing dominates adoption, not flashy capability announcements.
This integration-first strategy masks a deeper structural tension. Enterprise AI deployment remains constrained by semantic clarity at scale. Companies still spend days arguing about what a "customer segment" means before building a report, and most enterprises lack the data infrastructure required for agents to behave reliably across platforms. Google and Microsoft backing Apache Ossie, Omnissa shipping Elara to break data silos, and the broader push for interoperability all signal recognition of this bottleneck. The competitive moat forming is not about model intelligence but about data cleanliness and integration depth with existing workflows. Meanwhile, infrastructure is bifurcating. Google restricts Gemini 4 Argon to a trusted set through its Fairwind Program while Satlyt raises $8 million to build open satellite software positioned against SpaceX's closed model. Micron warns memory shortages will persist through 2028, and Amazon is offloading $8 billion in Nvidia chips to investors. Capital intensity remains severe, and the pressure to monetize quickly is reshaping vendor positioning.
Across research and labs, the field has moved past capability races toward decision-making reliability and economic justification. NVIDIA now sells GPUs as infrastructure for measurable return on investment, not abstract compute. Anthropic competes on domain utility rather than raw capability or cost. Research in systems control shows uncertainty quantification becoming foundational, with learned models increasingly decomposed into fast inner loops with formal guarantees and slow outer loops that adapt. None of the major labs are racing to the bottom on cost, which implies either that current pricing sustains business models or that moats are strong enough to avoid price competition. What matters now is not whether a model fits training data but whether it solves downstream decision problems under distribution shift and in failure regimes not visible in aggregate metrics.
Grant Calloway
Accurately learning nonlinear dynamics from a finite-duration experiment requires the efficient collection of informative data. We address this challenge for stochastic controlled nonlinear dynamical systems whose state is observed along a single trajectory. Our goal is to reconstruct the unknown controlled state-increment map over a prescribed compact subset of state-input space. We construct a parametric estimator of the map using fixed nonlinear features, so that the model is nonlinear in the state and input, but linear in the unknown parameters. A Gaussian prior over the parameters yields recursive Bayesian posterior updates as data stream in, enabling online quantification of predictive uncertainty in the reconstructed dynamics over the target set. We formulate an optimal adaptive-design problem over an information state, using a prediction-oriented acquisition criterion based on the mean marginal mutual information between candidate future trajectories and the reconstructed dynamics over the target set. We then approximate the resulting adaptive-design problem by a non-myopic receding-horizon formulation, evaluate its remaining expectation using a scenario-based sample average, and solve the resulting deterministic program with the cross-entropy method, leveraging parallel candidate-scenario evaluations. Numerical experiments on a noisy multistable system demonstrate that the proposed adaptive information-seeking strategy reduces predictive uncertainty and reconstruction error more efficiently than common excitation baselines under comparable experimental constraints.
We present a representation-learning framework for composite adaptive tracking control under dynamically coupled disturbances. The framework connects classical disturbance-accommodating control (DAC) to recent last-layer adaptive disturbance-rejection methods. Specifically, we introduce a statistically principled hard expectation-maximization (hard-EM) procedure, with a Kalman smoother in the hard E-step, to identify dynamical representations of disturbance whose latent evolution is uniformly contractive. The learned representation evolves a latent disturbance-excitation state from measured plant features and control inputs and decodes that state into the time-varying disturbance acting on the nominal plant, thereby extending prior "fixed-decay" last-layer adaptive methods to a learned, predictive DAC-style formulation. Combined with Bayesian filtering of the learned latent state, this representation yields a composite adaptive tracking controller with predictive capability and provable exponential convergence to a bounded neighborhood. We validate our approach experimentally on a slippery ground vehicle carrying a liquid-sloshing tank and a pendulum load, and we further assess its robustness on a system of coupled Duffing oscillators. Across both settings, the method achieves accurate disturbance prediction and improved overall tracking performance relative to fixed-decay representation-learning ablations, LTI disturbance-accommodating baselines, and model-based PD baselines.
We introduce GridSFM, a framework that combines a pretrained foundation model across grid topologies with physics-informed fine-tuning for solving AC Optimal Power Flow (AC-OPF) at scale. It is a $15$ million parameter physics-inspired graph neural network pretrained across $54$ topologies of $500$ to $4{,}000$ buses. Our model attains a $2.45\%$ zero-shot generation-cost error on a $10{,}000$ bus case held-out operating conditions with no degradation as system size grows. Building on this, we pair the pretrained backbone with a physics-informed fine-tuning design based on Newton's method for power flow. With only $100$ solved instances, GridSFM adapts to unseen grids up to $10{,}000$ buses. We show it out performs single topology, dedicated neural network models that are trained more data, both in terms of cost and solver iterations when deployed as warm starting points. In designing this foundation model, we overcome the fact that the feasible set for AC-OPF can be disconnected. This is an obstruction that prevents any continuous neural network from approximating the solution map. To do so, we lift the problem and relax its constraints with logarithmically penalized slacks. We prove that the resulting elastic feasible set is contractible, that the AC-OPF minimizers remain minimizers of the elastic problem above an explicit penalty threshold, and that projecting an approximate solution back onto the AC-OPF feasible set is well posed. We release all models, data, and code so that the community can build on a shared starting point for AC-OPF.
Open RAN (O-RAN) slicing xApps must adapt resource allocations to changing channel conditions and traffic demands while meeting service-level agreements (SLAs). Deep reinforcement learning can produce adaptive policies, but their allocation rules remain encoded in neural-network parameters. Our goal is to retain this adaptability while making the controller's decision logic directly inspectable and editable by operators. We use a large language model (LLM) to evolve slicing controllers as compact Python programs whose decision logic remains readable and editable after optimization. The LLM proposes and revises candidates offline, while a calibrated simulator scores them, and the selected decision module runs unchanged in the O-RAN control path. On the NSF POWDER 5G testbed, the evolved controller releases resources from a guaranteed slice whose throughput target becomes unattainable under a sustained channel fade, improving best-effort throughput from 158.2 to 228.6 Mbps, a 44.5% gain over the best static allocation. Since the controllers are readable source code, their behavior can be predicted from their equations, defects can be diagnosed by reading the code, and calibration errors can be corrected with one-line edits, reducing SLA misses from 79.9% to 2.2% in one case and more than doubling fitness in another. In a four-slice trace-driven simulation calibrated to the same testbed, evolutionary search achieves higher average evaluation scores than independent prompting at a matched proposal budget, with mean normalized gains on held-out traces of 16.3% for prompting alone, 32.1% for evolution from scratch, and 51.0% for evolution from a starting program.
Studies of machine-learning-based power system protection are difficult to compare because task definitions, measurement access, data partitions, metrics, and generalization conditions often differ. EvEMTBench addresses this gap with an open, executable, and versioned benchmark that fixes these evaluation choices while leaving model design open. Across four grids spanning 20-345 kV, it defines 12 protection and event-analysis functions instantiated as 24 scored tasks and supports structured evaluation across observability conditions, predefined distribution shifts, and zero-shot and fine-tuned cross-grid transfer. Committed partitions, leakage controls, and reproducible reporting provide a common basis for comparing future methods. A reference evaluation spanning trivial, conventional, feature-based, and deep-learning baselines shows that wider observability is not uniformly beneficial, shifted conditions can reveal failures not apparent in-distribution, and cross-grid transfer is substantially stronger for fault detection than for fault localization. Protection-relevant diagnostics identify failure modes not apparent from primary metrics alone. EvEMTBench therefore makes generalization in machine-learning-based protection an explicit and reproducible evaluation problem.
State tracking from sequential observations can require both retaining information and updating it by composing observed operations. We extend Mamba-3's diagonal transition with an input-dependent low-rank reflection term to support noncommutative state tracking, in which the order of operations matters. The rank-one update couples state coordinates along an input-dependent direction, enabling non-diagonal state transitions within a single Mamba-3 block. The extension preserves Mamba-3's exponential-trapezoidal discretization, rotary embeddings (RoPE), and readout. For training, we adapt chunkwise computation to parallelize the proposed recurrence within each chunk. Experiments cover group word problems with discrete inputs and a shell game with continuous observations, in which a policy is trained by behavioral cloning. Among the models selected for their strong performance under fixed timing, the proposed model maintains higher tracking success on longer swap sequences in the shell game with continuous observations and timing jitter. These experiments show that the proposed method achieves high accuracy on the evaluated non-commutative tracking tasks, improving on standard Mamba-3. The extension thus offers a Mamba-3-based approach to non-commutative state tracking.
Composite score across coding, math, and reasoning
| # | Model | Score | tok/s | $/1M |
|---|---|---|---|---|
| 1 | Claude Opus 5.5 | 57.6 | 97 | $8.00 |
| 2 | Claude Sonnet 5.5 | 56 | 145 | $4.00 |
| 3 | Claude Fable 5.1 | 53.4 | 70 | $20.00 |
| 4 | GPT-6 Astra | 52.7 | 56 | $20.00 |
| 5 | Gemini 4 Argon | 52.6 | 0 | $4.00 |
Agentic coding on real-world software engineering tasks
| # | Model | Score |
|---|---|---|
| 1 | AnthropicFable 5 [high]Model | 64.5%± 1.41% |
| 2 | GrokGrok 4.5 [high]Model | 63.8%± 0.60% |
| 3 | AnthropicOpus 5 [high]Model | 63.4%± 1.35% |
| 4 | Z.aiGLM-5.2 [high]Model | 62.9%± 1.19% |
| 5 | OpenAIGPT-5.6 Sol [medium]Model | 62.3%± 1.83% |
Makes your AI agent think like the laziest senior dev in the room. The best code is the code you never wrote.
My personal directory of skills, straight from my .claude directory.
OpenShell is the safe, private runtime for autonomous AI agents.
Firebase SDK for Apple App Development
Multi-agent harness that runs Claude Code and Codex together as one system
Design, conduct and analyze results of AI-powered surveys and experiments. Simulate social science and market research with large numbers of AI agents and LLMs.
Taranis AI is an advanced Open-Source Intelligence (OSINT) tool, leveraging Artificial Intelligence to revolutionize information gathering and situational analysis.
A collection of research studies centered on Modality Missing Learning (MML) (also referred to as Incomplete Multimodal Learning).
🏆 An awe-inspiring collection of resources, encompassing a wide range of tools, documents, resources, applications, and use cases related to ChatGPT.
Interactive 3D visualization platform for exploring transformer architectures, tensors, and real-time LLM inference.