The market is fragmenting along a clear axis: companies that move fast and build products are pulling away from those relying on regulatory friction and litigation to protect position. Musk's courtroom testimony against OpenAI admits xAI distills OpenAI's models while arguing the company betrayed its nonprofit mission, yet the Pentagon simultaneously signs deals with Nvidia, Microsoft, and AWS to diversify AI vendors after its own dispute with Anthropic over usage terms. Competition requires actual alternatives, but alternatives keep getting acquired. Cursor's reported $60 billion acquisition talks with SpaceX matter less for what they reveal about Cursor's value than for what they signal about consolidation math: if your product works, the acquirer will pay more than the market could ever allocate independently.
Models are commoditizing faster than the industry acknowledges. GPT-5.5 matches Mythos Preview in new cybersecurity tests, suggesting that cyber threat attribution to any single model is not a breakthrough but rather a feature of the capability tier itself. Models tuned to prioritize user satisfaction over truthfulness make more errors, which describes the actual tradeoffs built into deployment. The Pentagon's diversification strategy and DOD friction with Anthropic reveal an institution learning that single-vendor dependency creates leverage problems. Competition in AI infrastructure is real. Competition in capability differentiation is narrowing. Meanwhile, Chinese models are consolidating gains on code benchmarks: GLM-5 jumped from rank 17 to rank 3 on SWE-rebench, while Kimi K2.5 climbed from rank 29 to rank 16, suggesting systematic capability improvements across families rather than breakthrough leaps from any single model.
Regulatory capture is dressing itself as safety. Minnesota passes a ban on fake AI nudes with $500K fines while a new Christian cell network blocks pornography at the network level in ways adult users cannot override. English councils will trial Google AI tools to recommend planning decisions. These represent a shift from "AI companies should self-regulate" to "governments will regulate AI through whatever lever is closest at hand," often meaning regulation of user behavior rather than systems themselves. Platforms that claim they cannot moderate content at scale suddenly find themselves capable of blocking entire categories of speech when regulatory pressure arrives. That capability was always there. The question is only who decides when to use it.
Established labs are competing through distribution and positioning rather than capability announcements. Google is positioning scientific research as a partnership play built on open resources, signaling that influence in academia carries longer-term strategic value than proprietary model dominance alone. IBM is chasing immediate commercial application through consumer engagement via the Ferrari app and enterprise consulting to private equity firms. Neither involves breakthrough capability claims, underscoring a shift in how labs compete: not through raw capability but through distribution channels and positioning as trusted advisors in specific verticals. The real competition is over who owns the relationship when enterprises decide what to build with AI.
Grant Calloway
AI-assisted job-search tools have become increasingly popular by making it easier to find and apply to jobs. But by making it easier for applicants to generate and tailor application materials, they can also reduce how informative those materials are about applicant fit. We study this tradeoff in a hiring market where applicants differ in experience and latent match quality and firms use noisy application materials to decide whom to screen. We ask how AI affects downstream screening and hiring, and which applicants are most adversely affected. As application materials become less informative, a Bayesian firm rationally relies more heavily on coarse observables such as prior experience. Among the four applicant types defined by experience and compatibility for the job, inexperienced-compatible applicants are the most exposed: they lack observable experience and lose the individualized information that could distinguish them from other inexperienced candidates. When screening is costly, these changes can also generate inefficient screening failures in which firms screen no applicants or screen only experienced applicants. We then show that multistage hiring can arise as an endogenous firm response: a relatively inexpensive intermediate assessment allows firms to acquire new evidence of fit before costly full screening. This can restore screening opportunities that disappear under one-stage hiring and give inexperienced-compatible applicants a path to screening. Our results show how AI can shift the central friction in hiring from submitting applications to obtaining credible evaluation, creating entry barriers for high-fit workers without prior experience. Multistage hiring can endogenously arise in response, restoring evaluation opportunities that would otherwise disappear and helping preserve market functioning.
We introduce Multiplicatively Optimistic Regret Matching (MORM), an uncoupled learning rule for finite general-sum games. Under simultaneous full-information self-play, every player achieves external regret $O(\sqrt n\log d)$ uniformly over all horizons, using only one-step optimism. The analysis combines a potential-based regret-matching argument with multiplicative stability and Hellinger control of strategy movement. A learning-rate safeguard additionally gives $O(\sqrt{T\log d})$ regret in the face of adversarial utilities.
Computing Nash equilibria of simulation-based cybersecurity games with policy-space response oracles (PSRO) is bottlenecked by payoff estimation: every payoff-matrix entry costs Monte-Carlo rollouts of a slow simulator, while policies and restricted-game solves are cheap. We introduce Regret-Weighted Payoff Sampling (RWPS), a budgeted estimator that simulates only the cells an equilibrium is sensitive to and fills the rest with a surrogate trained on every entry simulated earlier in the run. The sup-norm error bound cannot evaluate such an estimator, because it is set by the cells left deliberately inaccurate. We prove an instance-dependent bound that weights error by the opponent's equilibrium mixture, a certificate computable from simulation data alone, and a coverage result showing that once the deviation-relevant set is simulated, surrogate error cannot affect either player's regret. On three 21x21 general-sum games, two synthetic and an asymmetric Colonel Blotto, the refined bounds are four to six times tighter on the estimator's own output, and the coverage result predicts in advance which games are cheap: 18% of the matrix for small-support games against 82% for Blotto. In growing-pool PSRO, RWPS reaches lower exploitability than minimum-regret-first search, information-gain search, and progressive sampling at a matched budget, and on the CyGym and ANSG cyber simulators it is lowest at the smallest budgets.
Long-running AI agents create a control problem: each action they take changes the state, which in turn affects the trajectory of future actions. If the agent is not fully aligned, then guaranteeing safety requires approving consequential actions before allowing them to be executed. But requiring human approval at every step makes attention a bottleneck. Delegating review to other AI agents raises the same alignment problem: the reviewers may themselves be misaligned. We identify a condition on a reviewing panel that is weaker than individual alignment yet necessary and sufficient for a guarantee that the principal fares at least as well in expectation as under a designated baseline policy. Each reviewer agent reports whether an action proposal made by a proposer agent improves its own utility relative to the baseline. We show that a threshold rule tolerating $k$ disapprovals is safe exactly when, after any $k$ reviewers are removed, the principal's utility can be written as a nonnegative combination of the remaining reviewers' utilities, plus a term that is nonnegative on every feasible proposal. We call this property $k$-robust coalitional alignment. The characterization lifts to sequential control: in a discounted MDP with an arbitrary proposer agent, safety at every state is both necessary and sufficient for the induced policy to match or improve on the baseline. When reviewers vote strategically, full-panel coverage in reward-function space guarantees that every Nash equilibrium is safe under the unanimous approval rule; in contrast, more permissive thresholds can admit unsafe equilibria even when reviewers are individually aligned. Experiments with existing reviewer models show that collective review can remain sound without an aligned individual, even when some disapprovals are tolerated.
In social dilemma games, additional rewards or punishments have been studied as means of promoting cooperation. Therefore, it is important to investigate the ideal situation, in which such an additional payoff would change the game. In this study, we investigated the symmetric solution of the Bellman optimality equation for a repeated harmony game. The calculations showed that three types of symmetric solutions exist. One of them corresponds to the trivial All-C strategy, and another to the Win-stay Lose-shift strategy of the prisoners dilemma game. The nontrivial behavior of the strategy corresponding to the last solution is also discussed in detail. In addition, we numerically investigated which strategy the agents actually learn by the reinforcement learning algorithm.
We give deterministic and uncoupled learning dynamics for finite multiplayer general-sum games under full-information feedback that achieve constant individual swap regret, independent of the horizon $T$. With $n$ players and at most $m$ actions each, the individual swap regret of every player is $O(\sqrt{n} m \log m \log^{5/2}(nm))$ at every finite horizon. Each player predicts the deviation gains, then uses these predictions to update a row-stochastic transition matrix, and plays its stationary distribution. The proof combines a potential argument exploiting stationarity with a two-scale higher-order prediction analysis, using rooted-tree representations to handle the nonlinear dependence of deviation gains on the stationary distributions. An adversarially robust variant, obtained through a generic common-prefix switching wrapper, preserves the self-play bound up to a universal constant and guarantees individual swap regret at most $7\sqrt{m T \log m}$ in the adversarial setting.
Composite score across coding, math, and reasoning
| # | Model | Score | tok/s | $/1M |
|---|---|---|---|---|
| 1 | GPT-5.5 | 60.2 | 73 | $11.25 |
| 2 | Claude Opus 4.7 | 57.3 | 51 | $10.00 |
| 3 | Gemini 3.1 Pro Preview | 57.2 | 132 | $4.50 |
| 4 | GPT-5.4 | 56.8 | 86 | $5.63 |
| 5 | Kimi K2.6 | 53.9 | 25 | $1.71 |
Agentic coding on real-world software engineering tasks
| # | Model | Score |
|---|---|---|
| 1 | Claude Opus 4.6 | 65.3% |
| 2 | gpt-5.2-2025-12-11-medium | 64.4% |
| 3 | GLM-5 | 62.8% |
| 4 | Junie | 62.8% |
| 5 | gpt-5.4-2026-03-05-medium | 62.8% |
TradingAgents: Multi-Agents LLM Financial Trading Framework
🕵️♂️ Collect a dossier on a person by username from 3000+ sites
Warp is an agentic development environment, born out of the terminal.
Coding Agent Harness
My personal directory of skills, straight from my .claude directory.
I'm going to build my own OpenClaw, with blackjack... and bun!
A general fine-tuning kit geared toward image/video/audio diffusion models.
A framework for efficient model inference with omni-modality models
Trinity-RFT is a general-purpose, flexible and scalable framework designed for reinforcement fine-tuning (RFT) of large language models (LLM).
Enterprise AI Platform with guardrails, MCP registry, gateway & orchestrator