The competitive dynamics in AI have shifted from regulation and safety to a far more immediate concern: preventing rivals from understanding how your customers actually use your tools. Google's narrative of frictionless human-AI collaboration through its Declaration of Independence campaign sits uneasily against Midjourney's legal demand that studios disclose their own AI usage patterns, a move that exposes the gap between how AI adoption is marketed and how it is practiced. Alibaba's classification of Claude Code as high-risk appears less a safety judgment than a market signal aligned with Beijing's preferences, a tactic that only functions if competitors are not deploying identical strategies elsewhere. Mistral's positioning as the open-source counterweight to OpenAI, backed by substantial capital, suggests the frontier model market is consolidating around a handful of players with distinct distribution strategies and geographic bets. The underlying story is one of companies racing to lock in users and developers while simultaneously erecting barriers to competitive intelligence.
This tension between opacity and consolidation finds a methodological parallel in how researchers are approaching statistical inference itself. Archived papers in causal inference, computational estimation, and synthetic data validation all grapple with a shared problem: preserving identification guarantees when standard assumptions break down in practice. Knowledge Cascade transfers hyperparameters from cheap models to expensive ones via scaling laws; HERO calibrates noisy crowdsourced labels using historical gold annotations; task exchangeability provides guarantees for synthetic data in scientific studies. The pattern across these papers favours transparent assumptions, closed-form solutions where possible, and explicit characterization of when methods diverge from ideal benchmarks, prioritizing interpretability over parameter count. The preference is revealing: when stakes are high and data is imperfect, researchers choose interpretability.
Meanwhile, the practical engineering effort is concentrating not on models themselves but on the infrastructure that connects them to the world. Chrome DevTools MCP, Unity MCP, and page-agent solve a genuine problem: giving language models actionable control over systems they previously could not touch. GitHub's trending set shows developers building agent multiplexers, skill repositories, and self-hosted infrastructure for privacy-sensitive workloads, with Chrome and game engine integrations gaining traction while viral token-optimization jokes remain commentary rather than solutions. The real work is happening in the connectors, not the connectors' users.
Grant Calloway
No lab headlines.
The nonparametric maximum likelihood estimator (NPMLE) of a Gaussian location mixture maximizes the likelihood over the infinite-dimensional space of mixing distributions. The maximizing mixing distribution can be nonunique, and the classical bound on its number of atoms grows linearly with the sample size $n$. We show that a vanishingly small random perturbation of the likelihood yields exact polylogarithmic sparsity. The resulting randomly reweighted NPMLE maximizes a weighted likelihood whose independent weights, taken to be Gamma in our analysis, concentrate around one as $n$ grows. With high probability, it is unique, has $O\{(\log n/\log\log n)^d+\log n\}$ atoms in dimension $d$, nearly maximizes the ordinary likelihood, and estimates the mixture density at a Hellinger rate that is parametric up to logarithmic factors. This sparsity holds for the estimator itself, not for an approximation of it, and requires no support penalty. The proof rests on an effective-dimension principle for positive kernel mixtures: low-dimensional variation of the fitted values controls the support of every extreme point of the set of maximizers. Numerical illustrations verify that the reweighted NPMLE has Hellinger risk and support size comparable to those of the ordinary NPMLE.
We consider estimating the conditional distribution of a multivariate outcome given covariates when its coordinates may be continuous, binary, categorical, ordinal or rankings, and are conditionally dependent on one another. Different statistical methods have been developed for each outcome type, and most of them target a summary of the conditional distribution, such as the mean of each coordinate, rather than the joint distribution of the outcome vector. We develop generalized engression models, a unified nonparametric distributional regression framework for outcomes of any type. The proposed method builds upon engression, a scoring-rule-based deep generative model, and introduces a data-type-specific link function and a stochastic perturbation that smooths the loss, enabling gradient-based training even with discontinuous links. We establish universal representation results for continuous, discrete and mixed outcomes. In simulations and in two applications, 242 species in a community ecology benchmark and a 17-dimensional mixed-type health outcome, the method matches type-specific models on marginal scores, improves on them on the joint distribution, and matches or exceeds purpose-built state-of-the-art joint species distribution models. Software is available in Python.
We develop a statistically explicit sentiment index for Google Play user reviews and establish the mathematical results supporting its construction. Normalized star ratings and text-sentiment scores are treated as noisy measures of latent review valence and fused by covariance-aware inverse-variance weighting. Review-level estimates are aggregated with bounded helpfulness and recency weights, then shrunk toward a population mean using estimated precision rather than an arbitrary review-count threshold. App-level rating histograms provide a distributional diagnostic for samples returned under different API sort orders; because star ratings are discrete, classical continuous Kolmogorov-Smirnov critical values are not used. A local-level state-space model and the Kalman filter provide a denoised temporal trend. Full proofs cover the BLUE and Gaussian maximum-likelihood result, Gaussian-conjugate shrinkage, the Glivenko-Cantelli and Donsker theorems, count transformations via the delta method, and exact Gaussian Kalman filtering. A worked three-review example shows how textual complaints can materially reduce an apparently perfect star-only score.
We study false-alert control when screening for text generated by artificial intelligence (AI). The screening procedure selects document prefixes and detectors from observed evidence and may stop before exhausting its inspection budget. We give two finite-sample constructions under document-level exchangeability between human calibration documents and a new null document, with no restriction on dependence among tokens within a document. Construction A registers a finite family of prefix-detector scores and allocates a false-alert budget across their conformal ranks. A union bound protects any executed subset of that family. Construction B calibrates the complete-path maximum of a development-fixed adaptive policy. Each partial-path maximum is bounded by the complete maximum, so a terminal conformal rank protects early stopping without splitting the error budget. We prove marginal control of any false alert across the permitted inspection path and derive necessary calibration counts for rejection. We also state oracle testing, distribution-shift, and independent-audit bounds with their additional assumptions. Both constructions protect stopping within their specified scope; neither proof constructs an e-process or justifies multiplying conformal ranks. Detection power and computational savings remain questions for empirical evaluation.
Generative AI systems are increasingly used, but aligning their outputs with user requirements poses a continuing challenge. Here, we aim to ensure that the distribution of an attribute of an AI-generated output aligns with a user-specified target. This is motivated by examples such as fairness, where we want to ensure that a protected attribute (e.g., gender, race, or age categories) follows a desired distribution, and synthetic data generation, where we want the generated data to be representative of a target distribution. We study the practically important black-box access setting, where a user can repeatedly query a generative AI model. The goal is to return $m\ge 1$ outputs whose joint attribute distribution is as close as possible to this target. For both exact and approximate alignment, we develop algorithms that minimize the expected number of queries to the generator, and we further demonstrate their optimality as the number of requested outputs $m \rightarrow \infty$. Experiments on text-to-image generation and geocoded persona generation tasks show that our post-processing algorithms improve statistical attribute alignment, complementing prompting-based interventions.
We propose Sufficiently Reduced Distributional Regression (SRDR), a generative method that combines conditional distribution estimation with nonlinear sufficient dimension reduction (SDR). It builds on a characterization of sufficiency through strictly proper scoring rules: a dimension reduction is sufficient if and only if predicting the response from the reduced covariates incurs no loss in expected score relative to the full covariates. Sufficient dimension reduction thus becomes a risk minimization problem. SRDR jointly trains a dimension reduction map and a generative prediction model by minimizing the energy score, which can be estimated by sampling without density evaluation or adversarial training. The framework extends to multi-environment data and to classification. We prove that the estimated conditional distributions converge in energy distance to the true ones, which implies that the learned representation is asymptotically sufficient. In simulations and applications to CT slice localization, superconductivity, and digit classification, SRDR recovers low-dimensional sufficient structure and matches or outperforms state-of-the-art nonlinear SDR methods in representation quality and predictive performance.
Composite score across coding, math, and reasoning
| # | Model | Score | tok/s | $/1M |
|---|---|---|---|---|
| 1 | Claude Fable 5 | 59.9 | 62 | $20.00 |
| 2 | Claude Opus 4.8 | 55.7 | 57 | $10.00 |
| 3 | GPT-5.5 | 54.8 | 92 | $11.25 |
| 4 | Claude Opus 4.7 | 53.5 | 47 | $10.00 |
| 5 | Claude Sonnet 5 | 53.4 | 79 | $6.00 |
Agentic coding on real-world software engineering tasks
| # | Model | Score |
|---|---|---|
| 1 | OpenAIgpt-5.5-2026-04-23-xhighModel | 62.7%± 0.91% |
| 2 | JunieJunieAgent | 61.6%± 0.64% |
| 3 | OpenAICodexAgent | 60.4%± 1.37% |
| 4 | AnthropicClaude CodeAgent | 59.6%± 1.98% |
| 5 | OpenAIgpt-5.5-2026-04-23-mediumModel | 58.9%± 0.78% |
Use Codex from Claude Code to review code or delegate tasks.
🪨 why use many token when few token do trick — Claude Code skill that cuts 65% of tokens by talking like caveman
JavaScript in-page GUI agent. Control web interfaces with natural language.
Open-source AI hackers to find and fix your app’s vulnerabilities.
Chrome DevTools for coding agents
Windows tmux alternative for AI agents — split terminals for Claude Code, Codex, Gemini CLI with MCP browser automation. No WSL required.
TongFlow : An Open-Source Multi-Modal GenAI Workflow Studio
Peekaboo is a macOS CLI & optional MCP server that enables AI agents to capture screenshots of applications, or the entire system, with optional visual question answering through local or remote AI models.
MegaDetector is an AI model that helps conservation folks spend less time doing boring things with camera trap images.
📑 PageIndex: Document Index for Vectorless, Reasoning-based RAG