Anthropic's Dario Amodei dining with Trump signals a direct play for political favor at the exact moment the company faces real competitive pressure from Meta's recent announcements and corporate America's visible pivot toward cheaper open-source models. The day's story is not about AI capability or safety but about who controls access to it and what that means for power. Microsoft's release of the .NET SDK for the AG-UI agent protocol, built with CopilotKit and published under MIT license, represents a different kind of consolidation: standardizing how agents talk to applications locks in architectural choices that benefit whoever moves first. ASML's Christophe Fouquet holds a monopoly on advanced chipmaking equipment that no amount of open-source software can replicate, which means the bottleneck for AI scaling is not code but physics and geopolitics. The Financial Times reports that US businesses are already adopting Chinese alternatives to OpenAI and Anthropic, which means the narrative of American AI dominance is fraying while executives are still talking to regulators about safety frameworks.
AWS is playing infrastructure middleman to frontier models while quietly shipping the plumbing that makes those models useful at scale. OpenAI's GPT-6 Astra landing on Amazon Bedrock, paired with Amazon Quick desktop going general availability, follows a consistent pattern: AWS doesn't need to build the frontier model itself but needs to be the layer where frontier models become accessible, deployable, and embedded in workflows that drive compute consumption. Lower friction for developers to build on AWS infrastructure means more workloads, more data residency, more lock-in. The signal from AWS is not about winning on model capability; it's about owning the distribution layer and making sure builders have no reason to leave once they start.
Developer priorities have split decisively between agent infrastructure and local-first alternatives to cloud services. Paperclip dominates not because it's novel but because it solves a concrete problem at scale: managing multiple agents in production workflows. Hindsight, openrig, and shodh-memory all attack the same bottleneck that agent systems now leak money and latency through stateless designs. The second pattern is more ideological. VoiceStudio, PipePipe, hwdsl2's self-hosted stack, and genoffice share a common thesis: users should own their compute. These repositories aren't trying to be better than their commercial equivalents; they're trying to be free and local. The traction here signals that developers are tired of dependency chains that end in someone else's servers, with Univer and dream-num's approach suggesting the next frontier is bundling these local-first tools into coherent application suites rather than leaving users to stitch together single-purpose binaries.
Grant Calloway
Reasoning models often generate very long reasoning traces, making inference computationally expensive. Existing approaches typically improve efficiency either through inference-time early-stopping mechanisms or by explicitly encouraging shorter reasoning during training, for example through reinforcement learning with length penalties. We show that substantial efficiency gains can instead emerge from a different kind of supervision: \textit{confidence}. Using a self-supervised procedure, we fine-tune reasoning models to predict their confidence in the answer at intermediate points along their own reasoning trajectories using only 600 training problems. Confidence is used only as a training target: the loss contains no objective for reasoning length, efficiency, or stopping. At inference, the fine-tuned models use the standard generation procedure, with no confidence elicitation or early-stopping mechanism. Despite this, self-supervised confidence fine-tuning makes reasoning more efficient, reducing generated tokens by up to 25\% at matched accuracy across Gemma, Qwen, Nemotron, and GPT-OSS models on mathematical, scientific, and coding reasoning benchmarks, with efficiency gains comparable to methods that explicitly optimize for shorter reasoning. Analysis of reasoning episodes further shows that confidence supervision largely preserves the base models' high-level reasoning composition rather than selectively suppressing particular behaviors. Our results suggest that efficient reasoning may emerge as a downstream consequence of learning metacognitive signals, without being directly optimized.
We give a gap-free differentially private algorithm for the principal component analysis (PCA) problem with Gaussian data.
Recent literature has shown a strong connection between optimization and sampling. We develop the corresponding first-order theory for diffusion models. First, the SDE-based reverse-time flows of overdamped and underdamped Langevin diffusions contract relative Fisher divergences at explicit exponential rates whenever the stationary potential of the forward process is strongly convex---a condition on the noising process one chooses, not on the data. This is a unique advantage of SDE-based reverse diffusion, absent in the reverse process based on ODEs. Second, we incorporate discretization and establish averaged first-order stationarity bounds---the sampling analog of averaged gradient-norm guarantees in nonconvex optimization---for samplers of both overdamped and underdamped diffusion models. As in nonconvex optimization, the convexity-free certificate is local: it guarantees score consistency, not global mode weights.
Generative AI systems are increasingly used, but aligning their outputs with user requirements poses a continuing challenge. Here, we aim to ensure that the distribution of an attribute of an AI-generated output aligns with a user-specified target. This is motivated by examples such as fairness, where we want to ensure that a protected attribute (e.g., gender, race, or age categories) follows a desired distribution, and synthetic data generation, where we want the generated data to be representative of a target distribution. We study the practically important black-box access setting, where a user can repeatedly query a generative AI model. The goal is to return $m\ge 1$ outputs whose joint attribute distribution is as close as possible to this target. For both exact and approximate alignment, we develop algorithms that minimize the expected number of queries to the generator, and we further demonstrate their optimality as the number of requested outputs $m \rightarrow \infty$. Experiments on text-to-image generation and geocoded persona generation tasks show that our post-processing algorithms improve statistical attribute alignment, complementing prompting-based interventions.
Large language models (LLMs) implicitly infer attributes of their users and adapt their behavior accordingly, yet these beliefs remain difficult to inspect and causally manipulate. We introduce Belief Self-Distillation (BSD), a unified read-write framework that bridges linear and causal probing by learning a compact user representation that can be both decoded and written back into the model. The frozen LLM acts as its own teacher, distilling beliefs from natural conversations without external annotations. Unlike conventional probing, BSD isolates not only information present in activations, but a state whose causal role can be directly tested. Across multiple model families, BSD faithfully recovers user beliefs and enables substantially stronger interventions than matched hidden-state steering. Crucially, we find that refusal depends not only on the request, but on the model's inferred user intent: changing this belief alters refusal while holding the request fixed. We further uncover a striking cross-model regularity: independently trained LLMs converge on a shared geometry for representing their users. Together, these results reveal implicit user models as readable and causally writable internal states with direct implications for AI safety, shaping how models condition safety decisions on whom they believe they are interacting with.
Low-rank adapters (LoRA) make it cheap to fine-tune a large language model once per task, but combining several independently trained adapters into one model remains difficult: merging the updates in weight space causes interference, retraining on all task data is expensive, and routing between separate adapters gives up the goal of a single combined model. We trace the difficulty to two choices that every composition method makes implicitly. A LoRA update admits infinitely many equivalent factorizations; the choice among them is invisible while an adapter serves alone, but it determines what a learned interaction between adapters can see. A coupling between an old skill and a new one can likewise point in either direction, and the direction decides whether the old skills keep computing what they computed before. We introduce READ (Read-only Expansion of Adapter Deltas), which fixes both choices: each adapter is rewritten into a balanced canonical form that preserves its update exactly, and the coupling grows in one direction only, so a new skill can read the input subspaces of old skills but cannot write into their output subspaces. The only trainable object at each append is the new skill's row of the coupling matrix, and the composed update folds into the base weights with no inference cost, routing, or task-specific rules. We evaluate READ across four benchmark suites and two model families, adding skills one at a time. Across several families, READ improves every suite average over the strongest published baselines built from the same adapters---by more than twenty points on SuperGLUE and more than seven points on the domain suite---and nearly all complete addition sequences end above every direct baseline. Factor coordinates and coupling direction, which a lone adapter never exposes, are what decide whether composed skills survive.
Composite score across coding, math, and reasoning
| # | Model | Score | tok/s | $/1M |
|---|---|---|---|---|
| 1 | Claude Opus 5.5 | 57.6 | 95 | $8.00 |
| 2 | Claude Fable 5.1 | 53.4 | 67 | $20.00 |
| 3 | GPT-6 Astra | 52.7 | 63 | $20.00 |
| 4 | Claude Opus 5 | 50.8 | 0 | $10.00 |
| 5 | Claude Fable 5 | 49.6 | 0 | $20.00 |
Agentic coding on real-world software engineering tasks
| # | Model | Score |
|---|---|---|
| 1 | AnthropicFable 5 [high]Model | 64.5%± 1.41% |
| 2 | GrokGrok 4.5 [high]Model | 63.8%± 0.60% |
| 3 | AnthropicOpus 5 [high]Model | 63.4%± 1.35% |
| 4 | Z.aiGLM-5.2 [high]Model | 62.9%± 1.19% |
| 5 | OpenAIGPT-5.6 Sol [medium]Model | 62.3%± 1.83% |
The open-source app everyone uses to manage agents at work
Hindsight: Agent Memory That Learns
VoiceStudio is the open-source, fully-local ElevenLabs alternative — voice cloning, voice design, video dubbing, dictation, transcription & audiobook creation in 646 languages.
Learn it. Build it. Ship it for others.
An open-source Android app to let you browse YouTube and other services freely.
RL training framework for diffusion and omni-modality models
📚 Collection of token-level model compression resources.
Cognitive brain for Claude, AI agents & edge devices — learns with use, runs offline, single binary. Neuroscience-grounded 3-tier architecture with Hebbian learning.
Label Studio is a multi-type data labeling and annotation tool with standardized output format
Additional packages (components, document stores and the likes) to extend the capabilities of Haystack