OpenAI is doubling down on the regulatory safety narrative while quietly shipping product integrations that expand GPT-6 Astra's footprint into enterprise workflows. The Australian Youth Safety Blueprint reads as a preemptive policy document, the kind of thing that gets cited in congressional testimony and regulatory filings but rarely changes how products actually work. Meanwhile, the Hex partnership demonstrates where the real energy is: embedding AI agents into tools people use daily, where adoption happens through workflow convenience rather than persuasion. Google's MilleMiglia addresses a concrete logistics problem with no accompanying policy framework, which is how infrastructure tends to get built. Anthropic's move into embedded evaluation with Accenture suggests the real differentiation play isn't in model capability anymore but in making evaluation and governance legible to enterprises who need to justify AI spending to compliance teams. The GitHub podcast framing of RAG and MCP as settled questions worth debating signals that the infrastructure layer is consolidating around certain patterns, and the open question now is which abstractions win adoption. What's absent from this set is more revealing than what's present: no major claims about reasoning improvements, no benchmark leaderboard shuffling, no announcements about scaling laws or training efficiency breakthroughs. The labs are signaling that the competition has moved past model announcements into distribution, integration, and the boring work of making AI useful in ways that survive contact with actual business requirements.
Sloane Duvall
A curated reference of models from major AI labs, with open/closed weight status, input modalities, and context window size. American labs tend towards closed weights models and Chinese labs tend toward open weights models.
None
None
None
None
None
None
None
None
None
None
None
None
None
None
None
None
None
None
None
None
None
None
None
None