Google is doubling down on video and real-time interaction as the frontier of generative AI deployment, launching both long-form coherent video generation and Gemini 3.8 Live with Live Avatar in the same cycle. The strategic signal is clear: text-to-text has commoditized, and the labs racing to lock in users are moving toward modalities that demand higher computational spend and stickier product integration. Meanwhile, GitHub is signaling that chat interfaces have hit a wall for certain workflows, introducing canvases as an alternative UI and deploying security-focused agents for fuzzing tasks. This suggests the industry is fragmenting around use cases rather than consolidating around a single interaction model. Infrastructure plays are consolidating around specific verticals: IBM is acquiring Logiq Consulting to bundle cybersecurity with digital transformation for regulated sectors, while simultaneously expanding banking infrastructure through Swift integration and tokenized deposit capabilities. NVIDIA's pandemic preparedness coalition announcement sits oddly alongside its GeForce NOW game launch, but the real move is positioning NVIDIA's compute as essential to both scientific research and entertainment delivery. Hugging Face's vision-language acceleration and Anthropic's economics project on agent trading suggest the next competitive frontier is not raw capability but efficiency and autonomous economic behavior. The labs are no longer racing to prove they can do X; they are racing to prove they can do X profitably at scale, with the right UI for the right customer, in the right vertical.
Sloane Duvall
A curated reference of models from major AI labs, with open/closed weight status, input modalities, and context window size. American labs tend towards closed weights models and Chinese labs tend toward open weights models.
None
None
None
None
None
None
None
None
None
None
None
None
None
None
None
None
None
None
None
None
None
None
None
None