The AI industry's data appetite is colliding with resistance at every layer, forcing a shift from extraction to negotiation. Where companies once simply scraped and trained, they now face font poisoning, opt-out friction, book-buying raids, and watermarking blowback. The real story is not whether these defenses work, but that they exist at all, and that companies are already accounting for them in their business models.
Amazon's decision to train on Twitch content by default, with an opt-out option, reveals the true calculus. Mike Minton's candid admission that opt-in would fail exposes the mechanics: the company knows users don't want this, but the revenue math assumes most won't bother to disable it. This is not a technical problem or a safety framework. It is a choice about whose interests get the default setting. Meanwhile, Anthropic's watermarking system is triggering the opposite complaint, users angry that their own use of Claude at work or school will be flagged. The company is solving for one problem and creating another: if watermarks work, they reveal Claude's presence in professional contexts where clients or employers might object. If they don't, they are theater.
Valuation momentum in AI coding and infrastructure companies tells a different story about where capital sees actual leverage. Cognition jumped from $26 billion to seeking $40 billion in months. Lovable hit $13.3 billion on $500 million ARR. Blacksmith's valuation jumped nearly tenfold in under a year. These are not companies training on scraped data or fighting watermark battles. They are companies selling tools to enterprises that need to deploy AI without drowning in the data and talent costs that make frontier models expensive. The frontier labs control the models. The infrastructure and tooling companies control the workflow. And the intermediaries, the platforms deciding which model gets called for each job, may control the revenue. That three-way split is where the real competition is now.
Sloane Duvall