
The internet is starting to train on itself, and that’s a bigger problem than it sounds.

In 2024, researchers writing in Nature showed what happens when AI models are trained over and over on their own output. The edges go first: the uncommon cases, the specific facts, the exceptions that don’t fit the average. The answers still sound fluent, but the model’s picture of reality gets narrower, and less reliable, with each generation. The researchers called this “model collapse.”
I’ll call it what it really is: data inbreeding. Keep a bloodline closed and the offspring still look like the original family. The resemblance holds, but the vitality doesn’t, and over subsequent generations things can get really ugly, pun intended.
Now imagine those offspring being copied into other families. That’s essentially what happens when a model feeds on its own output and that output gets scraped into the next model’s training data.
Synthetic data isn’t the real problem; replacement is. Follow-up research found that collapse hits hardest when generated data replaces real data. When real data stays in the mix, the damage stays limited. There’s a human floor, and as long as it holds, things stay healthy. That’s the real operating problem: running AI on secondhand signals isn’t a strategy.
The fresh web is already half machine-made. By late 2024, newly published articles were split roughly 50/50 between AI and human authors, and that ratio has held near that level into 2026. That isn’t the whole internet, but it’s the layer crawlers scrape next. Much of the web is also machine-translated at scale. Verified human writing has quietly become a scarce good, and it’s starting to be priced like one. Labs that keep a human floor and add synthetic data on top stay sharper. Everyone training on an unlabeled commons is basically guessing.

What does this mean for your company? Every team feels it differently. Sales needs to know what prospects actually want to buy. Marketing needs messaging that fills the pipeline and still holds up in search and generative engines, both of which penalize generic, fluent-sounding content. Product needs to know what to build and what to kill. Customer experience and account teams need to understand what customers think, feel, want, and are likely to do next.

Scraped pages, intent feeds, and an LLM’s built-in assumptions are all inference. Inference is useful, but it isn’t highly trustworthy data directly from the horse’s mouth. The data you can’t replace is what customers and the market tell you directly. That means collecting periodic, but continual, feedback from customers and the market (make them short pulses rather than punishing surveys). Then put the answers next to your CRM activity, campaigns, support tickets, product usage, and win/loss data. Research that stays a side project stays a side project. Inside an operating system, it becomes how the company sees.
The remedy is what breeders call outcrossing, and for a company, that means research. The failure mode is substitution. A model that only sees inferred or generated stand-ins will stay fluent on the average case and thin on the exception that happens to be true: the buyer who doesn’t match the persona, the objection that isn’t on the battle plan, the “satisfied” account that’s quietly shopping around. Those exceptions are exactly where deals are won and lost. So first-party research isn’t an annual voice-of-customer ritual, again, it’s a continual pulse, new fresh stock in the bloodline, so to speak. It belongs in your CRM, your analytics, and your models, as labeled, dated signals rather than an afterthought in a prompt.
This is why Reaction is built the way it is. We started as a research company. Teams used the platform to run customer feedback and market research they could actually field and act on. Analytics, marketing automation, and a full CRM came later, because insight goes to waste if it can’t live alongside the core work that’s being done in a company. That’s what a growth operating system is: research built into the core, not bolted onto a CRM as a clunky afterthought. If you put AI on top of that system, give it the campaign, the call, the ticket, and what real people said when you asked.
Keep the bloodline open because you’re not really short on compute, what you’re actually short on (and where real risk lives) is losing touch with the people who buy.
Curious what this looks like in practice? Schedule a demo and we’ll show you.
Subscribe for new articles straight to your inbox
Reaction helps organizations unify audience intelligence, coordinate engagement, and simplify growth operations across the customer lifecycle.