Why the AI Giants Ran Out of Data to Steal
Eli opens with the blunt premise that if you use Google, you are training its AI, and frames this as the logical next phase of an industry that has already scraped the entire public internet, mined internal corporate knowledge bases, and even monitored its own employees to find training material. The problem, he says, is that even after all that scraping, there simply is not enough data left to keep improving these models. That scarcity is what pushed Google, in his telling, to change the rules on ordinary users rather than admit the well has run dry.
He contrasts Google with more overtly AI-branded services like OpenAI or Anthropic, where users know they are interacting with an AI system and can calibrate their caution accordingly. Google's danger, he argues, is precisely that it does not feel like an AI company. People have relied on it for search, documents, and quick answers for up to two decades, so when AI functionality gets slotted into a familiar workflow, users barely notice the shift, let alone the new data policy attached to it.
"They have scraped the entire internet in order to get information to train the models, and what they're finding out is there is just not enough data."
The One PDF That Becomes a Metric Crap Ton of Leaked Data
To make the risk tangible, Eli builds a scenario around an employee staring at a confusing report and casually uploading a PDF to Google just to get a plain-language explanation of the numbers. On its own, he admits, one executive uploading one document does not matter. The danger appears only at scale, when a company with 150,000 employees each upload just one proprietary document per week, an assumption he calls conservative, and that volume compounds over months and years into a staggering amount of internal information sitting inside Google's AI systems.
He frames this as the classic technology-world mistake of reasoning about risk one-off rather than at scale, insisting that the habit of treating an AI query as a harmless shortcut, distinct from the official corporate-approved AI tool, is exactly how proprietary data quietly leaks out over a decade of accumulated small decisions.
"Imagine you have 150,000 employees that all upload to Google one proprietary document per week. How long before a metric crap ton of your proprietary information is now sitting in Google's AI LLMs?"
Reading the Fine Print: What Google's Policy Actually Says
Eli then walks through the specific language behind the change, noting that Google's updated privacy settings now allow it to store more user data, including images, files, audio, and video, explicitly to improve its AI models. He quotes Google's own email to customers stating that saved media, like search and service history, is used to develop and improve Google services and technologies, including AI models and safety measures. He stresses that this is not a guardrail or a safety subroutine sitting outside the model, it is a direct admission that user-submitted media becomes training material.
He also cites Google's help documentation, which describes using history to train generative AI models with the assistance of human reviewers, and points out that while some storage is temporary and tied to making a product function, Google's own language allows saved media to be retained specifically for AI training. For Eli, this confirms a familiar pattern: a useful, inexpensive tool earns trust, then slowly the company defaults users into new data-sharing arrangements they never explicitly agreed to.
"Your saved media is also used to develop and improve Google services and technologies, including AI models and safety measures."
The Zuckerberg and Altman Problem: Can You Trust the People Running This
This is where the episode turns sharply personal. Eli revives his long-running characterization of Mark Zuckerberg's approach to privacy, the idea that Zuckerberg will get what he wants from users regardless of their stated refusal, pointing to Facebook and Instagram's history of privacy overreach and the infamous Meta executive comment about the Rohingya massacre in Myanmar being an acceptable cost of doing business. He extends the same suspicion to Sam Altman, referencing Dario Amodei's description of him as difficult to work with, and Altman's sister's lawsuit against him, arguing that a person's personal conduct is a legitimate data point when deciding whether to hand him control of an organization's entire intelligence layer.
He frames Altman's ambition for OpenAI to function as a utility, one system running HR decisions, legal decisions, marketing decisions, even cafeteria menus, as a warning sign rather than a selling point. The core question he poses to listeners in IT leadership roles is whether they have the risk tolerance to trust what he calls the worst people on the planet with the systems their organizations depend on, given what those same people have already shown they are willing to do to users, employees, and even their own families.
"Can you trust that if that's what Sam Altman is willing to do to his own sister? Imagine what he's willing to do to your company."
The Case for Running AI In-House
Eli's practical conclusion is a push toward local, in-house AI deployment rather than reliance on big tech's cloud AI services. He points to Cisco rolling out internal AI agents to 90,000 employees as one model, and notes the growing popularity of open-source Chinese models specifically because they can be run internally without exposing data to an outside company's training pipeline. For him, this is not a fringe or paranoid position but a straightforward response to a pattern he has tracked for years, that big tech companies start out protecting privacy, then gradually encroach as the product becomes indispensable and switching costs rise.
He closes by inviting listeners to weigh the scale problem directly, imagining their own organization's headcount multiplied by a single weekly upload over a decade, and to decide for themselves whether that level of exposure is one they can tolerate.
"I think that's going to be one of the big pushes for deploying more and more local, more internal AI systems."
