← Back

Eli the Computer Guy

Google AI Training Models on User Data - Privacy Loses to Greed

Eli the Computer Guy15 min
Google privacyAI training datacorporate data riskbig tech ethicslocal AI deployment

Reacting to a TechCrunch report, Eli the Computer Guy walks through what he sees as an inevitable turn in the AI industry: after scraping the open internet and their own corporate archives, companies like Google have run out of data to feed their models and are now quietly rewriting privacy rules to harvest yours instead. He argues Google is especially dangerous because it does not feel like an AI company, it is the search bar, the document uploader, the tool people have trusted for twenty years, which makes it far easier for users and entire corporations to hand over sensitive information without registering what they are doing. The episode builds into a broader warning aimed at CIOs and CTOs, that trusting Silicon Valley's biggest players with an organization's proprietary data is a risk tolerance question, not a convenience question, and he names Sam Altman and Mark Zuckerberg directly as reasons why that trust is misplaced.

Why the AI Giants Ran Out of Data to Steal

Eli opens with the blunt premise that if you use Google, you are training its AI, and frames this as the logical next phase of an industry that has already scraped the entire public internet, mined internal corporate knowledge bases, and even monitored its own employees to find training material. The problem, he says, is that even after all that scraping, there simply is not enough data left to keep improving these models. That scarcity is what pushed Google, in his telling, to change the rules on ordinary users rather than admit the well has run dry.

He contrasts Google with more overtly AI-branded services like OpenAI or Anthropic, where users know they are interacting with an AI system and can calibrate their caution accordingly. Google's danger, he argues, is precisely that it does not feel like an AI company. People have relied on it for search, documents, and quick answers for up to two decades, so when AI functionality gets slotted into a familiar workflow, users barely notice the shift, let alone the new data policy attached to it.

"They have scraped the entire internet in order to get information to train the models, and what they're finding out is there is just not enough data."

The One PDF That Becomes a Metric Crap Ton of Leaked Data

To make the risk tangible, Eli builds a scenario around an employee staring at a confusing report and casually uploading a PDF to Google just to get a plain-language explanation of the numbers. On its own, he admits, one executive uploading one document does not matter. The danger appears only at scale, when a company with 150,000 employees each upload just one proprietary document per week, an assumption he calls conservative, and that volume compounds over months and years into a staggering amount of internal information sitting inside Google's AI systems.

He frames this as the classic technology-world mistake of reasoning about risk one-off rather than at scale, insisting that the habit of treating an AI query as a harmless shortcut, distinct from the official corporate-approved AI tool, is exactly how proprietary data quietly leaks out over a decade of accumulated small decisions.

"Imagine you have 150,000 employees that all upload to Google one proprietary document per week. How long before a metric crap ton of your proprietary information is now sitting in Google's AI LLMs?"

Reading the Fine Print: What Google's Policy Actually Says

Eli then walks through the specific language behind the change, noting that Google's updated privacy settings now allow it to store more user data, including images, files, audio, and video, explicitly to improve its AI models. He quotes Google's own email to customers stating that saved media, like search and service history, is used to develop and improve Google services and technologies, including AI models and safety measures. He stresses that this is not a guardrail or a safety subroutine sitting outside the model, it is a direct admission that user-submitted media becomes training material.

He also cites Google's help documentation, which describes using history to train generative AI models with the assistance of human reviewers, and points out that while some storage is temporary and tied to making a product function, Google's own language allows saved media to be retained specifically for AI training. For Eli, this confirms a familiar pattern: a useful, inexpensive tool earns trust, then slowly the company defaults users into new data-sharing arrangements they never explicitly agreed to.

"Your saved media is also used to develop and improve Google services and technologies, including AI models and safety measures."

The Zuckerberg and Altman Problem: Can You Trust the People Running This

This is where the episode turns sharply personal. Eli revives his long-running characterization of Mark Zuckerberg's approach to privacy, the idea that Zuckerberg will get what he wants from users regardless of their stated refusal, pointing to Facebook and Instagram's history of privacy overreach and the infamous Meta executive comment about the Rohingya massacre in Myanmar being an acceptable cost of doing business. He extends the same suspicion to Sam Altman, referencing Dario Amodei's description of him as difficult to work with, and Altman's sister's lawsuit against him, arguing that a person's personal conduct is a legitimate data point when deciding whether to hand him control of an organization's entire intelligence layer.

He frames Altman's ambition for OpenAI to function as a utility, one system running HR decisions, legal decisions, marketing decisions, even cafeteria menus, as a warning sign rather than a selling point. The core question he poses to listeners in IT leadership roles is whether they have the risk tolerance to trust what he calls the worst people on the planet with the systems their organizations depend on, given what those same people have already shown they are willing to do to users, employees, and even their own families.

"Can you trust that if that's what Sam Altman is willing to do to his own sister? Imagine what he's willing to do to your company."

The Case for Running AI In-House

Eli's practical conclusion is a push toward local, in-house AI deployment rather than reliance on big tech's cloud AI services. He points to Cisco rolling out internal AI agents to 90,000 employees as one model, and notes the growing popularity of open-source Chinese models specifically because they can be run internally without exposing data to an outside company's training pipeline. For him, this is not a fringe or paranoid position but a straightforward response to a pattern he has tracked for years, that big tech companies start out protecting privacy, then gradually encroach as the product becomes indispensable and switching costs rise.

He closes by inviting listeners to weigh the scale problem directly, imagining their own organization's headcount multiplied by a single weekly upload over a decade, and to decide for themselves whether that level of exposure is one they can tolerate.

"I think that's going to be one of the big pushes for deploying more and more local, more internal AI systems."

Key takeaways

  • Google has changed its privacy settings to allow storage of images, files, audio, and video specifically to train its AI models, per its own emailed language to customers.
  • AI companies have exhausted the public internet and internal corporate archives as data sources, driving them to loosen privacy rules to access ordinary user data.
  • Google's danger lies in its familiarity, users treat it as a search or query tool rather than an AI service, making casual uploads of sensitive documents easy to overlook.
  • Individually harmless actions like one employee uploading one PDF become a major data exposure risk at organizational scale, especially across large workforces over time.
  • Eli argues that Sam Altman and Mark Zuckerberg's documented personal and corporate conduct should factor into whether organizations trust them with core intelligence infrastructure.
  • Running AI systems in-house or using open-source models, as Cisco and several Chinese labs enable, is presented as the safer alternative to relying on big tech's cloud AI tools.

Resources mentioned

  • TechCrunch article on Google AI training on user data
  • Cisco internal AI agent deployment
  • Open-source Chinese AI models