arrow_backNeural Digest
OpenAI and Microsoft logos amid legal documents
Policy

OpenAI & Microsoft Knew Web Scraping Was a 'Doom Loop'

The Verge AI10h ago
auto_awesomeAI Summary

Newly unsealed documents from the New York Times' lawsuit against OpenAI and Microsoft show the companies privately acknowledged their AI training methods could trigger a 'doom loop' damaging the broader web ecosystem. Internal documentation went further, describing their large-scale data scraping as the 'largest theft of labor in human history.' These revelations significantly undercut public narratives that both companies have maintained around responsible AI development.

Key Takeaways

  • Unsealed court documents from the NYT lawsuit show OpenAI and Microsoft internally flagged their scraping practices as a potential 'doom loop' for the web.
  • Internal company documentation described their AI training data collection as the 'largest theft of labor in human history.'
  • The disclosures directly contradict both companies' public messaging around ethical and responsible AI development.

Court documents reveal OpenAI and Microsoft internally warned their data practices would devastate the web.

trending_upWhy It Matters

These revelations could significantly strengthen the NYT's legal case and embolden other publishers and creators pursuing similar copyright and compensation claims against AI developers. If courts rule that large-scale web scraping constitutes unlawful data use, it could force a fundamental restructuring of how AI models are trained, raising costs and slowing development timelines across the industry. Regulators in the US and EU, already scrutinising AI data practices, may use these admissions as evidence that self-regulation has failed. Smaller AI startups that rely on similar scraping pipelines could face existential legal exposure if precedents are set against OpenAI.

FAQ

What is the 'doom loop' OpenAI and Microsoft warned about?

The term referred to an internal concern that scraping vast amounts of web content to train AI models would degrade the quality and sustainability of the open web over time. As AI-generated content proliferates, the web risks becoming a low-quality feedback loop of machine-generated data.

What is the New York Times lawsuit against OpenAI and Microsoft about?

The NYT sued OpenAI and Microsoft alleging that their AI models were trained on millions of copyrighted NYT articles without permission or compensation. The case is one of the most high-profile copyright challenges facing the AI industry.

Could these documents change the outcome of the lawsuit?

Internal admissions that the companies were aware their practices could be characterised as large-scale theft are potentially damaging to their legal defence. They may complicate arguments that scraping was done in good faith or constitutes fair use under copyright law.

This summary was AI-generated. Neural Digest is not liable for the accuracy of source content. Read the original →
Read full article on The Verge AIopen_in_new
Share this story

Related Articles