Mozilla Report: How Common Crawl’s Data Infrastructure Shaped the Battle Royale over Generative AI

Mozilla Foundation

3 min read Original article ↗

Mozilla investigates Common Crawl’s influence as a backbone for Large Language Models: its shortcomings, benefits, and implications for trustworthy AI



(BERLIN, GERMANY | FEBRUARY 6, 2024) —
When OpenAI rolled out its text generator ChatGPT in 2022, few paid attention to the outsized importance of its chief training dataset, Common Crawl.

Now, Mozilla’s new study “Training Data for the Price of a Sandwich: Common Crawl’s Impact on Generative AI” shows how Common Crawl laid the infrastructural foundation that shaped today’s burgeoning generative AI boom.

Common Crawl is a small nonprofit organization that is largely unknown to the broader public, but has become pivotal to the development of generative AI as the largest freely available source of web crawl data.

The study explores Common Crawl’s role in generative AI, the benefits and risks of its popularity among AI builders, and how AI builders have often used its data uncritically. The researcher highlights Common Crawl’s approach to tackling data quality, the limitations of its data, and what builders’ dependence on Common Crawl means for furthering trustworthy AI.

Common Crawl’s massive dataset is more than 9.5 petabytes large and makes up a significant portion of the training data for many Large Language Models (LLMs) like GPT-3, which powers the free version of ChatGPT. Over 80% of GPT-3 tokens (a representation unit of text data) stemmed from Common Crawl. Many models published by other developers likewise rely heavily on it: the study analyzed 47 LLMs published between 2019 and October 2023 that power text generators and found at least 64% of them were trained on Common Crawl. Its popularity among AI builders comes with positive and negative implications, Mozilla’s research argues. It improved the openness and transparency of LLM research and development, but it also led to many models being trained on biased and toxic data as well as on copyrighted materials.

Indeed, most recently, Common Crawl was cited as a key player in the copyright infringement case of the New York Times against OpenAI and Microsoft. The New York Times highlighted that its content made up a significant proportion of Common Crawl’s data at the time OpenAI launched ChatGPT, and therefore it very likely made up a significant proportion of GPT-3’s training data as well. This added to the growing list of copyright cases between content producers and generative AI companies.

For the report, Mozilla researchers conducted in-depth interviews with Common Crawl’s director and main crawl engineer and analyzed online documentation and discussions from the project.

Says Stefan Baack, Mozilla researcher and the report’s author: “Common Crawl has helped to make generative AI more transparent and audible, but it is a problematic source to train LLMs that needs to be used with care. Yet, this care is often lacking among AI builders. With our report, we highlight the consequences of using Common Crawl uncritically and show what both Common Crawl and AI builders can do to make generative AI more trustworthy.”