Skip Navigation

Two authors file a proposed class action lawsuit against Apple, alleging Apple knowingly used a dataset of pirated books to train its AI models

storage.courtlistener.com /recap/gov.uscourts.cand.455858/gov.uscourts.cand.455858.1.0_1.pdf
  1. Apple Intelligence is a set of generative AI programs and technologies designed and

maintained by Apple. 2. Apple—one of the world’s most valuable companies—has invested substantial capital and engineering resources into Apple Intelligence. It regards Apple Intelligence as a breakthrough innovation that will make its users’ experiences “profoundly different” across various product applications. Through Apple Intelligence, Apple hopes to add trillions to its market capitalization in coming years. 3. But Apple is building part of this new enterprise using Books3, a dataset of pirated copyrighted books that includes the published works of Plaintiffs and the Class. Apple used Books3 to train its OpenELM language models. Apple also likely trained its Foundation Language Models using this same pirated dataset. 4. Apple is building another part of its Apple Intelligence empire by using Applebot, a software program that copies mass quantities of webpages (also known as “scraping”). Apple scraped data with Applebot for nearly nine years before disclosing that it intended to train its AI systems on this scraped data. Scrapers like Applebot can also reach “shadow libraries” that host millions of other unlicensed copyrighted books, including, on information and belief, Plaintiffs’ and Class Members’ copyrighted works. 5. The Foundation Language Models within Apple Intelligence depend on the contents of their training datasets. The Foundation Language Models operate by copying and later simulating creative expression found in copyrighted works. For this reason, the inclusion of expressive high-quality material—especially copyrighted material—in Apple’s AI training datasets is deliberate and commercially significant. For instance, to access even more copyrighted material to develop its valuable generative AI products, Apple entered into a multimillion-dollar licensing agreement with Shutterstock. But not with Plaintiffs or the Class. 6. Plaintiffs and the Class are authors who have registered copyrights for their published works. They did not consent to the use of their works in any Apple Intelligence model, including the Foundation Intelligence Models and OpenELM language models. 7. The licensing market for AI training data is burgeoning. Nevertheless, Apple did not compensate creators for use of their copyrighted works and concealed the sources of their training datasets to evade legal scrutiny. On information and belief, Apple continues to retain a private AI training-data library including thousands of pirated books to train its future models, without seeking Plaintiffs’ or Class Members’ consent or providing them compensation. 8. In sum, Apple has copied the copyrighted works of Plaintiffs and the Class to train AI models whose outputs compete with and dilute the market for those very works—works without which Apple Intelligence would have far less commercial value. This conduct has deprived Plaintiffs and the Class of control over their work, undermined the economic value of their labor, and positioned Apple to achieve massive commercial success through unlawful means.

Comments

0