Apple Faces Copyright Claims Over AI Training Data Amid Open Source Disputes
On March 18, Apple was once again named as a defendant in a copyright infringement lawsuit filed by Chicken Soup for the Soul, LLC. The suit alleges that Apple utilized "The Pile" dataset—which contains pirated books—for artificial intelligence training. This extensive litigation also targets other global technology leaders, including Meta, xAI, Google, Anthropic, OpenAI, Perplexity, and NVIDIA. At the heart of the case is the "Books3" shadow library module within the dataset, which houses a vast collection of copyrighted literary works.

In response to the claims, Apple reiterated its commitment to developing AI datasets legally and ethically since 2024. While Apple researchers did use "The Pile" dataset in the open-source project OpenELMs, the company clarified that this was solely for public research and was not employed in its core Apple Intelligence system. Legal analysts, however, point out a potential complication: since Apple's foundational model received assistance from Google Gemini, Apple could face intricate joint liability if Google is found to have breached regulations, due to their technical supply chain relationship.
Currently, companies such as Perplexity have defended their web scraping practices, while Apple continues to emphasize the transparency and compliance of its own model training. As AI regulations become more stringent, this class-action lawsuit, which targets the foundational training data, represents a significant escalation in creators' pushback against what they see as "data exploitation" by tech giants. It is also expected to compel the industry to reassess the compliance costs and technical limits of implementing "data traceability" in model development.
Related article
DeepMind CEO Hassabis: I sleep six hours a day, usually feel energetic around 1 a.m.
Fortune recently featured an interview with Demis Hassabis, CEO of Google DeepMind, revealing his unconventional approach to rest and productivity. Hassabis disclosed that he sleeps very little, structuring his waking hours into two distinct work blo
OpenAI, Anthropic Vie for Market Share Despite Revenue Shortfalls
Despite recent reports suggesting OpenAI missed revenue targets, creating pressure on tech stocks this Tuesday, private AI lab investors remain resilient. Seasoned backers have confirmed they will not reduce investment despite negative media coverage
California AV Compliance: A New Era of Tickets, Geofences, and 1M Miles
Guident operates an AuveTech shuttle in South Florida, managing a four-mile route in West Palm Beach and a one-mile route in Boca Raton using its remote monitoring technology. | Credit: GuidentCalifornia is redefining the regulatory landscape for dri
Related Special Topic Recommendations
Comments (0)
0/500
On March 18, Apple was once again named as a defendant in a copyright infringement lawsuit filed by Chicken Soup for the Soul, LLC. The suit alleges that Apple utilized "The Pile" dataset—which contains pirated books—for artificial intelligence training. This extensive litigation also targets other global technology leaders, including Meta, xAI, Google, Anthropic, OpenAI, Perplexity, and NVIDIA. At the heart of the case is the "Books3" shadow library module within the dataset, which houses a vast collection of copyrighted literary works.

In response to the claims, Apple reiterated its commitment to developing AI datasets legally and ethically since 2024. While Apple researchers did use "The Pile" dataset in the open-source project OpenELMs, the company clarified that this was solely for public research and was not employed in its core Apple Intelligence system. Legal analysts, however, point out a potential complication: since Apple's foundational model received assistance from Google Gemini, Apple could face intricate joint liability if Google is found to have breached regulations, due to their technical supply chain relationship.
Currently, companies such as Perplexity have defended their web scraping practices, while Apple continues to emphasize the transparency and compliance of its own model training. As AI regulations become more stringent, this class-action lawsuit, which targets the foundational training data, represents a significant escalation in creators' pushback against what they see as "data exploitation" by tech giants. It is also expected to compel the industry to reassess the compliance costs and technical limits of implementing "data traceability" in model development.
DeepMind CEO Hassabis: I sleep six hours a day, usually feel energetic around 1 a.m.
Fortune recently featured an interview with Demis Hassabis, CEO of Google DeepMind, revealing his unconventional approach to rest and productivity. Hassabis disclosed that he sleeps very little, structuring his waking hours into two distinct work blo
OpenAI, Anthropic Vie for Market Share Despite Revenue Shortfalls
Despite recent reports suggesting OpenAI missed revenue targets, creating pressure on tech stocks this Tuesday, private AI lab investors remain resilient. Seasoned backers have confirmed they will not reduce investment despite negative media coverage





Home






