In a recent copyright lawsuit brought by authors against Meta, it was revealed that Meta was caught using torrents to scrape unauthorized resources from LibGen, a shadow library considered to have serious piracy issues. Continue reading the report on Meta’s legal troubles stemming from using pirated content from torrent sites to train its AI.

▲Image source: Meta
Meta faces legal trouble over using pirated content from torrent sites to train its AI.
Even though they already possess the vast amounts of “data” generated daily by users on their massive social platforms, including Facebook, Instagram, and Threads, it’s clear that Meta, like many other companies developing similar technology, still needs various more in-depth “learning resources” to improve AI performance.
Previously, Meta had mentioned that it used the Books3 dataset to train its Llama large language model. However, in a recent copyright lawsuit filed by authors against Meta, we learned that Meta was caught using torrents to download unauthorized resources from LibGen, a shadow library considered to have serious piracy issues.
Even foreign media also disclosed that the lawsuit mentions CEO Mark Zuckerberg, who appears to have authorized the use of these copyright-disputed materials through such means, according to related notes.

This lawsuit, filed in 2023 against Meta by writers including novelists Richard Kadrey, Christopher Golden, and Sarah Silverman, and known as “Kadrey et al. v. Meta Platforms,” has recently had many related details revealed by foreign media.
This includes engineers mentioning they suspected that “torrenting from a [Meta-owned] corporate laptop doesn’t feel right.” There is also information that “MZ” (believed to be an abbreviation for CEO Mark Zuckerberg) favored authorizing such behavior.

The disclosures also include Meta’s attempts to refute the claims, including the argument that such content counts as public content and therefore qualifies under “public availability.” They also mentioned that they only “use text to statistically model language and generate original expression,” so it constitutes a reasonable use of these resources.
At the same time, however, it was seen by foreign media and even the judge as somewhat like “protesting too much,” demanding excessive confidentiality over litigation-related materials, which drew a warning about “unreasonably broad sealing requests.” In a notably harsh tone, the judge called the company’s cover-up “preposterous” and ruled that the original documents should be made public.
I have to say, although this lawsuit initially seemed like Meta could safely get through because the plaintiff hadn’t provided sufficient evidence, given the various circumstances revealed so far, it looks like Meta needs to be more on guard, doesn’t it?
Citation source:WiredVia:9to5Mac|
Further Reading:
Lightweight version of NotebookLM? Gemini Live AI’s feature supporting uploads of YouTube or files to start conversations has been unearthed.
Source: KOCPC Chinese