• About Us
King of Computer Media
  • Home
  • Tech News
  • AI News
  • Apps & Tutorials
  • Mobile & Telecom
  • Lifestyle
  • About Us
No Result
View All Result
  • Home
  • Tech News
  • AI News
  • Apps & Tutorials
  • Mobile & Telecom
  • Lifestyle
  • About Us
No Result
View All Result
King of Computer Media
No Result
View All Result

Home - AI Trends and Related News - A US court ruled that AI companies do not need permission from original authors to use legally obtained books to train artificial intelligence.

A US court ruled that AI companies do not need permission from original authors to use legally obtained books to train artificial intelligence.

KOCPC Editor by KOCPC Editor
June 26, 2025 - Updated on August 4, 2026
in AI Trends and Related News, Latest Technology News

As generative AI rapidly rises, high-quality raw text data sources used for training AI large models have become increasingly important—yet they also involve considerable copyright disputes. Recently, the Northern District of California court ruled against an AI company… Anthropic The copyright lawsuit involved made a ruling that may differ from what the general public expects: confirming that legally purchased books can be used for AI training, falling under the “Fair Use” principle in U.S. copyright law.

A US court ruled that AI companies do not need permission from original authors to use legally obtained books to train artificial intelligence.

The case began in August 2024, with three American authors and journalists—Andrea Bartz, Charles Graeber, and Kirk Johnson—jointly filing a lawsuit. They accuse Anthropic of using databases from well-known piracy sites such as LibGen and Books3, and even directly scanning physical books into digital form, to train its AI language model Claude, constituting serious copyright infringement.

According to foreign media reports, Anthropic has acknowledged downloading data from millions of books for training purposes. Some of these sources were physical books purchased and scanned by the company itself, while a large portion also came from the aforementioned illegal online resources. However, Anthropic emphasized that its training practices serve a “transformative purpose” — that is, not merely copying, but creating entirely new works, which meets the legal requirements for “fair use.”

Three Key Court Rulings: Defining the Legal Boundaries of AI Training Data

This case was presided over by Judge William Alsup of the U.S. District Court for the Northern District of California, who issued detailed rulings in his judgment on three major points of contention, providing key guidance for the AI industry on the application of law.

1. Using scanned data from paper books for AI training constitutes “fair use”

Judge Alsop first pointed out that viewing AI training as “copying behavior” does not hold. The substantive act of AI model training is learning the statistical correlations between large volumes of text, rather than storing or reproducing specific content.

Regarding the plaintiff’s argument that “AI training may lead to the mass production of works in a similar style, affecting the market for original works,” the judge dismissed it. He cited a specific rationale: “This view is akin to arguing that teaching children to write would increase the number of competing writers and therefore harm authors’ interests. The purpose of copyright law is not to eliminate competition, but to promote creativity and the advancement of knowledge.”

The judge further noted that if the content generated by an AI model does not directly copy the original work, it does not constitute plagiarism or infringement.

Two, legally purchasing and scanning books for internal training also constitutes “fair use.”

Regarding Anthropic’s practice of purchasing books and then cropping, scanning, and digitally storing them for internal research, the court also determined that this conduct serves a “transformative” purpose and meets the conditions for fair use.

The judge emphasized that the company already owns the books, and the digitized books are used solely for internal AI training and research, without involving external distribution or sales. “This is an act of format conversion, not an infringement of copyright. The purpose lies in space management and search convenience, and it does not constitute a violation of the authors’ distribution rights or derivative work rights.”

This ruling is expected to serve as a precedent for many research institutions that are also undertaking internal data organization and digital transformation.

III. Using Pirated Data to Train AI Does Not Constitute “Fair Use”

The most controversial part comes from Anthropic’s admission that it used pirated sources such as Books3 and LibGen to collect over 7 million books as training material. Although the company claimed in litigation that this admission carried “malice,” it simultaneously emphasized that its purpose remained creative.

However, the court did not accept this. The judge explicitly stated that using illegally downloaded content to construct a central database is not “transformative” and is essentially a “substitute for paid books,” seriously violating copyright law.

Moreover, even if Anthropic later purchased some of the books, that does not offset the initial infringement. The judge noted with gravitas: “If academic research could serve as a free pass for illegal copying, the entire publishing market would cease to exist.”

The ruling also criticized Anthropic for failing to promptly delete illegally sourced data, indicating that its purpose of use had exceeded a reasonable scope.

What data can AI companies legally use? The court provides specific standards.

This ruling marks the first time that the legality of data use in AI model training has been given clear judicial definition. In summary, Judge Alsup’s decision establishes clear boundaries on the following points:

  • ✅ Legally purchased data used internally for transformative purposes constitutes fair use.

  • ❌ Content obtained from illegal websites, even if used for research, does not constitute fair use.

  • ✅ Digitization is a format conversion for legitimate purposes and does not constitute distribution or reproduction.

  • ❌ Storing and reusing pirated materials constitutes repeated infringement.

The Verge Quoting Anthropic’s official response to the ruling: “We are pleased that the court recognized the transformative nature of AI training. The purpose of training Claude models has never been to copy or replace original works, but to inspire creativity and advance scientific progress.” However, the portion involving pirated use will continue to be litigated, and the court will make further determinations on the amount of damages. While subsequent book purchases may help reduce legal liability, they are not enough to fully offset the earlier infringement.

Source: KOCPC Chinese

Tags: aiAnthropic

Recent Posts

  • The Xiaomi Pad 8S Pro has passed network access certification and will debut with the self-developed XRING O3 chip.
  • The entire Google Pixel 11 lineup has been leaked! Official promotional renders of the Pixel 11 Pro XL have also surfaced
  • Are Chinese phone battery capacities falsely labeled? A brief look at the “capacity locking” phenomenon in Chinese silicon-carbon batteries.
  • NCC is leaderless, recklessly sending out national-level alert messages!?
  • What does “QR” in QR Code mean?

Recent Comments

No comments to show.
  • About Us

We welcome partnership inquiries and product review opportunities from smartphone manufacturers, iPhone accessory brands, and app developers.koc kocpc.com.tw|Privacy Policy |Hosting & Maintenance: Fast Line Taiwan, A-Chang Digital Technology

No Result
View All Result
  • Home
  • Tech News
  • AI News
  • Apps & Tutorials
  • Mobile & Telecom
  • Lifestyle
  • About Us

We welcome partnership inquiries and product review opportunities from smartphone manufacturers, iPhone accessory brands, and app developers.koc kocpc.com.tw|Privacy Policy |Hosting & Maintenance: Fast Line Taiwan, A-Chang Digital Technology