AI Training Data Litigation: What the Google Publisher Suit and the Anthropic Settlement Mean for Your Company

Quick answer: Yes — training an AI model on lawfully acquired copyrighted books can qualify as fair use, but retaining pirated copies of that data, or stripping copyright management information from it, is a separate legal violation that fair use does not excuse. Companies that document data provenance and avoid retaining improperly sourced material significantly reduce their exposure.

On July 14, 2026, a coalition of publishers and authors — including Hachette, Cengage, Elsevier, and novelist Scott Turow — filed a class action lawsuit against Google, alleging that the company used their copyrighted works to train its Gemini models without permission. The complaint goes further: it alleges Google intentionally removed or altered copyright management information embedded in the works to obscure the fact that Gemini had been trained on the material.

The Google suit is the latest entry in a rapidly expanding body of AI copyright litigation, and it arrives just as courts are beginning to settle some of the foundational legal questions raised by generative AI. Chief among them: does training a model on copyrighted material infringe the underlying copyright, or does it qualify as fair use?

The most instructive answer so far comes from Bartz v. Anthropic, 787 F.Supp.3d 1007 (N.D. Cal. 2025). There, a federal court held that training an AI model on lawfully acquired copyrighted books is a transformative use protected by fair use. That was a significant win for AI developers. But the same court drew a hard line around how the training data was obtained and stored: Anthropic's practice of retaining pirated digital copies of books was not excused by the fair use finding. That distinction proved expensive. The case ultimately settled for $1.5 billion, with individual payouts estimated near $3,000 per infringed work.

UMG Recordings, Inc. v. Suno, Cause No. 1:24-CV-11611 (D. Mass. filed June 24, 2024) is a parallel case involving music recordings worth watching closely. Discovery in that case reportedly surfaced audio fingerprinting evidence showing millions of copyrighted recordings embedded in Suno's training data. A summary judgment ruling on whether AI music training without a license constitutes fair use — once expected in the summer of 2026 — has been pushed back repeatedly and is now scheduled for January 2027 before Chief Judge F. Dennis Saylor IV in the District of Massachusetts. The outcome will matter well beyond the music industry, since the ruling on the fair use issue will likely apply to text, image, and code models as well.

Taken together, these cases are converging on a workable, if still evolving, legal framework: the act of training a model on copyrighted works may be defensible as transformative fair use, but a company's data acquisition and retention practices are judged independently and can create liability even when the training itself would otherwise be protected. In practice, this means the sourcing pipeline of materials used to train LLM’s is now a distinct area of legal risk that requires its own diligence.

What This Means for Businesses Building or Deploying AI

For companies developing proprietary AI models, fine-tuning foundation models on internal or third-party datasets, or licensing AI tools from vendors, several practical steps follow from these cases. First, document the provenance of training data used internally or by vendors, including licensing terms and any representations about how datasets were assembled. Second, avoid retaining copies of copyrighted works obtained outside of a clear license, even temporarily, since retention itself may be an independent basis for liability. Third, when licensing AI tools from outside vendors, ask pointed questions about the vendor's training data sourcing and seek contractual protections, including indemnification, in the event the vendor is later found to have used improperly obtained data. Fourth, monitor AI-fair use cases closely as this is a rapidly evolving area of caselaw.

Wittliff Cutter Saba's technology and IP litigation team advises companies on AI-related intellectual property risk, from vendor diligence to litigation defense. If your business is building, fine-tuning, or licensing AI models, we can help you assess your exposure before it becomes a lawsuit. And if it does, we can be there to represent you in court.

Frequently Asked Questions

Is training an AI model on copyrighted books considered fair use?

Under Bartz v. Anthropic, 787 F. Supp. 3d 1007 (N.D. Cal. 2025), a federal court held that training an AI model on lawfully acquired copyrighted books can qualify as a transformative fair use. However, the same court held that retaining pirated copies of those books is a separate, unprotected violation — a distinction that cost Anthropic a $1.5 billion settlement.

What is copyright management information (CMI), and why does it matter in AI litigation?

CMI is the metadata attached to a copyrighted work — such as the author, title, and license terms — that identifies its ownership and use restrictions. The 2026 lawsuit against Google over Gemini alleges the company removed or altered this information to obscure that the model had been trained on the underlying works, a claim that is legally distinct from ordinary copyright infringement.

What should a company do to reduce legal exposure from AI training data?

At minimum: document the provenance and licensing terms of training data used internally or by vendors; avoid retaining copies of copyrighted works obtained outside a clear license, even temporarily; negotiate indemnification when licensing AI tools from outside vendors; and monitor developments in AI fair-use litigation, since the law in this area is still evolving.

Sources: Bartz v. Anthropic, 787 F. Supp. 3d 1007 (N.D. Cal. 2025); UMG Recordings, Inc. v. Suno, Cause No. 1:24-CV-11611 (D. Mass., filed June 24, 2024); Hachette, Cengage, Elsevier, et al. v. Google (N.D. Cal., filed July 14, 2026).
Visuable

Award-winning Top Squarespace Expert agency creating high-end Squarespace 7.1 websites with Fluid Engine. We deliver branding, copywriting, SEO and AIO strategies, membership, course, subscription, scheduling and ecommerce platforms (Squarespace and Shopify). Since 2015, we’ve helped 1,500+ brands across B2B, Food & Drink, Technology, Creative, Health & Wellness, personal brands and non-profits grow online. Meet us: visuable.co.

http://www.visuable.co
Next
Next

Wittliff Cutter saba Earns Chambers Honors for Commercial Litigation, Intellectual Property