Latest Trending Discover Timelines Categories
All explainers

Technology explainer

When Can AI Training Use Copyrighted Material Legally?

Whether AI training may use copyrighted works depends on the country, lawful access, the copies made, licensing terms, exceptions such as fair use or text-and-data mining, and effects on markets for the works. One judgment or settlement rarely supplies a universal answer.

Quick summary

AI training normally requires collecting, processing and often temporarily or persistently copying large amounts of material. Copyright law regulates particular acts in particular jurisdictions. The central questions are what was copied, how it was obtained, whether permission or a statutory exception applies, and whether the use substitutes for protected markets.

Separate the legal questions

Training data, model weights and generated outputs are different objects. A court may find that acquiring a dataset was unlawful while reaching another conclusion about analysis of lawfully acquired copies. An output that reproduces protected expression may raise infringement questions even if the training process itself qualifies for an exception.

Four routes that may authorize use

  • Public domain: copyright has expired or never applied to the material.
  • Permission or licence: the rightsholder authorizes specified uses under agreed terms.
  • Statutory exception: a jurisdiction may permit fair use, research, quotation, temporary copying or text-and-data mining under conditions.
  • Unprotected elements: facts, ideas and methods are generally treated differently from protected expression, although extracting them may still involve copying a work.

Why lawful access matters

Material visible online is not automatically free of copyright. Access may be constrained by subscription terms, technical controls or database rights. Copies obtained from pirate repositories create legal issues distinct from copies purchased, licensed or made available for mining. Provenance records help show where each dataset came from and under what terms.

How exceptions are evaluated

In the United States, fair use is a multi-factor, case-specific analysis that can consider purpose, transformation, the nature and amount of the work and market effects. Other countries use narrower enumerated exceptions, including text-and-data-mining rules with eligibility, access or opt-out conditions. The same training activity can therefore have different answers across borders.

Outputs and market harm

A model that learns statistical relationships is not necessarily storing a browsable library, but models can sometimes reproduce memorized passages or imitate commercially significant material. Courts may examine whether outputs are substantially similar to protected expression and whether the system competes with licensing or sales markets. Technical safeguards can reduce memorization but do not settle the legal test.

Reality check

A settlement binds the parties and may avoid a definitive ruling. A trial-court decision can be limited to its facts and may be appealed. “The court said AI training is legal” is usually too broad unless the jurisdiction, dataset, acquisition method, claim and procedural stage are specified.

What responsible operators document

Useful records include dataset sources, licences, opt-outs, filtering, retention, model purpose and tests for regurgitation. Legal review should cover every jurisdiction and product use involved. This explainer describes the analytical framework, not legal advice for a particular model or dispute.

First appeared in

Anthropic Will Pay $1.5 Billion for Pirated Books. The Bigger AI Question Is Still Open

A new version of NewTqnia is ready.