Quick summary
AI training normally requires collecting, processing and often temporarily or persistently copying large amounts of material. Copyright law regulates particular acts in particular jurisdictions. The central questions are what was copied, how it was obtained, whether permission or a statutory exception applies, and whether the use substitutes for protected markets.
Separate the legal questions
Training data, model weights and generated outputs are different objects. A court may find that acquiring a dataset was unlawful while reaching another conclusion about analysis of lawfully acquired copies. An output that reproduces protected expression may raise infringement questions even if the training process itself qualifies for an exception.
Four routes that may authorize use
- Public domain: copyright has expired or never applied to the material.
- Permission or licence: the rightsholder authorizes specified uses under agreed terms.
- Statutory exception: a jurisdiction may permit fair use, research, quotation, temporary copying or text-and-data mining under conditions.
- Unprotected elements: facts, ideas and methods are generally treated differently from protected expression, although extracting them may still involve copying a work.
Why lawful access matters
Material visible online is not automatically free of copyright. Access may be constrained by subscription terms, technical controls or database rights. Copies obtained from pirate repositories create legal issues distinct from copies purchased, licensed or made available for mining. Provenance records help show where each dataset came from and under what terms.
How exceptions are evaluated
In the United States, fair use is a multi-factor, case-specific analysis that can consider purpose, transformation, the nature and amount of the work and market effects. Other countries use narrower enumerated exceptions, including text-and-data-mining rules with eligibility, access or opt-out conditions. The same training activity can therefore have different answers across borders.
Outputs and market harm
A model that learns statistical relationships is not necessarily storing a browsable library, but models can sometimes reproduce memorized passages or imitate commercially significant material. Courts may examine whether outputs are substantially similar to protected expression and whether the system competes with licensing or sales markets. Technical safeguards can reduce memorization but do not settle the legal test.
Reality check
A settlement binds the parties and may avoid a definitive ruling. A trial-court decision can be limited to its facts and may be appealed. “The court said AI training is legal” is usually too broad unless the jurisdiction, dataset, acquisition method, claim and procedural stage are specified.
What responsible operators document
Useful records include dataset sources, licences, opt-outs, filtering, retention, model purpose and tests for regurgitation. Legal review should cover every jurisdiction and product use involved. This explainer describes the analytical framework, not legal advice for a particular model or dispute.