Technology explainer
When Can AI Training Use Copyrighted Material Legally?
The legality of AI training depends on jurisdiction, how material was acquired, what was copied, and how the resulting system affects the market. A settlement may resolve one dispute without answering the broader legal question.
AI models learn statistical patterns from large collections of text, images, audio, or code. When those collections include copyrighted works, two separate questions arise: whether the copies were lawfully obtained, and whether using them for training is legally permitted.
What does copyright protect?
Copyright protects original expression, not facts, ideas, or general styles. It gives rights holders control over activities such as copying, distribution, and some derivative uses.
Why does AI training involve copying?
Training systems usually create or process copies of works so software can analyse them. Courts may examine both how those copies were acquired and what the training process does with them.
What is fair use?
In the United States, fair use is evaluated through several factors, including purpose, the nature of the work, the amount used, and market effect. No single factor decides every case.
Why does lawful acquisition matter?
A court may treat the use of lawfully purchased material differently from books or files obtained through piracy. The training purpose does not automatically erase problems in acquisition.
Does a legal settlement create a general rule?
No. A settlement resolves claims between particular parties. It may avoid a final ruling on questions that affect the wider industry.
Why is the issue still unsettled?
Different cases involve different datasets, model designs, outputs, and markets. Laws also vary across countries, so one judgment may not apply globally.
First appeared in
Anthropic Will Pay $1.5 Billion for Pirated Books. The Bigger AI Question Is Still Open