copyright and training data
/ KOP-ee-ryt and TRAY-ning DAY-tuh /
Modern AI models learn by ingesting staggering amounts of text, images, code, and music — much of it scraped from the open internet, and much of that protected by copyright, the legal right of a creator to control copying of their work. "Copyright and training data" is the tangle of questions this raises: Is it legal to feed a billion copyrighted images into a model without asking? Does the resulting model, or its output, infringe? And who, if anyone, owes the original creators?
Several distinct legal questions hide inside the controversy, and they have different answers. (1) The input question: was copying the works to train the model itself an infringement, or is it permitted as "fair use" / "text and data mining"? (2) The memorization question: does the model sometimes spit out near-verbatim copies of specific works (it sometimes can), which is a clearer problem? (3) The output question: who owns what the model generates, and can a generated image that closely imitates a living artist's style infringe? (4) The authorship question: can purely AI-made work even be copyrighted at all? Different countries are answering these differently, and many cases are still being litigated.
Why it matters: the answers will shape who profits from AI and whether human creators get paid or pushed out. The honest state of things as of the mid-2020s is unsettled — there is no global consensus, courts are issuing their first rulings, and reasonable people disagree. Beware confident claims in either direction ("it's obviously theft" / "it's obviously fair use"); the law genuinely hasn't decided, and "the model learned from it like a human would" is a rhetorical analogy, not an established legal principle.
A novelist discovers her books were part of a dataset used to train a chatbot, and that the bot can summarize her plots in detail. She sues. The case forces a court to decide: was copying her books to train the model an infringement, or transformative fair use? As of the mid-2020s, courts around the world are issuing their first, often conflicting, answers.
The same facts are being judged differently across jurisdictions — the law is still being written.
Two things often get conflated: training on copyrighted data, and a model reproducing it. The first is legally murky and hotly contested; the second — a model emitting a near-exact copy of a specific protected work — is a much clearer infringement risk. Lumping them together muddies a debate where the distinctions matter.