Limitations, Hallucination & Risks

copyright and LLMs

Modern models are trained on enormous amounts of text and images, much of it written or drawn by people who hold copyright and were never asked. That raises two distinct legal questions. First, is it lawful to train on protected works without permission — defended by some as fair use or research, contested by authors, artists, and news organizations who say their labor was taken. Second, can a model's output itself infringe, by reproducing a passage or a recognizable style closely enough to substitute for the original.

None of this is settled, and the answers vary by country and are being fought out in courts and legislatures right now. For someone using these tools, the practical caution is twofold: outputs can occasionally echo training text closely enough to create exposure, and the provenance of generated material is murky, so it pays to check before publishing commercially. This is a legal and ethical frontier, not a solved engineering problem, and the rules you build on today may shift.