By Stuart Kerr, Technology Correspondent, LiveAIWire
AI training data lawsuits reached a new scale when Reddit filed suit against Anthropic in 2024, alleging that the AI company used Reddit posts to train its Claude models without authorisation, compensation, or adequate consideration of the rights of the users who created the content. The lawsuit joined a rapidly growing body of litigation challenging the foundational data practices of the AI industry, alongside cases brought by the New York Times against OpenAI, Getty Images against Stability AI, and dozens of individual authors and artists against multiple AI companies. Collectively, these cases are forcing courts in the United States and Europe to answer questions that existing intellectual property law was not designed to address.
The legal questions at the centre of these AI training data lawsuits are genuinely novel and technically complex. Copyright law protects specific expression rather than ideas or facts, and the question of whether training a machine learning model on copyrighted text constitutes infringement depends on how courts characterise what the model does with that text. Does a large language model store and reproduce copyrighted expression in a way that violates the reproduction right? Or does it learn patterns and relationships more analogous to how a human reader absorbs information from books they have read?
The Fair Use Question Behind AI Training Data Lawsuits
In the United States, the primary legal battleground is fair use doctrine, which permits certain uses of copyrighted material without licence or compensation based on a four-factor test examining purpose, nature of use, amount used, and market effect. AI companies including Anthropic and OpenAI argue that training on publicly available internet content is transformative fair use analogous to a human reading and learning from published material. Copyright holders argue that training data use is not transformative, that AI models can and do reproduce training data in their outputs, and that the market effect of AI on the market for licensed content is substantial and negative.
The New York Times case against OpenAI, filed in December 2023, includes evidence that ChatGPT can reproduce substantially verbatim paragraphs from Times articles when prompted in certain ways, a finding that directly challenges the claim that AI models do not reproduce training data in outputs. OpenAI disputes the characterisation of these examples as representative of normal model behaviour, arguing they represent edge cases that its safety guidelines are designed to prevent. The court’s eventual assessment of this contested evidence will significantly shape how fair use applies to AI training.
The Licensing Alternative
Several AI companies and content providers have pursued negotiated licensing agreements as an alternative to fighting AI training data lawsuits through the courts. OpenAI has signed agreements with the Associated Press, Axel Springer, and several other publishers that provide financial compensation in exchange for access to content archives for training.
Reddit’s own agreement with Google, announced in February 2024 and valued at approximately 60 million dollars per year, provides a data point for the commercial value that social media platforms can extract from their content archives through AI licensing rather than litigation. The Authors’ Licensing and Collecting Society in the UK has been developing collective licensing frameworks that would allow AI companies to license content from authors and publishers through a single negotiated arrangement.
UK and EU Legal Frameworks
The legal landscape outside the United States is different in important respects. In the EU, the Copyright in the Digital Single Market Directive includes a text and data mining exception that permits mining of publicly accessible content for research purposes but allows rights holders to opt out of mining for commercial purposes. Several major publishers have exercised this opt-out, creating a class of content that AI companies cannot legally mine under EU law without a licence.
The UK has a similar research-purpose exception but has been debating whether to introduce a broader text and data mining exception for commercial AI development, a policy question that has generated sustained lobbying from both the AI industry and the creative sector. The Intellectual Property Office consultation on AI and copyright produced responses that reflected the depth of disagreement between these constituencies, and the government’s eventual policy position will significantly affect the UK’s relative attractiveness as an AI development hub. This same jurisdictional fragmentation runs through LiveAIWire’s coverage of AI power concentration, where regulatory divergence between major markets creates uneven leverage for the largest AI developers.
What This Means for You
If you are a creator of any kind, whether a writer, photographer, musician, software developer, or visual artist, your work has almost certainly been used to train AI models, a concentration of value LiveAIWire has traced in our coverage of the AI shadow workforce, where similar questions of consent and compensation apply to human labour rather than published text. Whether this use was legal, and whether you have any right to compensation, depends on questions that courts and legislators are only now beginning to answer.
Supporting the advocacy organisations representing creators’ interests in these debates, including the Society of Authors, the NUJ, and the ALCS in the UK, is a way of ensuring that creator perspectives are adequately represented in policy processes that will shape the legal framework for AI and creative work for years to come.
Where the AI Training Data Lawsuits Are Headed
The timeline for resolution of the major AI training data lawsuits is uncertain, but the legal landscape is likely to look materially different within two to three years as courts in the US, UK, and EU produce initial rulings that establish precedent. The outcome most favourable to AI companies is a broad fair use ruling that permits training on publicly available internet content without consent or compensation; the outcome most favourable to rights holders is a ruling that training constitutes infringement requiring licence and compensation.
The most likely outcome is something in between, perhaps distinguishing between different types of content or different types of training use, that creates a framework requiring negotiated licensing for some categories of content while permitting other uses. This same tension between innovation speed and legal accountability echoes what LiveAIWire has traced in our coverage of the AI right to be forgotten, where enforcement technology also lags well behind the legal right it is meant to support.
About the Author
Stuart Kerr is Technology Correspondent at LiveAIWire, covering artificial intelligence, cybersecurity, and the social impact of emerging technology. He publishes daily at LiveAIWire.com.