Lecture 33: Copyright and Artificial Intelligence — Limits

Block: Limits. Topic: Copyright. Target audience: interested adults. Total duration: 50 minutes (Part 1: 20 min, Part 2: 20 min, Part 3: 10 min).

Note: The lecture relies on publicly available sources (bibliography at the end). The legal situation and case law are in flux; where relevant I point out uncertainties and open questions.

Part 1 (approx. 20 minutes): Basic understanding — central questions and simple analogies

The debate about copyright and AI focuses on two core questions: Is an AI allowed to train on copyrighted works, and who owns the content generated by AI? Both questions depend on how the law defines infringement, use, and human authorship.

An illustrative analogy: Imagine a student reads hundreds of thousands of books, newspaper articles and artworks to learn writing and stylistic devices. Afterwards the student composes independent texts. Is that writing a copy of a particular book, or is it a new work shaped by the learning? Formally similar processes occur when training large models: algorithms analyze large amounts of texts and images, extract statistical patterns, and generate new outputs from them.

Whether the "reading" and "learning" (that is, the automatic ingestion and processing of texts and images) is legally permitted depends on national and supranational law. In the European Union the so-called Digital Single Market (DSM) framework provides exceptions for text and data mining (TDM); these exceptions are not fully uniform and distinguish between research and commercial applications. In the USA the question often arises under the fair use doctrine, while authorities such as the U.S. Copyright Office maintain that works without a human author are not registrable.

Practical examples: A AI developer copies text corpora from the web, including copyrighted books. Whether this is legally permissible can depend on a TDM exception, a judicial fair use assessment, or a license. Another example: An image generator uses millions of licensed and unlicensed images for training. Image agencies and individual authors have since initiated legal action; courts now have to resolve a range of questions.

Important conclusion from this overview: There is no single, globally valid answer. Rather, the situation is a mixture of statutory exceptions, procedural decisions and administrative practice — and these differ between the EU, the USA and other jurisdictions.

Part 2 (approx. 20 minutes): Deepening — key terms, legal institutions and current conflicts

Essential terms and legal principles

Copyright: This right protects intellectual creations in literature, science and the arts. Protection arises automatically upon creation in many legal systems; certain requirements of originality and human creation may be required.

Text and data mining (TDM): TDM denotes automated procedures for searching and analyzing large datasets. The EU Copyright Directive 2019/790 (DSM Directive) contains mandatory exceptions for TDM for research organizations and optional exceptions for other users; in practically relevant cases it is therefore necessary to examine whether the conditions of the respective exception are met.

Fair use (USA): In the United States fair use is a flexible, fact-based balancing test. Courts examine four factors to decide whether a use is permissible despite the absence of a license. Earlier cases such as the dispute between the Authors Guild and Google Books show that courts can consider mass processing akin to TDM permissible in certain contexts if a transformative benefit exists.

Copyright eligibility of AI works: Administrative view (e.g. U.S. Copyright Office) generally requires a human contribution for a work to be eligible for protection. Fully autonomously generated content without a significant human creative input is often not registrable as protected.

Current conflicts and litigation (status based on publicly available reports)

Since 2023 several notable lawsuits have been filed in which authors and rights holders claim that large models were trained on protected works without permission. Examples include cases against companies that develop or operate AI models; image agencies and authors point to unauthorized use of works, and rights holders seek compensation or injunctive relief. Many of these proceedings are still ongoing or have been partially settled out of court; therefore there is uncertainty about final binding principles.

In parallel, authorities and international institutions are preparing statements or guidelines: In the USA the Copyright Office has published guidance on registration and the role of human authorship. The EU regulates TDM at the directive level, with varying implementation across member states. The World Intellectual Property Organization (WIPO) monitors and analyzes developments and has produced several overviews on AI and intellectual property.

Classifying technical and legal questions

Legally important is the distinction between input and output: The ingestion (input) of copyrighted works for the purpose of training can be assessed separately from the question whether a particular generated text or image file constitutes an unauthorized copy of a protected work. Judicial analyses often examine whether an output reproduces substantial parts of a protected work or whether it should be regarded as a new, independent creation.

Contractual issues and licenses also play a central role: Many providers work with licensed datasets or enter license agreements with rights holders. The legal disputes therefore concern not only abstract principles but concretely which uses of which datasets are permitted or must be remunerated.

Policy options and directions of development

From a legal perspective several development paths can be distinguished: (1) expansion of exceptions (e.g. broader TDM permissions), (2) stronger regulation through remuneration obligations for the use of training data, (3) clear codification of the necessity of human authorship for eligibility, and (4) specialized rules for generative AI (e.g. reporting obligations, transparency duties towards rights holders). Various political actors propose or consider combinations of these options; so far there is no internationally unified approach.

In all proposals conflicts of interest must be considered: protection of authors and collecting societies on one side, and promotion of innovation and scale effects in AI development on the other. Courts and legislators will attempt to balance these interests; the exact design remains open.

Uncertainty notice: Because many cases have not yet been finally decided and legislative changes are underway, the concrete legal situation in individual applications often remains unclear. A reliable case-by-case assessment remains necessary.

Part 3 (approx. 10 minutes): Applications, limits and short thought exercises

Concrete applications and their legal risks:

1) Operator of a chatbot trained on web corpora: Risk if outputs reproduce long passages from a single copyrighted work; less risk if outputs are transformative and do not reproduce the substance of a single work. Practically relevant: logging, filter mechanisms and licensing can reduce risk.

2) Image generation system that used millions of images: Rights holders may claim that their works were used without permission. Providers should review which data sources they use, whether they hold licenses and which measures (e.g. blocklists, attribution, compensation mechanisms) can be implemented.

3) Scientific research: In the EU there are special TDM exceptions for research organizations; this facilitates the use of text and image corpora for purely scientific purposes, but requires a careful review of national implementation and any limitations.

Small thought exercises for reflection (briefly discuss each):

a) Suppose you operate a language model that produces answers in the style of a well-known living author. What risks do you see, and what technical or legal measures could you take? (Expected: review of rights of imitation/personality rights, license agreements, style imitation vs. direct quotation obligations.)

b) You own photos that are publicly accessible in a large internet archive. Can you simply demand that these photos be removed from training datasets? (Expected: review of the platform's terms of use, legal actions, balancing with transformative use and practical enforcement hurdles.)

c) Briefly discuss the difference between "training use" and "output" using a concrete example: A model outputs a text that reproduces a sequence of 300 words from a copyrighted novel almost verbatim. Which criteria are relevant to assess whether this is an unauthorized use? (Expected: substance of the reproduction, likelihood of recognizability, transformative use, possible licenses.)

Concluding practical advice:

Organizations should provide transparency about data sources used, carry out risk assessments, consider licenses or opt-out mechanisms as appropriate, and seek legal advice in case of legal uncertainty. At the regulatory level it should be observed how courts and legislators strike the balance between copyright and the interest in innovation.

Sources (selected, actually used sources)

1) Directive (EU) 2019/790 of the European Parliament and of the Council on copyright and related rights in the Digital Single Market (DSM Directive). Publisher: EUR-Lex / European Union. Date of adoption: 17 April 2019. URL: https://eur-lex.europa.eu/eli/dir/2019/790/oj

2) U.S. Copyright Office — "Copyright and Artificial Intelligence." Publisher: U.S. Copyright Office. (Collection of guidance and publications on how the Copyright Office deals with AI-generated content). Status and documentation on the website. URL: https://www.copyright.gov/policy/artificial-intelligence/

3) Reuters — "Authors Guild sues OpenAI, Microsoft, alleging copyright infringement" (report on the lawsuit by authors' associations against AI providers). Publisher: Reuters. Date of publication: 18 September 2023. URL: https://www.reuters.com/legal/authors-guild-sues-openai-microsoft-alleging-copyright-infringement-2023-09-18/

4) Reuters — "Getty Images sues Stability AI alleging copyright infringement" (report on Getty Images' lawsuit against Stability AI). Publisher: Reuters. Date of publication: 28 March 2023. URL: https://www.reuters.com/technology/getty-images-sues-stability-ai-alleging-copyright-infringement-2023-03-28/

5) World Intellectual Property Organization (WIPO) — "WIPO Technology Trends: Artificial Intelligence." Publisher: WIPO. Collection and analysis on interfaces between AI and intellectual property. URL: https://www.wipo.int/tech_trends/en/artificial_intelligence/

6) Court decision: Authors Guild v. Google (Second Circuit) — decision addressing mass digitization in the context of Google Books and the question of fair use. Second U.S. Court of Appeals decision, 16 October 2015. Source: Law.justia (full text of the decision). URL: https://law.justia.com/cases/federal/appellate-courts/ca2/13-4829/13-4829-2015-10-16.html

7) Supreme Court of the United States — Andy Warhol Foundation for the Visual Arts, Inc. v. Lynn Goldsmith, No. 21‑869, decision (arguments and conclusions regarding the fair use doctrine in image uses). Publisher: Supreme Court of the United States. Date of decision: 2023. URL: https://www.supremecourt.gov/opinions/22pdf/21-869_6k47.pdf

Note on the sources: The selection above forms the basis for the verifiable statements made in the lecture. Many current proceedings are still pending; the sources reflect the state of public reporting and official guidance up to June 2024. There are numerous additional publications (academic articles, policy papers, further press reports) that are not listed here but can be used for more in-depth research.