Copyright and AI models: can a model itself contain a copy?

Why the AI model cannot be ignored from a legal standpoint
The production chain of generative AI roughly consists of several phases. First, data is collected and used as training data. Next, a model is trained. That model is then integrated into an AI system that allows users to generate output via prompts.
Legal discussions often emphasize the beginning and end of that chain. Is it permissible to use copyrighted material for training? And when does the output infringe upon an existing work?
However, the trained AI model itself sits in between. This is precisely where a third copyright question can arise: can a protected work be encoded into a model's parameters during the training process in such a way that the model itself constitutes a reproduction?
This is more than a theoretical discussion. If a protected work is indeed stored within the model, the focus shifts from incidental output to the technical source from which that output can be generated. For companies that develop, offer, or incorporate models into their own products, this can make a fundamental legal difference.
Article 14 of the Copyright Act: copyright is not dependent on the medium
The notion that a copyrighted work is only copied when a human can directly see or read the copy has long been outdated in copyright law.
Article 14 of the Copyright Act plays an important role here. This provision clarifies that the recording of a work, or a part thereof, onto an object that can make the work audible or visible can also qualify as a reproduction.
The historical background of this is remarkably relevant today.
At the beginning of the twentieth century, for example, musical works could be stored on piano rolls. Such a roll did not contain traditional musical notation. The music was encoded in patterns of perforations that were converted into piano playing by a machine. Anyone looking only at the roll would not see a piece of music in a directly understandable form.
At the time, this led to the argument that the musical work itself was not present on the medium. After all, the medium only contained technical patterns.
Copyright law was specifically developed to prevent a new technical form of recording from causing a protected work to suddenly disappear from a legal perspective. The focus is not on readability for the human eye, but on whether the work has actually been recorded and can be made perceptible with the aid of a technical device.
This technology neutrality is particularly relevant for AI.
A protected work does not need to exist in a system as a recognizable image, text, or audio file for a copyright-relevant reproduction to occur. An indirect, encoded, or otherwise technically configured reproduction can also fall under the right of reproduction.
Statistical patterns are not automatically irrelevant to copyright
During training, AI models store information in a vast number of parameters. These parameters contain statistical relationships that subsequently help determine the output the model produces.
A common assumption is that a model therefore cannot contain works. The model is thought to store only statistics: numbers, probabilities, patterns, and relationships.
However, it does not legally follow that the stored information is unprotected.
A statistical pattern can be highly abstract, but it can also be extremely specific. A model may learn a general rule about language, images, or music. Such a rule need not contain anything from an individual protected work. But a pattern can also become so specific that concrete, protected characteristics of a particular work can be reconstructed from it.
The problem lies in the leap from "the model consists of statistical patterns" to "therefore, the model cannot contain copyright-protected expression." The second conclusion does not automatically follow from the first.
Generalization and memorization are distinct phenomena.
To understand this discussion, the distinction between generalization and memorization is useful.
In generalization, a model extracts broader information from training data. For example, it learns structures, relationships, or general characteristics that are not necessarily tied to one specific protected work.
In memorization, information about a specific work is captured much more concretely within the model. This allows the model to reproduce that work, or relevant protected parts of it, at a later time.
Incidentally, this does not mean that "memorization" and "reproduction" are legally identical. Memorization is originally a technical term. Computer science research, for instance, often looks for output that corresponds almost verbatim to material from the training data.
Copyright law can be broader. A reproduction does not always have to be one hundred percent identical. A copyright-relevant adaptation can also incorporate protected material without being an exact copy.
For a legal analysis, one must therefore ask more than just whether a model has technically "memorized" something. The relevant question is whether copyright-protected features of a work have been captured in such a way that it could constitute a reproduction.
A model contains more than just loose building blocks.
Another argument is that an AI model contains, at most, small fragments or loose elements from training data. Individual facts, words, ideas, stylistic elements, and other non-protected building blocks may be used freely on their own.
That premise remains important. The mere fact that a system can eventually create something that resembles a protected work using free building blocks does not mean that the protected work itself is already present in the model.
However, the analysis changes when the model contains more than just loose building blocks.
If the statistical structure of the model is so specific that a particular work can be regenerated with a simple prompt, it becomes harder to argue that only unprotected fragments are present. The model may then function as a technically encoded storage form from which protected expression can be retrieved.
The difference is significant. A box of general building blocks does not automatically contain the building that can be made with them. But when the information to reconstruct a specific protected result is embedded in the technical structure, the situation comes much closer to a compressed or encoded reproduction.
Compression does not automatically render a copy harmless.
Digital information does not need to be stored in its original form.
We have known this for a long time from digital compression. For example, an image can be technically converted and highly compressed, while the characteristic content can still be reconstructed. The technical form in which information is stored, therefore, does not necessarily determine its copyright status.
The same may be relevant for AI.
The fact that information is incorporated into a model in a distributed, fragmented, or mathematically encoded way does not rule out the presence of protected elements. If that information is retained latently and can later be made visible or audible again, the method of storage alone may be insufficient to rule out a reproduction.
For AI companies, a functional approach is therefore particularly relevant: what can actually be reconstructed based on the stored model parameters?
The AI model as a new type of medium
Terminology sometimes makes this discussion more complicated than necessary.
Words like "learning," "remembering," "neurons," and "recalling" make AI easier to understand, but they are metaphors. An AI model is not a human brain. It is a technical system in which information is captured and processed by means of parameters.
For copyright law, that distinction is important.
Human memory is not treated as a copyright-protected data carrier. A trained AI model, by contrast, is a stable technical object whose parameters can be stored, copied, distributed, and processed within software systems.
This allows an AI model to function as a special type of medium in certain situations.
The difference from a traditional medium is that a model can do more than simply reproduce stored information. It can also calculate new output based on the patterns formed during training. A model can therefore simultaneously carry information and generate new content.
That makes it legally more complex, but not necessarily fundamentally different. If a specific protected work is actually captured within the model, that capture does not disappear just because the same model can perform billions of other calculations as well.
The output can reveal what is inside the model
One of the most difficult aspects of modern AI models is their opacity. Even for developers, it is not always possible to determine exactly what information is stored in the parameters, where, or in what way.
This gives rise to an important question regarding the burden of proof.
Suppose a user enters a simple prompt and the AI system subsequently generates an output that is strikingly similar to a protected work on which the model was trained. Must the rights holder be able to technically point out exactly in which parameters that work is stored and how that storage occurred during the training process?
In practice, such a requirement would be exceptionally burdensome. The model can consist of billions of parameters, and the rights holder generally has no access to its technical inner workings.
The output can therefore be an important legal indicator.
When a specific reproduction can be generated from a model following a simple prompt, it may indicate that relevant information about the underlying work was already present in the model. In that case, the output is not only the potential endpoint of an infringement but can also have evidentiary value regarding what was captured in the model during training.
This aligns with a broader principle in copyright law regarding evidence. When the similarity between two works is so striking that independent creation becomes unlikely, the alleged infringer may be required to provide an explanation for that similarity. In the context of generative AI, this can be relevant when it is established that the work in question was also part of the training data.
This does not mean that every similarity between AI output and an existing work proves that a copy exists within the model. The circumstances of the specific case remain decisive. It does mean, however, that the technical complexity of a model does not automatically serve as a legal shield.
Why a single internal copy can have major consequences
The extent of memorization in generative AI models is not easy to determine. Technical estimates vary and often focus primarily on near-verbatim reproductions. The copyright concept of reproduction can be broader.
Even when only a relatively small portion of a massive training dataset remains in a model, the absolute number of works involved can be large.
There is a second aspect to consider: scalability.
A traditional copy is a single copy. A reproduction stored within a model can potentially function as a kind of master copy from which output is repeatedly generated. If an AI system lacks adequate restrictions, users can repeatedly request output that reproduces or adapts the stored material, either in whole or in part.
For generative AI companies, the risk therefore does not only concern what happens during training. The combination of the underlying model, the user interface, permitted prompts, and any output filters is also relevant.
The TDM exception does not automatically solve the model issue
Text and data mining, often abbreviated as TDM, also plays an important role in AI training.
One possible legal approach is that certain reproductions necessary for text and data mining may fall under a statutory exception. However, equating generative AI training with TDM is a subject of debate.
Even if it is assumed that the TDM exception applies to the training process, it does not automatically follow that every copy that remains in the model after training is also permitted.
Article 15o of the Copyright Act links the retention of a reproduction made for text and data mining to the period necessary for that text and data mining. Based on this premise, it is arguable that a protected work that remains permanently in the model parameters after training is completed is not necessarily justified by the same exception.
The distinction between temporary reproductions during the training process and a permanent reproduction in the trained model is therefore essential.
For AI startups relying on a TDM basis, it is risky to treat training and model storage as a single legal event. The question is not only what happens during the analysis of training data, but also what actually remains in the model afterwards.
What does this mean for startups and scale-ups?
For tech companies, it is crucial that copyright analysis does not stop at the dataset. Anyone developing or operating generative AI must also look at the trained model and the relationship between that model and its output.
In practice, five questions are particularly relevant:
- What remains in the model after training? A process that only generalizes abstract patterns raises different questions than a model that stores specific works in a reproducible way.
- What output can be generated with simple prompts? Regular or strikingly precise reproductions may indicate a problem that runs deeper than just the behavior of a single user.
- What technical limitations are built in? Filters and other measures can be relevant to prevent protected output from being easily reproduced or provided to users.
- On what legal basis is the training founded? Justification for certain training activities does not automatically mean that a permanent copy in the final model is also justified.
- What is the company's role in the AI value chain? It makes a difference whether a startup trains a model itself, distributes an existing model, or builds its own AI system on top of one. Models, systems, and outputs are legally distinct components of the same production chain.
For founders and investors, this also means that an AI model should not be viewed solely as a technical asset. The origin of the training data, the characteristics of the trained model, and the methods used to prevent protected output can all be part of that asset's legal risk profile.
Technology neutrality is essential, especially in AI
Generative AI makes it tempting to seek entirely new legal rules for new technical concepts. However, a significant part of copyright law was specifically designed to withstand technological change.
A work does not necessarily disappear in a legal sense just because it is converted into perforations, digital code, compressed data, or statistical parameters. The decisive factor may be whether protected expression has actually been fixed and can be made perceptible again from that fixation.
It does not follow that every AI model contains copyright-protected works. Nor does a model consist entirely of copies of its training data. The relevant position lies between these extremes: models can contain both unprotected general patterns and concrete, protected information.
For startups and scale-ups, that nuance is precisely what matters. The question is not whether an AI model contains "just statistics," but what those statistics represent and can reproduce in a specific case.
Conclusion: look beyond training and output to the model itself
The copyright debate surrounding generative AI has three relevant levels: the training phase, the trained model, and the final output. Those who look only at input and output may miss a legally significant part of the chain.
Article 14 of the Copyright Act supports a technology-neutral approach. A work does not need to be directly readable, visible, or recognizable in a file to be considered fixed. Even a statistical or highly compressed representation can be relevant under copyright law if protected characteristics are retained within it and can be made perceptible again.
For AI startups and scale-ups, this makes model behavior itself a legal focal point. Not every form of learning by a model constitutes a reproduction, but the argument that a model cannot, by definition, contain protected works because it consists only of parameters and statistical patterns is too simplistic. Especially with generative AI, one must look not only at what goes in and what comes out, but also at what actually remains in the model.



















