Copyright and AI models: can a model itself contain a copy?

In the context of generative AI, copyright discussions often focus on training data and output. However, from a legal perspective, the AI model itself may be the crucial missing link: it can contain copyrighted works stored in a statistical format. For AI startups, scale-ups, and other tech companies, this directly impacts training, product design, filters, licensing, and liability risks.
No items found.
Insights
Caylun J. Scholtens
17.08.2026

Why the AI model cannot be ignored from a legal standpoint

The production chain of generative AI roughly consists of several phases. First, data is collected and used as training data. Next, a model is trained. That model is then integrated into an AI system that allows users to generate output via prompts.

Legal discussions often emphasize the beginning and end of that chain. Is it permissible to use copyrighted material for training? And when does the output infringe upon an existing work?

However, the trained AI model itself sits in between. This is precisely where a third copyright question can arise: can a protected work be encoded into a model's parameters during the training process in such a way that the model itself constitutes a reproduction?

This is more than a theoretical discussion. If a protected work is indeed stored within the model, the focus shifts from incidental output to the technical source from which that output can be generated. For companies that develop, offer, or incorporate models into their own products, this can make a fundamental legal difference.

Article 14 of the Copyright Act: copyright is not dependent on the medium

The notion that a copyrighted work is only copied when a human can directly see or read the copy has long been outdated in copyright law.

Article 14 of the Copyright Act plays an important role here. This provision clarifies that the recording of a work, or a part thereof, onto an object that can make the work audible or visible can also qualify as a reproduction.

The historical background of this is remarkably relevant today.

At the beginning of the twentieth century, for example, musical works could be stored on piano rolls. Such a roll did not contain traditional musical notation. The music was encoded in patterns of perforations that were converted into piano playing by a machine. Anyone looking only at the roll would not see a piece of music in a directly understandable form.

At the time, this led to the argument that the musical work itself was not present on the medium. After all, the medium only contained technical patterns.

Copyright law was specifically developed to prevent a new technical form of recording from causing a protected work to suddenly disappear from a legal perspective. The focus is not on readability for the human eye, but on whether the work has actually been recorded and can be made perceptible with the aid of a technical device.

This technology neutrality is particularly relevant for AI.

A protected work does not need to exist in a system as a recognizable image, text, or audio file for a copyright-relevant reproduction to occur. An indirect, encoded, or otherwise technically configured reproduction can also fall under the right of reproduction.

Statistical patterns are not automatically irrelevant to copyright

During training, AI models store information in a vast number of parameters. These parameters contain statistical relationships that subsequently help determine the output the model produces.

A common assumption is that a model therefore cannot contain works. The model is thought to store only statistics: numbers, probabilities, patterns, and relationships.

However, it does not legally follow that the stored information is unprotected.

A statistical pattern can be highly abstract, but it can also be extremely specific. A model may learn a general rule about language, images, or music. Such a rule need not contain anything from an individual protected work. But a pattern can also become so specific that concrete, protected characteristics of a particular work can be reconstructed from it.

The problem lies in the leap from "the model consists of statistical patterns" to "therefore, the model cannot contain copyright-protected expression." The second conclusion does not automatically follow from the first.

Generalization and memorization are distinct phenomena.

To understand this discussion, the distinction between generalization and memorization is useful.

In generalization, a model extracts broader information from training data. For example, it learns structures, relationships, or general characteristics that are not necessarily tied to one specific protected work.

In memorization, information about a specific work is captured much more concretely within the model. This allows the model to reproduce that work, or relevant protected parts of it, at a later time.

Incidentally, this does not mean that "memorization" and "reproduction" are legally identical. Memorization is originally a technical term. Computer science research, for instance, often looks for output that corresponds almost verbatim to material from the training data.

Copyright law can be broader. A reproduction does not always have to be one hundred percent identical. A copyright-relevant adaptation can also incorporate protected material without being an exact copy.

For a legal analysis, one must therefore ask more than just whether a model has technically "memorized" something. The relevant question is whether copyright-protected features of a work have been captured in such a way that it could constitute a reproduction.

A model contains more than just loose building blocks.

Another argument is that an AI model contains, at most, small fragments or loose elements from training data. Individual facts, words, ideas, stylistic elements, and other non-protected building blocks may be used freely on their own.

That premise remains important. The mere fact that a system can eventually create something that resembles a protected work using free building blocks does not mean that the protected work itself is already present in the model.

However, the analysis changes when the model contains more than just loose building blocks.

If the statistical structure of the model is so specific that a particular work can be regenerated with a simple prompt, it becomes harder to argue that only unprotected fragments are present. The model may then function as a technically encoded storage form from which protected expression can be retrieved.

The difference is significant. A box of general building blocks does not automatically contain the building that can be made with them. But when the information to reconstruct a specific protected result is embedded in the technical structure, the situation comes much closer to a compressed or encoded reproduction.

Compression does not automatically render a copy harmless.

Digital information does not need to be stored in its original form.

We have known this for a long time from digital compression. For example, an image can be technically converted and highly compressed, while the characteristic content can still be reconstructed. The technical form in which information is stored, therefore, does not necessarily determine its copyright status.

The same may be relevant for AI.

The fact that information is incorporated into a model in a distributed, fragmented, or mathematically encoded way does not rule out the presence of protected elements. If that information is retained latently and can later be made visible or audible again, the method of storage alone may be insufficient to rule out a reproduction.

For AI companies, a functional approach is therefore particularly relevant: what can actually be reconstructed based on the stored model parameters?

The AI model as a new type of medium

Terminology sometimes makes this discussion more complicated than necessary.

Words like "learning," "remembering," "neurons," and "recalling" make AI easier to understand, but they are metaphors. An AI model is not a human brain. It is a technical system in which information is captured and processed by means of parameters.

For copyright law, that distinction is important.

Human memory is not treated as a copyright-protected data carrier. A trained AI model, by contrast, is a stable technical object whose parameters can be stored, copied, distributed, and processed within software systems.

This allows an AI model to function as a special type of medium in certain situations.

The difference from a traditional medium is that a model can do more than simply reproduce stored information. It can also calculate new output based on the patterns formed during training. A model can therefore simultaneously carry information and generate new content.

That makes it legally more complex, but not necessarily fundamentally different. If a specific protected work is actually captured within the model, that capture does not disappear just because the same model can perform billions of other calculations as well.

The output can reveal what is inside the model

One of the most difficult aspects of modern AI models is their opacity. Even for developers, it is not always possible to determine exactly what information is stored in the parameters, where, or in what way.

This gives rise to an important question regarding the burden of proof.

Suppose a user enters a simple prompt and the AI system subsequently generates an output that is strikingly similar to a protected work on which the model was trained. Must the rights holder be able to technically point out exactly in which parameters that work is stored and how that storage occurred during the training process?

In practice, such a requirement would be exceptionally burdensome. The model can consist of billions of parameters, and the rights holder generally has no access to its technical inner workings.

The output can therefore be an important legal indicator.

When a specific reproduction can be generated from a model following a simple prompt, it may indicate that relevant information about the underlying work was already present in the model. In that case, the output is not only the potential endpoint of an infringement but can also have evidentiary value regarding what was captured in the model during training.

This aligns with a broader principle in copyright law regarding evidence. When the similarity between two works is so striking that independent creation becomes unlikely, the alleged infringer may be required to provide an explanation for that similarity. In the context of generative AI, this can be relevant when it is established that the work in question was also part of the training data.

This does not mean that every similarity between AI output and an existing work proves that a copy exists within the model. The circumstances of the specific case remain decisive. It does mean, however, that the technical complexity of a model does not automatically serve as a legal shield.

Why a single internal copy can have major consequences

The extent of memorization in generative AI models is not easy to determine. Technical estimates vary and often focus primarily on near-verbatim reproductions. The copyright concept of reproduction can be broader.

Even when only a relatively small portion of a massive training dataset remains in a model, the absolute number of works involved can be large.

There is a second aspect to consider: scalability.

A traditional copy is a single copy. A reproduction stored within a model can potentially function as a kind of master copy from which output is repeatedly generated. If an AI system lacks adequate restrictions, users can repeatedly request output that reproduces or adapts the stored material, either in whole or in part.

For generative AI companies, the risk therefore does not only concern what happens during training. The combination of the underlying model, the user interface, permitted prompts, and any output filters is also relevant.

The TDM exception does not automatically solve the model issue

Text and data mining, often abbreviated as TDM, also plays an important role in AI training.

One possible legal approach is that certain reproductions necessary for text and data mining may fall under a statutory exception. However, equating generative AI training with TDM is a subject of debate.

Even if it is assumed that the TDM exception applies to the training process, it does not automatically follow that every copy that remains in the model after training is also permitted.

Article 15o of the Copyright Act links the retention of a reproduction made for text and data mining to the period necessary for that text and data mining. Based on this premise, it is arguable that a protected work that remains permanently in the model parameters after training is completed is not necessarily justified by the same exception.

The distinction between temporary reproductions during the training process and a permanent reproduction in the trained model is therefore essential.

For AI startups relying on a TDM basis, it is risky to treat training and model storage as a single legal event. The question is not only what happens during the analysis of training data, but also what actually remains in the model afterwards.

What does this mean for startups and scale-ups?

For tech companies, it is crucial that copyright analysis does not stop at the dataset. Anyone developing or operating generative AI must also look at the trained model and the relationship between that model and its output.

In practice, five questions are particularly relevant:

  • What remains in the model after training? A process that only generalizes abstract patterns raises different questions than a model that stores specific works in a reproducible way.
  • What output can be generated with simple prompts? Regular or strikingly precise reproductions may indicate a problem that runs deeper than just the behavior of a single user.
  • What technical limitations are built in? Filters and other measures can be relevant to prevent protected output from being easily reproduced or provided to users.
  • On what legal basis is the training founded? Justification for certain training activities does not automatically mean that a permanent copy in the final model is also justified.
  • What is the company's role in the AI value chain? It makes a difference whether a startup trains a model itself, distributes an existing model, or builds its own AI system on top of one. Models, systems, and outputs are legally distinct components of the same production chain.

For founders and investors, this also means that an AI model should not be viewed solely as a technical asset. The origin of the training data, the characteristics of the trained model, and the methods used to prevent protected output can all be part of that asset's legal risk profile.

Technology neutrality is essential, especially in AI

Generative AI makes it tempting to seek entirely new legal rules for new technical concepts. However, a significant part of copyright law was specifically designed to withstand technological change.

A work does not necessarily disappear in a legal sense just because it is converted into perforations, digital code, compressed data, or statistical parameters. The decisive factor may be whether protected expression has actually been fixed and can be made perceptible again from that fixation.

It does not follow that every AI model contains copyright-protected works. Nor does a model consist entirely of copies of its training data. The relevant position lies between these extremes: models can contain both unprotected general patterns and concrete, protected information.

For startups and scale-ups, that nuance is precisely what matters. The question is not whether an AI model contains "just statistics," but what those statistics represent and can reproduce in a specific case.

Conclusion: look beyond training and output to the model itself

The copyright debate surrounding generative AI has three relevant levels: the training phase, the trained model, and the final output. Those who look only at input and output may miss a legally significant part of the chain.

Article 14 of the Copyright Act supports a technology-neutral approach. A work does not need to be directly readable, visible, or recognizable in a file to be considered fixed. Even a statistical or highly compressed representation can be relevant under copyright law if protected characteristics are retained within it and can be made perceptible again.

For AI startups and scale-ups, this makes model behavior itself a legal focal point. Not every form of learning by a model constitutes a reproduction, but the argument that a model cannot, by definition, contain protected works because it consists only of parameters and statistical patterns is too simplistic. Especially with generative AI, one must look not only at what goes in and what comes out, but also at what actually remains in the model.

Testimonials

What our clients say

Startups and scale-ups enjoy working with us. Here’s what they think of our expertise and approach:

We hired Startup-Recht to draft our general terms and service agreements. The result was fast, high-quality, and perfectly tailored to our needs thanks to the revision rounds. They really took the time to understand our business context. Professional, reliable, and a pleasure to work with.
Daan Witte
Co-founder AcuityAi
legal expertise for fast moving startups in regulated industries. Startup-Recht provides the legal foundation for us to innovate at Pabel AI.
Stan Haaijer
Co-founder Pabel B.V.
Good, energetic lawyers with clear and strong subject-matter expertise. They respond quickly and think proactively, finding solutions for innovative and sometimes complex issues within our sector: Open Source Consulting. The documents were delivered on time, and communication throughout was clear and prompt.We also had the documents reviewed by several other lawyers, who were impressed by their quality. Substantive feedback was addressed thoroughly and with great care. This gives us confidence in our new legal foundation.Thank you for the pleasant collaboration—looking forward to working together again soon.
Niels Verhage
Co-founder Rogue IT Consulting B.V.
Maarten and Caylun from Startup-Recht are supporting me in setting up my business. They do so in a very pleasant and professional manner. As an entrepreneur, it’s extremely valuable to be able to rely on their expertise in startups.I can reach out with questions whenever they arise and always receive a prompt response. In addition, they take all legal work off my hands and assist with drafting the right documents.In short, I am very happy with this collaboration and can highly recommend them.
Erik Maessen
Founder CoachChecker B.V.
We had a very pleasant collaboration. They thought along with us carefully, truly understood our vision, and supported us in a professional and approachable way. The communication was personal and clear throughout. Definitely highly recommended.
Luc de Graag
Co-founder Tikt.ai
We had an excellent experience working with Startup-Recht. Their team combines professionalism with a genuine understanding of startups’ needs, guiding us through every step with clarity and efficiency. They didn’t just answer our questions – they anticipated challenges and offered practical solutions that gave us real peace of mind. Highly recommended for any young company looking for reliable legal support.
Luis Martinez
Co-founder UpTo
Logo staallokaal
At Startup-Recht, the mix of young entrepreneurship and solid legal advice is pure gold. As an entrepreneur, you know you need to sort out your terms, but it rarely gets done—until Startup-Recht sits down with you. They guide you through what really matters and create terms that fit your company. The perfect balance between customer-focused and legally safe. Still in doubt? Have a coffee with the guys and you’ll be convinced.
Sybrandus Pietersma
Mede-eigenaar Staallokaal B.V.
Very satisfied with Startup-Recht. They helped us draft multiple contracts and general terms and managed to translate our services and workflow perfectly into strong legal documents. Everything was clearly explained, and they even covered points we hadn’t thought of. Fast communication, clear advice, and a top result.
Daniël Coenen
Mede-oprichter Digiswift B.V.
We engaged Startup-Recht to draft our terms and conditions and service agreement. The result was delivered quickly, of high quality, and fully tailored to our needs thanks to the revision rounds. In addition, Startup-Recht provided valuable input within the context of our business.

Professional, reliable, and a pleasure to work with.
Paul Brandsma
Mede-oprichter AcuityAi

Startup-Recht assisted me in a professional and careful manner. Their work was characterized by speed, transparency, and a smooth process – all at a very reasonable rate. I consider the collaboration trustworthy and highly recommendable.

Michael de Jong
Webdeveloper & Founder
Maarten and Caylun did an excellent job helping us draft strong legal terms and meet the right compliance standards. We didn’t have much prior knowledge, but they took the time to explain everything clearly and gave valuable advice for the future. Overall, we were very well supported and would definitely recommend Startup-Recht.
Robin Jonckers
Co-founder Copywise Ai
Caylun en Maarten van Startup-Recht

Meet your modern legal partner. Work becomes easier, faster, and more secure.

Book a consultation