February 10, 2026

Intellectual property rights over AI training data

Legal analysis of the ownership, protection, and licensing of data used to train AI models. Implications of the AI Act and the Copyright Directive.

Intellectual Property

Who owns the data that trains AI?

Intellectual property over training data is one of the most relevant legal debates in the AI ecosystem. For startups and companies developing machine learning models, the question is not abstract: it determines what they can do with their models, how they can commercialise them, and what legal risks they assume.

Europe has taken a clear position with two complementary regulatory instruments:

Copyright Directive (2019/790), Article 4. Establishes a mandatory text and data mining (TDM) exception for anyone with lawful access to content, provided the rights holder has not expressly reserved their rights. This means you can extract data from public sources to train models, unless the owner has opted out.

AI Act, GPAI obligations. Providers of general-purpose models must implement a copyright compliance policy and publish a sufficiently detailed summary of the content used for training. This transparency obligation is new and specific to the AI Act.

Not all training data have the same legal status:

Internally created data. If your company generates the data (annotations, labelling, synthetic data), ownership is clear: it belongs to the company, protected as a database and potentially by copyright on the selection and arrangement.

Public data with copyright. Texts, images, code, music — most online content has copyright. The TDM exception allows its use for training in Europe, but with conditions: lawful access and respect for the rights holder’s opt-out.

Personal data. GDPR applies regardless of use for AI training. You need a legal basis (consent, legitimate interest) and must comply with principles of minimisation, proportionality, and data subject rights.

Third-party data under licence. Commercial datasets have specific licence terms that may restrict their use for AI training. It is essential to review each licence.

  1. Copyright infringement through web scraping. If your model was trained on internet data and a rights holder demonstrates they reserved their rights (opt-out), you could face infringement claims. Several lawsuits are pending in European jurisdictions.

  2. GDPR non-compliance. Training models with personal data without an adequate legal basis can result in sanctions. The ChatGPT-Italy case illustrates the sensitivity of European authorities.

  3. Trade secret violation. If you use datasets containing confidential third-party information (even inadvertently), you could face claims under the Trade Secrets Directive.

  4. AI Act transparency obligations. Failing to publish the training content summary or implement a copyright policy can result in sanctions under the AI Act for GPAI providers.

Practical training data protection strategy

For your own data:

  • Document the creation and collection process
  • Register significant databases
  • Implement access controls
  • Include ownership clauses in contracts with annotators and data providers

For third-party data:

  • Audit the sources and licences of all datasets
  • Document compliance with the TDM exception
  • Implement opt-out detection mechanisms
  • Maintain data provenance records to comply with the AI Act

For personal data:

  • Conduct a Data Protection Impact Assessment (DPIA)
  • Establish the most appropriate legal basis
  • Implement anonymisation or pseudonymisation where possible
  • Document compliance for audits

The trained model: derivative work or new creation?

An open question is whether a trained model constitutes a derivative work of the training data. If it were, data rights holders could claim rights over the model. The majority view in Europe tends to consider that the model is not a derivative work if it does not directly reproduce the data, but case law is evolving.

For startups, the practical recommendation is clear: document your training process, ensure the legality of sources, and maintain records demonstrating compliance. At A2 we help AI companies structure their training data strategy with legal certainty. Consult with our team.

Contact us