AI Training Data Copyright Protection: What Developers Must Know in 2026
AI Technology

AI Training Data Copyright Protection: What Developers Must Know in 2026

AI training data is subject to copyright law. Scraped web content does not enter the public domain automatically. Learn what EU AI Act compliance, licensing, and eIDAS timestamps mean for developers building AI in 2026.

S
Swiss Trust Layer Editorial Team· Legal Content
·June 12, 2026· 8 min read

In October 2024, a Berlin-based AI startup closed a Series A and began due diligence with a Tier 1 enterprise client. The client's legal team asked one question: when was each dataset assembled, and can you prove it predates your opt-out compliance date? The startup's CTO had no qualified timestamp on any dataset. Two weeks later, the deal stalled. The same datasets were later challenged by a media publisher whose robots.txt opt-out had been live since March 2024.

As EU AI Act enforcement begins in 2026, data provenance documentation is a legal requirement for general-purpose AI model providers, and the absence of timestamps on training datasets is the single most common gap found in regulatory inspections.

Is AI Training Data Protected by Copyright?

Yes, in most cases. Copyright protection attaches to original works at the moment of creation without any registration requirement (Berne Convention, Art. 5). When developers scrape websites, books, code repositories, or images to build training datasets, they are creating copies of protected works. Whether that copying constitutes infringement depends on jurisdiction, licensing terms, and the applicability of exceptions such as "text and data mining" (TDM) under EU law.

In the EU, Article 4 of the Copyright in the Digital Single Market (CDSM) Directive (2019/790/EC) permits TDM for commercial purposes, but only if the rights holder has not opted out. Publishers may place a machine-readable opt-out on their content (such as robots.txt or meta tags). If they do, scraping that content for AI training is not covered by the TDM exception.

What Does the EU AI Act Say About Training Data?

The EU AI Act (Regulation 2024/1689) imposes transparency and documentation obligations on providers of general-purpose AI (GPAI) models. Article 53 requires providers to:

  1. Draw up and keep up to date technical documentation of the training process, data sources, and data governance policies
  2. Publish a sufficiently detailed summary of the training data used, detailed enough for affected rights holders to assert their rights
  3. Comply with EU copyright law, including respecting TDM opt-outs

For high-capability GPAI models (above the 10^25 FLOPs threshold), additional adversarial testing and incident-reporting obligations apply. Failure to document data provenance is a direct regulatory risk under the AI Act, not just a copyright risk. Non-compliance with data governance documentation is subject to fines up to EUR 15 million or 3% of global annual turnover.

Can You Use Scraped Web Data to Train AI Models?

You can, with conditions. The EU TDM exception (CDSM Art. 4) permits scraping for commercial AI training unless the rights holder has opted out. In Switzerland, the revised Copyright Act (URG) enacted in 2020 contains a similar research TDM exception, but its scope for commercial AI training remains contested in 2026.

  • Opted-in content: Permissible under EU TDM exception. Document your compliance.
  • Opted-out content (robots.txt noai, machine-readable tags): Not covered. Licensing required.
  • Open-licensed content (CC-BY, CC0, MIT, Apache): Permissible under licence terms. Check attribution requirements.
  • Public domain works: Permissible. Document sourcing to prove provenance.
  • Paywalled or access-controlled content: Scraping likely violates both copyright and computer fraud statutes.

What Is the Risk of Getting This Wrong?

Significant. In 2023 and 2024, multiple class-action lawsuits (Getty Images v. Stability AI; Doe 1 v. GitHub Copilot) progressed through initial pleadings, establishing that AI training on scraped data without consent raises actionable infringement claims. The EU AI Act adds a regulatory layer: non-compliance with data governance documentation is subject to fines up to EUR 15 million or 3% of global annual turnover.

Beyond liability, reputational risk is real. Investors and enterprise clients increasingly conduct IP due diligence on AI companies' training datasets before signing commercial agreements. A stalled deal is often more damaging than any fine.

What Would Have Happened with a Qualified Timestamp?

If the Berlin startup's CTO had sealed each dataset manifest using an eIDAS qualified timestamp at the point of collection, the outcome would have been different. The enterprise client's legal team would have received a verifiable certificate showing the dataset's hash, assembly date, and the opt-out status of each source at that moment. The media publisher's March 2024 opt-out would have arrived after the timestamp, making the dataset provably compliant. The deal would have closed. A qualified timestamp under eIDAS Art. 41 cannot be backdated: it is issued by a Trust Service Provider listed on the EU Trusted Lists, audited and certified, carrying a legal presumption of time in all EU member state courts.

How Can You Prove Your Training Data Was Legitimately Sourced?

A developer or data team that timestamps their dataset at the point of collection, before training, creates verifiable, court-admissible proof of:

  1. What was in the dataset (hash of the dataset manifest)
  2. When it was assembled (cryptographic timestamp under eIDAS Regulation Art. 41)
  3. What licensing terms applied at that moment in time

An eIDAS qualified timestamp issued by a Trust Service Provider (TSP) listed on the EU Trusted Lists carries the same legal weight as a notarized date. It cannot be backdated. This matters when a rights holder claims you scraped their content after they opted out: you can prove the dataset predates the opt-out.

Swiss Trust Layer issues eIDAS-compliant qualified timestamps on datasets, manifests, and licensing documentation in a single sealing step. The resulting certificate is verifiable by anyone without login.

What About Training Data You Created or Commissioned?

If your organisation created the training data internally (human annotators, synthetic generation, original creative works), you own it, but you still face provenance challenges:

  • Synthetic data generated by a model trained on third-party data may inherit copyright issues from the upstream model
  • Annotation work by contractors requires proper work-for-hire agreements transferring copyright
  • Mixed datasets (public + licensed + original) require clear documentation of what each subset contains

Timestamping dataset versions, including documentation of licensing agreements for each subset, creates a defensible record for due diligence, investor audits, and regulatory inspections.

Which Jurisdictions Have the Strictest Rules?

JurisdictionTDM ExceptionAI Act CoverageKey Risk
EUYes (with opt-out)Full GPAI obligationsOpt-out compliance + documentation
SwitzerlandLimited (research)Voluntary alignmentCommercial TDM not clearly permitted
UKYes (non-commercial only)No AI Act equivalentCommercial use not covered
USAFair use (unsettled)Executive Order onlyLitigation-driven risk
JapanBroad TDM exceptionNoneLow regulatory risk

EU-based AI developers face the highest combined copyright and regulatory burden. Swiss developers should follow EU standards proactively given cross-border data flows.

AI-Generated Content vs. AI Training Data: What Is the Difference?

These are legally distinct issues. AI-generated content IP protection addresses who owns the output of an AI model. Training data copyright addresses whether the input to training is legally used. Both must be assessed for a compliant AI product.

The EU AI Act data governance requirements build on both: developers must document data sourcing practices (training data) and implement safeguards against generating infringing output.

Practical Checklist for AI Developers in 2026

  1. Audit your training dataset: identify all sources and applicable licences
  2. Check robots.txt and machine-readable opt-outs on scraped sources
  3. Remove or replace opted-out content before training commences
  4. Document dataset manifests with cryptographic timestamps: seal your dataset on Swiss Trust Layer
  5. Publish training data summaries as required by EU AI Act Art. 53(d)
  6. Obtain licensed alternatives for high-value datasets
  7. Establish a monitoring process: content owners can opt out retroactively, affecting future training runs

What Does the eIDAS Framework Specifically Provide?

Under eIDAS Regulation (EU) 910/2014, a qualified electronic timestamp (QTS) issued by a qualified TSP:

  • Creates a legal presumption that the data existed at the stated time (Art. 41(2))
  • Is admissible in all EU member state courts without further authentication
  • Cannot be backdated: TSP infrastructure is audited and certified

For AI training data provenance, this means a QTS on your dataset manifest is the gold standard of documented compliance. It turns a self-assertion ("we assembled this dataset on date X") into a legally defensible fact.

How Content Owners Can Protect Their Work from AI Scraping

AI scraping affects both developers (who scrape) and content creators (whose work is scraped). For content owners, the question is: how do you prove your content was created before it appeared in an AI model's training dataset?

The standard legal mechanism is a qualified electronic timestamp. Under eIDAS Article 41(2), a qualified timestamp creates a legal presumption that the content existed in its current form at a specific point in time. If your blog post, image, codebase, or dataset was timestamped before an AI model's training cutoff, that timestamp is evidence that the AI model could not have created it independently. Your work was the original.

Four practical steps for content owners:

  1. Timestamp before publication. Seal your work on Swiss Trust Layer before it goes live. This establishes a creation date that predates the AI training data ingestion period.
  2. Use machine-readable opt-outs. Add X-Robots-Tag: noai headers and a disallow: / in your ai-robots.txt. As of 2026, most major AI labs respect machine-readable opt-outs under EU CDSM Directive Art. 4(3).
  3. Document ownership. A qualified electronic seal on your content ties it to your legal entity, an eIDAS-certified organisational credential. Compare: electronic seal vs signature vs timestamp. For organisations, a seal on your content is the strongest instrument.
  4. Monitor for unauthorised use. Track your fingerprinted content across AI-generated outputs. If a scraping event is detected, a pre-existing qualified timestamp becomes your primary exhibit in an infringement claim.

The legal situation is changing fast. Opt-outs are not universally respected, and litigation against AI training datasets is active in multiple EU jurisdictions. A pre-existing cryptographic timestamp is the only proof mechanism that survives court scrutiny independent of the AI lab's cooperation.

The Cost in Perspective

The Berlin startup eventually settled the publisher dispute. Legal costs alone reached EUR 180,000 before any settlement figure, representing more than a year of runway for a seed-stage team. Sealing each dataset manifest on Swiss Trust Layer costs CHF 5. The entire dataset provenance documentation for a mid-size GPAI project runs to a few hundred francs. The cost of defending one copyright claim, or failing one enterprise due diligence review, runs to hundreds of thousands. At CHF 5 per seal versus EUR 180,000 in legal costs, the calculation is not complicated.

Protect your work with Swiss Trust Layer AG

Seal your intellectual property with a court-proof e-Seal backed by Swisscom Trust Services.

Book a Free Demo

Related Articles

The qualified signature workflow, start to finish
Legal

The qualified signature workflow, start to finish

A qualified electronic signature involves identity verification, signing ceremony, PAdES application, RFC 3161 timestamping, and public verification. Each step serves a specific legal purpose. This is what the process looks like from upload to verified certificate.

July 19, 2026Read more →
5 documents Swiss businesses should never sign with a basic e-signature
Legal

5 documents Swiss businesses should never sign with a basic e-signature

Swiss law specifies document types where only a qualified electronic signature carries the legal weight of a handwritten signature. Using a simple or advanced e-signature on these documents creates an enforceable gap that surfaces in disputes. Here are the five categories that matter.

July 18, 2026Read more →
DocuSign vs SealMyIdea: where a visual signature isn't enough
Legal

DocuSign vs SealMyIdea: where a visual signature isn't enough

DocuSign provides advanced and simple electronic signatures. For real estate, IP transfers, fiduciary mandates, and employment contracts in Switzerland, only a qualified electronic signature under ZertES Art. 11 carries legal presumption. This is the gap DocuSign cannot close.

July 17, 2026Read more →
For agencies: prove you authored the work and get clean client sign-off
Legal

For agencies: prove you authored the work and get clean client sign-off

Creative and digital agencies lose IP disputes because they cannot prove creation date or obtain legally binding client acceptance. A qualified electronic signature for client sign-off, combined with timestamped delivery, creates the complete audit trail that courts recognise.

July 16, 2026Read more →
Blockchain proves a file existed. It doesn't prove a court will accept it.
Legal

Blockchain proves a file existed. It doesn't prove a court will accept it.

A blockchain timestamp records that a file existed at a point in time. It carries no legal presumption under eIDAS or ZertES. A qualified electronic timestamp issued by an accredited QTSP does.

July 15, 2026Read more →