
AI training data is subject to copyright law. Scraped web content does not enter the public domain automatically. Learn what EU AI Act compliance, licensing, and eIDAS timestamps mean for developers building AI in 2026.
In October 2024, a Berlin-based AI startup closed a Series A and began due diligence with a Tier 1 enterprise client. The client's legal team asked one question: when was each dataset assembled, and can you prove it predates your opt-out compliance date? The startup's CTO had no qualified timestamp on any dataset. Two weeks later, the deal stalled. The same datasets were later challenged by a media publisher whose robots.txt opt-out had been live since March 2024.
As EU AI Act enforcement begins in 2026, data provenance documentation is a legal requirement for general-purpose AI model providers, and the absence of timestamps on training datasets is the single most common gap found in regulatory inspections.
Yes, in most cases. Copyright protection attaches to original works at the moment of creation without any registration requirement (Berne Convention, Art. 5). When developers scrape websites, books, code repositories, or images to build training datasets, they are creating copies of protected works. Whether that copying constitutes infringement depends on jurisdiction, licensing terms, and the applicability of exceptions such as "text and data mining" (TDM) under EU law.
In the EU, Article 4 of the Copyright in the Digital Single Market (CDSM) Directive (2019/790/EC) permits TDM for commercial purposes, but only if the rights holder has not opted out. Publishers may place a machine-readable opt-out on their content (such as robots.txt or meta tags). If they do, scraping that content for AI training is not covered by the TDM exception.
The EU AI Act (Regulation 2024/1689) imposes transparency and documentation obligations on providers of general-purpose AI (GPAI) models. Article 53 requires providers to:
For high-capability GPAI models (above the 10^25 FLOPs threshold), additional adversarial testing and incident-reporting obligations apply. Failure to document data provenance is a direct regulatory risk under the AI Act, not just a copyright risk. Non-compliance with data governance documentation is subject to fines up to EUR 15 million or 3% of global annual turnover.
You can, with conditions. The EU TDM exception (CDSM Art. 4) permits scraping for commercial AI training unless the rights holder has opted out. In Switzerland, the revised Copyright Act (URG) enacted in 2020 contains a similar research TDM exception, but its scope for commercial AI training remains contested in 2026.
Significant. In 2023 and 2024, multiple class-action lawsuits (Getty Images v. Stability AI; Doe 1 v. GitHub Copilot) progressed through initial pleadings, establishing that AI training on scraped data without consent raises actionable infringement claims. The EU AI Act adds a regulatory layer: non-compliance with data governance documentation is subject to fines up to EUR 15 million or 3% of global annual turnover.
Beyond liability, reputational risk is real. Investors and enterprise clients increasingly conduct IP due diligence on AI companies' training datasets before signing commercial agreements. A stalled deal is often more damaging than any fine.
If the Berlin startup's CTO had sealed each dataset manifest using an eIDAS qualified timestamp at the point of collection, the outcome would have been different. The enterprise client's legal team would have received a verifiable certificate showing the dataset's hash, assembly date, and the opt-out status of each source at that moment. The media publisher's March 2024 opt-out would have arrived after the timestamp, making the dataset provably compliant. The deal would have closed. A qualified timestamp under eIDAS Art. 41 cannot be backdated: it is issued by a Trust Service Provider listed on the EU Trusted Lists, audited and certified, carrying a legal presumption of time in all EU member state courts.
A developer or data team that timestamps their dataset at the point of collection, before training, creates verifiable, court-admissible proof of:
An eIDAS qualified timestamp issued by a Trust Service Provider (TSP) listed on the EU Trusted Lists carries the same legal weight as a notarized date. It cannot be backdated. This matters when a rights holder claims you scraped their content after they opted out: you can prove the dataset predates the opt-out.
Swiss Trust Layer issues eIDAS-compliant qualified timestamps on datasets, manifests, and licensing documentation in a single sealing step. The resulting certificate is verifiable by anyone without login.
If your organisation created the training data internally (human annotators, synthetic generation, original creative works), you own it, but you still face provenance challenges:
Timestamping dataset versions, including documentation of licensing agreements for each subset, creates a defensible record for due diligence, investor audits, and regulatory inspections.
| Jurisdiction | TDM Exception | AI Act Coverage | Key Risk |
|---|---|---|---|
| EU | Yes (with opt-out) | Full GPAI obligations | Opt-out compliance + documentation |
| Switzerland | Limited (research) | Voluntary alignment | Commercial TDM not clearly permitted |
| UK | Yes (non-commercial only) | No AI Act equivalent | Commercial use not covered |
| USA | Fair use (unsettled) | Executive Order only | Litigation-driven risk |
| Japan | Broad TDM exception | None | Low regulatory risk |
EU-based AI developers face the highest combined copyright and regulatory burden. Swiss developers should follow EU standards proactively given cross-border data flows.
These are legally distinct issues. AI-generated content IP protection addresses who owns the output of an AI model. Training data copyright addresses whether the input to training is legally used. Both must be assessed for a compliant AI product.
The EU AI Act data governance requirements build on both: developers must document data sourcing practices (training data) and implement safeguards against generating infringing output.
Under eIDAS Regulation (EU) 910/2014, a qualified electronic timestamp (QTS) issued by a qualified TSP:
For AI training data provenance, this means a QTS on your dataset manifest is the gold standard of documented compliance. It turns a self-assertion ("we assembled this dataset on date X") into a legally defensible fact.
AI scraping affects both developers (who scrape) and content creators (whose work is scraped). For content owners, the question is: how do you prove your content was created before it appeared in an AI model's training dataset?
The standard legal mechanism is a qualified electronic timestamp. Under eIDAS Article 41(2), a qualified timestamp creates a legal presumption that the content existed in its current form at a specific point in time. If your blog post, image, codebase, or dataset was timestamped before an AI model's training cutoff, that timestamp is evidence that the AI model could not have created it independently. Your work was the original.
Four practical steps for content owners:
X-Robots-Tag: noai headers and a disallow: / in your ai-robots.txt. As of 2026, most major AI labs respect machine-readable opt-outs under EU CDSM Directive Art. 4(3).The legal situation is changing fast. Opt-outs are not universally respected, and litigation against AI training datasets is active in multiple EU jurisdictions. A pre-existing cryptographic timestamp is the only proof mechanism that survives court scrutiny independent of the AI lab's cooperation.
The Berlin startup eventually settled the publisher dispute. Legal costs alone reached EUR 180,000 before any settlement figure, representing more than a year of runway for a seed-stage team. Sealing each dataset manifest on Swiss Trust Layer costs CHF 5. The entire dataset provenance documentation for a mid-size GPAI project runs to a few hundred francs. The cost of defending one copyright claim, or failing one enterprise due diligence review, runs to hundreds of thousands. At CHF 5 per seal versus EUR 180,000 in legal costs, the calculation is not complicated.
Protect your work with Swiss Trust Layer AG
Seal your intellectual property with a court-proof e-Seal backed by Swisscom Trust Services.
Book a Free Demo