Will the European Commission make AI innovation significantly harder and more expensive by mandating ineffective copyright opt-out mechanisms for AI training? The question is technical, but the answer is critical to Europe’s innovation potential.
Disruptive Competition Project - Breaking News on Breaking Stuff
Main takeaways
Will the European Commission make AI innovation significantly harder and more expensive by mandating ineffective copyright opt-out mechanisms for AI training? The question is technical, but the answer is critical to Europe’s innovation potential.
First, some context. The AI Act imposes copyright rules on providers of general-purpose AI (GPAI) models, including an obligation to identify and comply with reservations expressed by rightsholders under Article 4 of the EU Copyright Directive (EUCD). Article 4 creates a text-and-data-mining (TDM) exception, which enables AI training on publicly available content unless a rightsholder objects in a “machine-readable” format, known as an ‘opt-out’.
Article 4 intentionally uses broad wording to avoid locking the law into rigid technical protocols or picking winners, allowing markets and experts to develop technical solutions and standards. This AI Act obligation was recently translated into a voluntary Code of Practice for GPAI providers, under which signatories commit to read and follow instructions expressed through the robust and universal Robots.txt Protocol.
Signatories also commit to identify and comply with other protocols adopted by standardisation organisations, as well as state-of-the-art, technically implementable, widely adopted protocols that have been “generally agreed” through inclusive EU-level discussions.
1. The Commission’s unbalanced processIn this context, the European Commission asked stakeholders in early 2026 for their views on opt-out solutions that could be included in a list of generally agreed mechanisms. The Commission expected to reach an agreement after two stakeholder workshops – an extremely ambitious timeline. As anticipated, stakeholders across the board expressed disagreement at the first workshop in June, which led the Commission to delay the second.
Surprisingly, the Commission excluded trade associations representing technology companies from the workshops, leading to significant underrepresentation of AI developers. The process also raises other concerns. The Commission, which is not a standardisation body and is therefore only empowered to facilitate discussions, should not exceed its mandate by unilaterally imposing specific technical standards.
None of the seven opt-out solutions explored so far is fit for purpose, and two receiving particular attention from the Commission would be entirely unworkable: the TDM Reservation Protocol (TDMRep) and the Creator Assertions Working Group Protocol (CAWG).
2. TDMRep: A redundant and technically unworkable protocolDeveloped by French publishers via a W3C community group, rather than a formal standards body, TDMRep is not a W3C standard and lacks balanced multistakeholder governance. It introduces TDM-reservation (a binary opt-out flag) and TDM-policy (a JSON file using ODRL to express terms like a “duty to compensate”).
In practice, TDMRep’s location-based controls offer no additional functionality over Robots.txt. Worse, conflicting instructions (such as ’TDM-reservation: 1’ – a binary exclusion – alongside ‘GPT-Bot: allow’) create severe legal uncertainty for web crawlers and rightsholders alike.
Furthermore, TDM-policy files attempt to create a veneer of machine-readability by embedding legal terms of service (ToS) in JSON files. But these preferences are not structured in a way that allows them to be translated into actionable instructions at web scale. That is why such files should be excluded from any mandatory implementation.
Controlled by a narrow group of publishers with vested interests, mandating TDMRep leaves developers vulnerable to arbitrary future requirements without technical oversight.
3. CAWG: Conflating provenance and copyright, compromising privacyThe CAWG protocol attempts to attach rights-reservation signals to individual digital assets – such as an image – using metadata layered onto the Coalition for Content Provenance and Authenticity (C2PA) specification. But C2PA is designed to record information about an asset’s provenance – where it came from and how it was modified – not to determine legal copyright ownership. A C2PA manifest can contain a camera or software signature, but that does not by itself verify who owns the underlying copyright (for example, a photo agency).
This conflation creates severe privacy and safety risks by attaching personal identity data directly to digital assets. C2PA deliberately excluded identity features because human rights groups warned that this would threaten journalists, whistleblowers, and dissidents. Forcing copyright enforcement into provenance tools risks fracturing the ecosystem and stalling C2PA adoption.
More generally, asset-based metadata mechanisms fail at scale. Embedded metadata lacks verification, inviting fraud as anyone can alter it. Standard web platforms routinely strip metadata during file compression, meaning persistence would require a disruptive overhaul of web architecture. Finally, static metadata snapshots cannot capture dynamic, transferable copyright – resulting in conflicting signals across duplicate files.
4. Rely on consensus-driven international technical standardsAs it stands, no solution other than the Robots.txt Protocol offers the maturity, scalability, and technical robustness needed to express opt-outs. The Commission’s process should not mandate ineffective solutions, which would only raise the cost of innovation in the EU.
Rather than introducing fragmented, overlapping, or legally mandated opt-out mechanisms that risk making the TDM framework unworkable, the EU should rely on industry-led, international consensus standards. Specifically, the EU should support ongoing work within international bodies, such as the Internet Engineering Task Force’s (IETF) AI Preferences Working Group, to update widely adopted protocols like Robots.txt, enabling more effective and granular rights reservations.
Establishing a standardised, machine-readable vocabulary at global level ensures that rights reservations can be processed automatically and reliably at scale across jurisdictions, which also benefits rightsholders. This is key to avoiding conflicting signals or regional protocols that would defeat the purpose of rights reservations.
ConclusionThe success of the EU’s framework for AI and copyright hinges on robust, universally agreed, and market-led standards. Instead of mandating inappropriate technical tools, the European Commission should support the ongoing work of standardisation bodies with the relevant expertise and ability to foster market adoption.
| # | Наименование новости | Тональность | Информативность | Дата публикации |
|---|---|---|---|---|
| 1 | OpenAI 'trained models' on books from 'sketchy AF' LibGen, unsealed court filings reveal | 0 | 9.57 | 23-09-2026 |
| 2 | Court Filing Shows Microsoft Exec Called AI Training the "Largest Labor Theft in Human History" | 0 | 5.73 | 20-09-2026 |
| 3 | Empleados de OpenAI y Microsoft temen el impacto de la IA en el periodismo: "El mayor robo de trabajo de la historia de la humanidad" | 0 | 7.13 | 18-09-2026 |
| 4 | El Gobierno afirma que está en «permanente contacto» con el sector cultural ante el avance de la IA | 0 | 6.82 | 27-09-2026 |
| 5 | Microsoft drafts ‘humanist’ AI code of conduct amid safety debate | 0 | 6.96 | 15-09-2026 |
| 6 | Canada: Ottawa Advances AI Strategy on Two Fronts | 0 | 10 | 16-09-2026 |
| 7 | AI giants probing tens of thousands of security incidents – Axios | 0 | 9.83 | 27-09-2026 |
| 8 | US and China establish ‘AI incidents’ hotline | 0 | 12.86 | 26-09-2026 |
| 9 | OpenAI pauses most powerful AI training after thousands of sandbox escapes uncovered | 0 | 9.94 | 27-09-2026 |
| 10 | Treat Spotify Like a Public Utility | 0 | 13.24 | 24-09-2026 |