Roundtable 36: Generative AI and the Challenges of Cross-Border Governance

0
1

Key Takeaways

  • Under the Berne Convention, storing a copyrighted work to train an AI model constitutes a reproduction, and any national exception must satisfy the three‑step test (certain special case, no conflict with normal exploitation, no unreasonable prejudice to authors).
  • The United States’ broad fair‑use doctrine and Japan’s “non‑enjoyment” exception fail the Berne test because they allow large‑scale commercial AI training without adequate consent or compensation, undermining emerging licensing markets.
  • The European Union’s DSM Directive comes closest to compliance by offering an opt‑out mechanism for text‑and‑data mining, but its effectiveness is limited by authors’ technical ability to reserve rights.
  • A new WIPO special agreement could clarify how the three‑step test applies to AI training, establish minimum standards for access, rights reservations, and compensation, and bridge the practical gap in the EU opt‑out system—though U.S. resistance remains a major obstacle.
  • The GDPR provides binding extraterritorial reach for EU‑related data processing, yet enforcement against offshore AI developers is hampered by detection difficulties and lack of EU assets; the voluntary CBPR framework lacks statutory fining or injunctive power, leaving individuals unprotected.
  • Mass web scraping for AI training conflicts with GDPR’s consent and legitimate‑interest bases, and the right to erasure raises unresolved questions about “machine unlearning” of already‑trained models.
  • The EU AI Act adopts a risk‑based classification that does not incorporate the precautionary principle adequately, leaving society‑wide, emergent harms (e.g., cognitive atrophy, systemic bias) insufficiently addressed.
  • Compared with EU GMO and heritable human genome editing regimes, the AI Act’s focus on pre‑market classification and economic incentives creates a structural ceiling for fundamental‑rights protections.
  • Effective AI governance will require hard‑law baselines—such as mutual‑recognition agreements, expanded GDPR adequacy decisions, and privacy provisions in trade agreements—that impose enforceable obligations on AI developers regardless of voluntary participation, coupled with rights‑based impact assessments and sector‑specific rules to curb novel societal harms.

Overview of AI training and reproduction right under the Berne Convention
When an AI model stores a copyrighted work as part of its training process, “a reproduction has occurred under Article 9 of the Berne Convention,” to which more than 180 countries are party. The convention grants authors “the exclusive right of authorizing the reproduction of these works, in any manner or form,” but Article 9(2) permits member states to create national exceptions provided the reproduction “does not conflict with a normal exploitation of the work and does not unreasonably prejudice the legitimate interests of the author.” The Agreed Statements concerning the WIPO Copyright Treaty clarify that “storage of a protected work in an electronic medium constitutes reproduction,” meaning the legal question is not whether the output resembles the original, but whether the input process—storing copyrighted works for training—is authorized.

Application of Berne Article 9(2) three‑step test and U.S. fair use
In the United States, courts have applied the fair‑use doctrine to AI training, as seen in Bartz v. Anthropic (2025), where a federal court held that Anthropic’s use of authors’ books was “exceedingly transformative.” Fair use evaluates four factors: purpose and character of the use, nature of the copyrighted work, amount copied, and market effect. However, the court “did not consider whether its interpretation of fair use satisfies the Berne Convention’s three‑step test.” Because fair use is not a “certain special case” but a broad, flexible standard, its application to mass commercial AI training fails the first step of the Berne test. Moreover, allowing free use “undermines an emerging market that constitutes foreseeable normal exploitation,” violating step two, and the resulting systemic harm to authors’ legitimate interests breaches step three. Thus, the current U.S. approach conflicts with its Berne obligations.

Japan’s Article 30‑4 approach and its shortcomings
Japan permits AI training under Article 30‑4 of its Copyright Act when the purpose is “non‑enjoyment,” meaning the work is analyzed purely as data rather than enjoyed for its expression. While this distinction attempts to satisfy the first step of the Berne test, “most AI training can be classified as non‑enjoyment,” so the exception is not truly limited in scope. Like the U.S. fair‑use regime, Japan’s law does not guarantee prior consent or a general remuneration right, leaving authors to prove individualized harm—a “structural burden‑of‑proof problem” because the injury is diffuse across millions of works. Consequently, Japan’s framework still fails the second and third steps of the Berne test.

EU DSM Directive opt‑out mechanism and practical gaps
The European Union’s Digital Single Market (DSM) Directive allows text and data mining for any purpose on lawfully accessible works, but rightsholders may opt out, stating that their work cannot be used for data mining. The directive “falls within ‘certain special cases’ because it specifically applies to text and data mining rather than serving as a general‑purpose exception.” The opt‑out protects normal exploitation by enabling rightsholders to decide whether their works are used for AI training, thereby bolstering the licensing market. Yet, “the opt‑out system is only effective if authors know where their works are being trained and can meet the technical requirements, such as machine‑readable rights reservations.” Many individual authors, especially those in developing countries or outside the tech sector, lack the resources to implement these reservations, rendering the protection formally available but practically inaccessible.

Need for a new WIPO special agreement and obstacles
Given that Article 9(2) was not drafted with generative AI in mind, the article suggests that “Participating members of Berne could adopt another WIPO special agreement clarifying how the three‑step test applies to AI training.” Such an agreement should recognize AI‑training licensing markets as part of normal exploitation, set minimum standards for access, rights reservations, and compensation, and borrow from the EU opt‑out framework while addressing its practical weaknesses—e.g., ensuring authors are aware of use and can exercise rights without specialized technical capacity. The primary obstacle is the United States, home to the world’s leading AI firms, which has a history of resisting international copyright obligations that conflict with its domestic fair‑use doctrine. Without U.S. participation, any new agreement would govern AI training everywhere except where most training occurs, limiting its effectiveness.

Cross‑border data flows under GDPR vs CBPR – extraterritorial reach
Section II contrasts the EU’s legally binding General Data Protection Regulation (GDPR) with the voluntary APEC/Global Cross‑Border Privacy Rules (CBPR). The GDPR’s Article 3(2) extends its scope to non‑EU processors when their data processing relates to “offering goods or services to data subjects within the EU” or “monitoring their behavior within the EU.” For generative AI, jurisdiction hinges on whether model training involves intentional tracking of EU individuals; indiscriminate web scraping often falls outside this reach, creating a loophole. Moreover, “European supervisory authorities generally lack the technical capacity to trace and attribute specific instances of web scraping to specific overseas developers,” meaning GDPR may have jurisdiction on paper but limited practical enforcement.

Enforcement challenges of GDPR and limitations of CBPR
When the GDPR identifies a violation, it possesses statutory tools under Article 58—including binding suspension orders, representative mandates, and fines enforceable through international legal cooperation. The CBPR framework, by contrast, relies on corporate self‑assessment verified by independent accountability agents that “lack independent statutory fining authority or injunctive powers.” Their remedies are limited to revoking certification or referring cases to domestic regulators like the U.S. FTC, whose reach does not extend to foreign AI developers without a U.S. presence. Consequently, “the CBPR has no equivalent statutory recourse even when clear violations are identified, leaving individuals whose data has been harvested defenseless outside of the EU or jurisdictions with equivalent hard‑law protections.”

Operational tensions: consent, legitimate interest, machine unlearning
Mass web scraping for AI training clashes with GDPR’s requirement that data processing rest on valid consent or a “legitimate interest” under Article 6(1)(f). In 2023, Italy’s Garante temporarily blocked ChatGPT, finding that OpenAI lacked an adequate statutory basis to collect personal data for model training. After OpenAI added age verification, updated its privacy policy, and introduced a data‑opt‑out mechanism, access was restored. This episode shows the GDPR’s power to suspend operations but also its tendency to yield procedural concessions rather than structural changes to data collection practices. Regarding erasure, Articles 16 and 17 raise the unresolved question of whether “erasure” necessitates purging raw training datasets or undergoing “machine unlearning” to strip the influence of specific data from an already‑trained model—a technical challenge that even hard‑law jurisdictions may struggle to enforce.

International data transfers and Schrems II implications for AI
The Court of Justice of the EU’s Schrems II decision holds that personal data cannot be transferred to third countries unless protections “essentially equivalent to EU law” are guaranteed. Generative AI complicates this because distributed cloud computing fragments training data across global servers, and the models themselves embody cross‑border data. An open question is whether each server‑to‑server movement constitutes a separate “transfer” requiring its own adequacy assessment, or whether the entire training process is treated as a single operation. If each transfer needs independent verification, compliance at AI‑training scale becomes practically impossible; if treated as one operation, a single adequacy assessment must account for every jurisdiction the data touches, creating a substantial compliance burden.

EU AI Act and precautionary principle – missing integration
Section III argues that the EU AI Act fails to properly integrate the precautionary principle, which the Court of Justice of the EU has recognized as a general principle of EU law requiring action “to prevent specific potential risks to public health, safety and the environment” even amid scientific uncertainty. The AI Act’s risk‑classification system relies on predetermined categories and a simple harm‑probability‑times‑severity formula, assessing only known, measurable risks and neglecting emergent, society‑wide harms such as “cognitive atrophy” among younger generations or systemic bias in AI‑driven hiring. As a result, the Act “does not meet this goal of robust fundamental rights protections” and instead offers a veneer of compliance that prioritizes economic interests over rights.

Comparison with GMO/HHGE regulation and structural limits
The precautionary principle is applied more strongly in EU regulation of heritable human genome editing (HHGE), which operates under a blanket moratorium allowing research but not market release, giving regulators time to evaluate costs and benefits. Even the EU’s GMO regime—requiring explicit permission for organisms that cannot occur through traditional breeding and permitting stricter national bans—offers more comprehensive precautionary safeguards than the AI Act. The AI Act’s grounding in Article 114 of the TFEU, which focuses on harmonized product safety, creates a “structural ceiling for fundamental‑rights protections” because it treats risk as a sliding scale rather than a binary rights‑based assessment. This market‑oriented approach leads other jurisdictions, such as Brazil and South Korea, to emulate the EU model despite possessing broader constitutional authority to protect fundamental rights, thereby exporting the EU’s limitations.

Consequences and recommendations for rights‑based AI governance
Without a rights‑based, precautionary framework, AI development risks exacerbating inequality, misinformation, and environmental harm while sidelining creators’ remuneration and individuals’ privacy. The article concludes that effective governance requires hard‑law baselines: mutual‑recognition agreements with strengthened statutory interoperability, expanded GDPR adequacy decisions covering compatible jurisdictions, and privacy provisions embedded in trade agreements. These mechanisms would create enforceable obligations that apply to AI developers regardless of voluntary participation. Complementary steps include mandating fundamental‑rights impact assessments for all AI systems, developing sector‑specific rules to address novel ethical implications, and supporting technical solutions for machine unlearning and transparent data provenance. Only through such binding, rights‑centric approaches can governments worldwide prioritize personal data, privacy, and creators’ interests over unfettered corporate control in the age of generative AI.

https://www.culawreview.org/roundtable-1/roundtable-36-scraping-by-generative-ai-and-the-limits-of-cross-border-governance

SignUpSignUp form

LEAVE A REPLY

Please enter your comment!
Please enter your name here