0 cumulative citations
View corpus contextA legal test for generative-AI copying: an output infringes only if it could not have been produced without a specific training work, and under formal modeling this rule implies a sharp divide — when organic creation is light-tailed, individual dependence fades and regulation is unlikely to constrain generation, but with heavy-tailed creation regulation can remain binding.
Citation observations
Cumulative provider counts captured on specific dates; providers are never combined.
Copyright law focuses on whether a new work is "substantially similar" to an existing one, but generative AI can closely imitate style without copying content, a capability now central to ongoing litigation. We argue that existing definitions of infringement are ill-suited to this setting and propose a new criterion: a generative AI output infringes on an existing work if it could not have been generated without that work in its training corpus. To operationalize this definition, we model generative systems as closure operators mapping a corpus of existing works to an output of new works. AI generated outputs are \emph{permissible} if they do not infringe on any existing work according to our criterion. Our results characterize structural properties of permissible generation and reveal a sharp asymptotic dichotomy: when the process of organic creations is light-tailed, dependence on individual works eventually vanishes, so that regulation imposes no limits on AI generation; with heavy-tailed creations, regulation can be persistently constraining.
Summary
Main Finding
The paper proposes a counterfactual, generative-account definition of infringement for AI outputs: an AI-generated work infringes an existing work c if that output could not have been generated without c being in the model’s training corpus. Modeling generators as closure (consequence) operators, the authors characterize the set of permissible (non-infringing) outputs and show a sharp asymptotic dichotomy: if new human creations are drawn from a light-tailed distribution, dependence on any single existing work vanishes as corpora grow and almost all generable outputs become permissible; if innovation is heavy-tailed, corpora can continue to contain essential, non-redundant works and a positive measure of outputs may remain infringing indefinitely.
Key Points
- New infringement criterion: counterfactual dependence. Output is infringing iff it is generable with the full corpus but not generable when a particular existing work is removed.
- Generators abstracted as closure operators (preservation, monotonicity, idempotence). This captures many mechanisms (interpolation, recombination, iteration).
- Example generator classes:
- Convex-hull generator (convex combinations of existing points),
- Splice generator (coordinate-wise recombination),
- Box generator (composition of splice and convex hull).
- Permissible set = outputs generable even after removing any single existing work. Violation set = generable outputs that rely essentially on at least one specific existing work.
- Structural properties:
- Permissible set grows monotonically with corpus size; closed under further generation (combining permissible outputs cannot create a violation).
- Comparative statics: adding a previously-violating work strictly expands the permissible set; adding a work that is already permissible leaves the permissible set unchanged.
- A sufficient condition for non-emptiness of the permissible set uses the Radon number (convex-geometry condition).
- Long-run asymptotics (main theorem):
- Light-tailed innovation: ratio of permissible to generable outputs → 1 almost surely as corpus grows (single works cease to be essential; redundancy builds).
- Heavy-tailed innovation: persistent essentialness; a nonzero fraction of outputs remains violations even for large corpora.
- Legal & conceptual distinction: this counterfactual notion can diverge from “substantial similarity.” An output might infringe under the counterfactual test without resembling the source, and vice versa.
Data & Methods
- Methodology: theoretical, formal model. No empirical dataset is analyzed.
- Formal ingredients:
- Representation of creations as points in R^d and corpora as measurable sets.
- Definition of generator g: a map from corpora to corpora satisfying preservation, monotonicity, idempotence (i.e., a closure operator).
- Definitions of c-permissible and c-violation sets, global permissible and violation sets as intersections/unions across corpus elements.
- Examples and constructions to illustrate intuition (convex hull, splice, box).
- Use of convex geometry (Radon number, Tukey-depth intuition for convex-hull case) to derive non-emptiness and geometric properties.
- Probabilistic asymptotic analysis: model of sequential corpus growth (new works drawn i.i.d. from a distribution with specified tail behavior) to derive almost-sure limits of permissible/generable ratios under light vs heavy tails.
- Comparative static propositions on how permissible set changes with additions to the corpus.
Implications for AI Economics
- When policy matters depends on the geometry of creative production:
- If creative output distributions are light-tailed (high redundancy), restricting outputs via the counterfactual rule would become essentially non-binding as data scale up. Regulatory limits on permissible outputs would not meaningfully constrain model outputs in saturated markets.
- If creativity is heavy-tailed (persistent novelty and a few unique works), output-based liability can remain binding indefinitely: single works retain leverage and outputs may continue to depend essentially on particular originals.
- Intellectual property design:
- The counterfactual criterion provides a principled alternative to “substantial similarity” that aligns liability with marginal dependence on training data rather than observed resemblance. This changes litigation incentives and the target of enforcement (which works are “essential” to a harmful output).
- In large-corpus settings where dependence vanishes, ex ante constraints on training (e.g., licensing) matter more for creator compensation than ex post output rules. Conversely, in heavy-tailed domains, ex post liability or licensing may remain critical to protect creators.
- Market structure and bargaining:
- Persistence of essential works (heavy tails) gives creators bargaining power and supports licensing markets/ex-post remedies; redundancy (light tails) weakens bargaining power and may encourage different business models (platform-as-aggregator, subscription).
- The results complement prior work (Gans 2024; Yang & Zhang 2025): policy choices (fair-use scope, copyrights on outputs) interact with corpus scale and innovation tails to affect welfare and incentives.
- Litigation and enforcement:
- The counterfactual test is conceptually attractive but may be practically hard to implement (requires reasoning about generative pathways, models’ architecture, and counterfactual retraining). Empirical estimation of whether an output is generable without a given work will be technically costly and litigation-intensive.
- The model highlights where enforcement resources may be most valuable: domains with heavy-tailed returns (e.g., blockbuster media, iconic visual styles) versus domains where redundancy renders enforcement low-value.
- Policy priorities and empirical program:
- Empirically estimating tail behavior of creative production becomes crucial: regulators should invest in measuring whether artistic/creative output distributions are heavy- or light-tailed in relevant domains.
- Policy interventions could be tailored: stronger ex post protections (or licensing mandates) in heavy-tailed domains; looser output restrictions and focus on market/redistribution solutions in light-tailed domains.
- Extensions and caveats:
- The model is abstract and static; welfare implications (consumer surplus, creators’ incentives, innovation) require embedding the criterion into dynamic economic models (production choices, licensing, investment in model quality).
- Practical application needs operational testing procedures to assess counterfactual generability for particular models and works.
- Strategic behavior by creators (e.g., intentionally producing “essential” works) or modelers (selective dataset curation) could alter the tail properties and thus policy outcomes.
Overall, the paper reframes infringement as a question of essential dependence in generation, links enforceability to statistical properties of human creativity, and provides a sharp guide for when output-based ownership rules will matter economically.
Assessment
Claims (6)
| Claim | Direction | Outcome | Confidence & Evidence | Details |
|---|---|---|---|---|
| Existing definitions of copyright infringement (focused on whether a new work is "substantially similar" to an existing one) are ill-suited to generative AI, because generative AI can closely imitate style without copying content. Governance And Regulation | negative | suitability of existing legal definitions for generative AI |
Reading fidelity
high
Study strength
medium
|
not reported
|
| A generative AI output should be considered infringing if and only if it could not have been generated without that specific work being present in the model's training corpus (proposed new criterion). Governance And Regulation | positive | legal criterion for infringement of AI-generated outputs |
Reading fidelity
high
Study strength
speculative
|
not reported
|
| Generative systems can be modeled as closure operators mapping a corpus of existing works to an output set of new works; under this formalization, an AI-generated output is permissible if it does not infringe on any existing work according to the proposed criterion. Governance And Regulation | positive | formal characterization of permissible AI outputs |
Reading fidelity
high
Study strength
medium
|
not reported
|
| There is a sharp asymptotic dichotomy: if the process generating organic (human) creations is light-tailed, dependence of generative outputs on any individual work eventually vanishes, implying that regulation (under the proposed criterion) ultimately imposes no limits on AI generation. Governance And Regulation | positive | extent to which regulation constrains AI generation asymptotically under light-tailed creation processes |
Reading fidelity
high
Study strength
medium
|
not reported
|
| By contrast, if the process of organic creations is heavy-tailed, dependence on individual works can persist, so regulation (under the proposed criterion) can be persistently constraining on AI generation. Governance And Regulation | negative | extent to which regulation constrains AI generation asymptotically under heavy-tailed creation processes |
Reading fidelity
high
Study strength
medium
|
not reported
|
| The capability of generative AI to imitate style without copying content is now central to ongoing litigation. Governance And Regulation | mixed | centrality of style-imitation capability in current litigation involving generative AI |
Reading fidelity
medium
Study strength
low
|
not reported
|