AI training data disclosure

Full Title:
CLEAR Act

Summary#

This bill (the CLEAR Act) would require people who use training datasets to build or release generative artificial intelligence (AI) models to file a notice with the Register of Copyrights listing copyrighted works used in that dataset. The notice must include a detailed summary of each copyrighted work and, if the dataset is publicly available, the dataset’s URL. The bill creates a public online database of those notices, gives copyright owners a private right to sue for failures to file, and sets civil penalties and other remedies.

  • Main change: Anyone who uses a training dataset to train or release a generative AI model must file a notice with the Copyright Office describing copyrighted works in the dataset.
  • Timing: For models first used or released on/after the law’s effective date, the notice must be filed at least 30 days before commercial use or release (commercial use includes internal organizational use). For models already in use before the law takes effect, notices must be filed within 30 days after the Register issues implementing rules. The law starts 180 days after enactment.
  • Enforcement: Copyright owners can sue users who fail to file. Courts may impose a civil penalty of at least $5,000 per failure, order injunctions stopping the use of the work until a notice is filed, and award attorney’s fees. There is a cap of $2.5 million per person per year on total civil penalties.
  • Public database: The Copyright Office must create and maintain a publicly available online database of all filed notices.
  • Regulations: The Register must issue rules within 180 days after the law takes effect to set the notice form and procedures.

What it means for you#

  • AI companies and developers: You must prepare and file a notice that lists and summarizes every copyrighted work in any training dataset you use to train or release a generative AI model. If the dataset is publicly available online, you must provide its URL. You must file the notice at least 30 days before commercial use or release. Failing to file can lead to lawsuits, penalties, injunctions, and legal fees.
  • Businesses that use AI internally: Internal commercial uses count. If your organization uses a model internally, you must still file the notice 30 days before that internal commercial use.
  • Copyright owners (creators and rights holders): You gain a private right to sue users who fail to file required notices. If you win, the court can stop the infringing use and award you attorney’s fees.
  • Members of the public / researchers: A public database will list notices about what copyrighted works were used in training datasets, which could increase transparency about model training sources.
  • Copyright Office (Register): Must create filing forms and rules and run a public online database. The Register will receive civil penalties paid to offset operating costs.
  • Start date: The law takes effect 180 days after it is enacted. Models already in use before that date need notices after the Register issues regulations (within the schedule set by the law).

Expenses#

No publicly available information.

  • The bill directs civil penalties paid by violators to the Copyright Office to offset its operating costs.
  • The Copyright Office must create and maintain a public online database and write implementing rules; those actions would have administrative costs (staff, IT). The bill does not include a formal cost estimate.
  • Companies will face compliance costs: preparing “sufficiently detailed” summaries, tracking dataset contents and URLs, legal review, and possibly delays in releases.
  • Copyright owners may incur litigation costs to enforce the filing requirement; successful plaintiffs can recover attorney’s fees.

Proponents' View#

  • The bill appears intended to increase transparency about what copyrighted works are used to train generative AI models.
  • This could give creators clearer notice when their registered works are used in training and a means to seek remedies if notices are not filed.
  • A public database could make it easier for the public, researchers, and policymakers to see common training sources and monitor practices.
  • Penalties paid to the Copyright Office could help defray the Office’s costs for running the database and processing notices.
  • Requiring pre-release or pre-commercial-use notice could encourage better record-keeping and accountability by model builders.

Opponents' View#

  • One concern is administrative burden and cost: companies must prepare detailed summaries for every registered copyrighted work in a dataset, which could be time-consuming and expensive.
  • The bill does not define “sufficiently detailed summary,” so compliance standards are unclear until the Register issues rules. This creates uncertainty for filers.
  • The filing deadline (30 days before commercial use or release) could delay product launches or impose operational friction, especially for iterative model development.
  • Enforcement depends on copyright owners filing suit, which could lead to many private lawsuits and added legal risk for model builders.
  • The statutory definition of “copyrighted work” covers only works that are registered (or scheduled under a specified code section), so unregistered works used in training would not be covered; this creates an uneven scope.
  • Publicizing dataset URLs and summaries may raise concerns about exposing proprietary datasets or trade secrets; the bill does not specify protections for confidential information.
  • It is unclear how courts will apply the per-instance $5,000 minimum penalty in complex cases (for example, where many works are involved), and how the annual $2.5 million cap will affect deterrence and remedies.