Copyright and model training are the subject both of disputes and of regulation that is being progressively filled in. European law permits automated analysis of texts and data, including protected works, provided lawful access to them exists and the rightholder has not expressly reserved their use – for content on a website, this reservation is applied in a machine-readable way. Providers of general-purpose models also have an obligation to publish a sufficiently detailed summary of the content used for training. Two practical points follow for businesses. As the owner of a website, they can reserve the use of their content, although this also reduces their visibility in generative tools. As a user of a tool, they should know the licence terms attached to the outputs created. The decision on reserving rights is therefore a business decision, not merely a technical website setting.
See also: Meta rules for AI bots, Paywall and AI access, Intellectual Property Rights to an AI Solution.