Training data management

Training data management covers the selection, review and documentation of the data on which a model is trained or fine-tuned. For high-risk systems, this is a legal requirement: the data must be relevant, sufficiently representative and as complete as possible in view of the intended use, and the company must be able to document its origin and how it was processed. There are three practical questions – where the data comes from and whether the company is allowed to use it for this purpose, whether it contains personal data and on what legal basis, and whether it corresponds to the environment in which the system will be deployed. A model trained on data from a different market or a different era produces skewed outputs, even when the data is technically clean.

See also: AI and Company Data Quality, Bias and discrimination in AI, Data preparation as an eligible expense.