Skip to content
Menu

Resources

Documenting training data for the EU AI Act

What the training-summary template asks and what a dataset vendor should hand over.

What the obligation actually is

Providers of general-purpose AI models in the European Union must publish a sufficiently detailed summary of the content used for training, following a template issued by the AI Office. It is a transparency obligation rather than a licensing one: it asks you to describe your data, not to prove you owned all of it.

In practice that turns your vendors into part of your compliance surface, because you cannot describe what they will not tell you.

What the summary asks about a purchased dataset

  • Who the provider is and what the dataset is called, at a specific version.
  • The nature and origin of the content, in enough detail for a reader to understand what kind of material it is.
  • Approximate scale, and the main types and modalities it contains.
  • The basis on which it was licensed for training.
  • Whether it contains personal data, and how that was handled.

What to require from a vendor

A per-asset provenance manifest rather than a category description. A machine-readable dataset manifest, for which Croissant is the emerging convention. A written statement of origin you can quote directly. A version identifier and checksums, so the thing you documented is the thing you trained on. And a named contact who will answer a compliance question during the licence term.

Ask for these before signing. Afterwards, you are asking a supplier for a favour rather than exercising a term.

Why placeholder content simplifies this

A corpus written entirely with invented organisations, invented figures and no real people has a short answer to the personal-data question, and a short answer to where the content came from. That is worth as much at audit time as it is at training time.

Read next