Public data source transparency
2026-08-31 | The summaries of public training data are encouraged to have a lot of information on public data sources used. However, we find that in general the information provided is lacking.
Article 53(1)(d) of the EU AI Act requires providers to disclose a public summary of the data used to train their models, following a provided template. On this site, we work to evaluate such summaries in order to determine to what extent they meet the requirements outlined in the template. In this blog post, we summarize our findings regarding evaluations of a single section of the public summaries, section 2.1, which concerns itself with the disclosure of information about public data sources used to train the model. We determine that, though disclosure in this section is sporadically done well, the provision of information by Big AI organizations in particular generally fails to meet disclosure ideals.
General information
First, a note on scope. Section 2.1 limits itself strictly to ‘publicly available’ datasets. ‘Publicly available’ here entails that a dataset is (1) compiled by a third party, (2) made publicly available for free, and (3) is readily downloadable, either as a whole or in predefined chunks. As such, it does not pertain to, and should not contain, (1) datasets created by the provider themselves and subsequently made publicly available (which belong either in Section 2.3, 2.4, 2.5, or 2.6 depending on content), (2) datasets that are not publicly available, and are instead acquired through a license agreement or some means from a third party (which belong in Section 2.2 about private datasets), (3) datasets made available by third parties, but not in a publicly available manner, such as datasets exclusively available through personal correspondence with the author (which belong in the same sections as the first category).
Section 2.1 splits its requirements for the provision of public datasets into two parts. First, it requires that each large dataset used in training (with ‘large’ defined as at least 3% of the training data for a given modality after preprocessing) be reported exactly, and that all datasets which are not considered ‘large’ under the above definition are described generally. That is, the datasets should be explicitly named in such a way that they can be exactly identified. A link to the dataset is required if it exists, and otherwise the dataset should be described in such a way to at least include the start and end dates of data collection. A general description for filtering should be provided in cases where only part of a given dataset was used.
Second, the section requires a general description of all other publicly available datasets through high-level prose. Suggested information includes the types of modality included in data, the nature of the data’s content, the linguistic characteristics of the data, and the times of collection. Examples of datasets which should be included Section 2.1 include datasets made available on public repositories (such as HuggingFace), specialized websites, as well as CommonCrawl snapshots.
Practical disclosure -- main issues
In practice, we see a distinct lack of disclosure for this section, especially among BigAI companies. Most often, organisations which train the biggest AI models only disclose use of the CommonCrawl. The general description of other datasets is commonly even left empty, with the implication being that essentially ‘only the Common Crawl’ was used. We view this as highly unlikely, given that raw CC is generally not considered high-quality data. Most open-data large language models incorporate large numbers of datasets of various sizes, and so the description should provide comprehensive information about such data. Given that the goal of the template is to facilitate rightsholders, we believe that disclosure should be done in such a way as to allow a rightsholder to determine which, if any, public data might be subject to licensing restrictions. The current narrow descriptions do not do so.
What is more, the CommonCrawl consists of many different snapshots, and its data is almost always filtered. Providers rarely indicate which snapshots were used in the training of their model and do not disclose the filtering methodologies they use. This runs in stark contrast to the suggestions provided by the template. We would encourage providers, particularly among BigAI, to rectify these problems to ensure that rightsholders are able to properly make use of the template as a source of information on what publicly available data current large language models are trained on.
The CommonCrawl is a crawl of the entire internet. By only listing it as a source without any information on how data was filtered, providers are answering a question analogous to “where did you get this information?” with a generic “the internet”, failing to meet even grade school standards of sourcing.
Practical disclosures -- minor issues
Besides these main issues, we also observed several smaller issues in this section in individual summaries. First, we observed cases where the list of large datasets was missing entirely (Phi-4) or contained only minimal/unnecessary information (Grok 4.5, Gemini 3 family, Muse Spark, Muse Image, Gemma 4, Nova 2 lite). Second, we observed limited instances of providers providing unneeded datasets in the list of large publicly available datasets (Domyn Large) and general description (SmolLM3). Third, we observed a case where the modalities indicated for public data were relatively unlikely (Muse Image, where a text-to-image model was declared to include public video and audio data). Last, we observed a single instance of some superfluous information being provided in the section for additional comments (Minimax M3).
Conclusion
In general, we see poor compliance with this particular aspect of the template of public training data. By providing little practical information, providers take away from the ability of rightsholders to identify key rights issues in public data sourcing. This, in turn, hurts the broader ecosystem of rights enforcement particularly as it pertains to LLMs.