Enforcement day is here! 39 public summaries assessed, 20 still missing

2026-08-07 | 2nd August 2026 marks start of enforcement, yet we found 39 Summaries and think 20 are missing

This work has been featured in an Euractiv article titled AI labs at odds with EU over half-hearted data disclosures with more information and quotes from GPAI providers and the AI Office.

With enforcement day having arrived, we are happy to report that many additional providers have submitted summaries in fulfilment of EU AI Act Article 53(1)(d).

We found 39 public summaries so far, and have evaluated 38 of them. Based on our preliminary analysis, we think 20 models are currently missing a public summary (not counting models for whom enforcement starts in 2027). The average scores for evaluation are predominantly C grade for transparency, and C- to D for usefulness. This means though the public summaries contain information, it is generally not sufficiently detailed nor does it give any actual actionable information. Since we based our analysis on the text of the AI Act and the explanatory note, we think this clearly shows that most providers are not following the established norms and are trying to box-tick their way out of actual compliance obligations. Our research in this work is based on our published peer-reviewed article at the top-ranked ACM FAccT conference in June, which is accessible from our website. We are engaging with the AI Office to share our findings and make recommendations.

Some of the specific issues we found

Discovery / Findability of Document(s)

Public summaries are VERY difficult to find. We had hoped that simple searches should be sufficient to surface them, but we required lots of manual effort to discover that a summary was made available on a dedicated page (e.g. Google, Anthropic, Mistral) that we had to 'accidentally' stumble upon. This is despite the AI Act and the guidance requiring summaries to be provided in the "context of model", which we interpreted as alongside other existing documentation. Some providers required us to follow several links until we reached the summary in an non-obvious manner (e.g. OpenAI has it in a generic AI Act help page not part of their model documentation), but we also found examples where providers clearly indicate 'legal documents' and where to find the public summary. We also found some providers using the public summary as a vehicle for their marketing claims or virtue signalling instead of actually giving the requested information.

Vagueness / Obfuscation of Information

We had difficulty in understanding the summaries because it was not clear which models are or aren't being covered by the public summary. For example, many providers only mention a single model but do not mention other derivatives which would also be covered by the summary. Others, such as Google use broad categories such as "Gemma/Gemini Pro family" and provide a single document even when there are significant differences in the models and the way they have been trained. We think obfuscates the ability to determine how specific data was used for a specific model, which makes it difficult to perform rights-interest based assessments e.g. copyright. In comparison, Anthropic and Mistral have published separate summaries for models within the 'same family' which often show important differences, both technical as well as legal aspects such as 'date of market placement'.

Lack of Information / Incompleteness

The objective of the summary was to make transparent what data was used. Yet, information was missing or incomplete in obvious ways. Most providers state they use commonly known datasets Common Crawl but fail to declare which version was used – this is important because the Crawl datasets are versioned based on time of collection and updated periodically. The impact of this is that we do not know what web documents were part of the training dataset e.g. by analysing the corresponding Crawl. Another obvious thing was that despite template explicitly asking for domains crawled, no provider has indicated which specific websites they crawl. For example, Reddit sued Anthropic in 2025 for allegedly ‘scraping’ user comments to train Claude [AP, 2025], while we found no information about domains scraped within the public summary. The summary also explicitly asks for the identifier for crawlers used and we found some companies, e.g. Mistral, still do not provide this information. We think that this is a clear actionable omission as rightsholders have no ability to determine when their content is being scraped, and are denied the ability to block or opt out from having their data being collected.

Ambiguity in Language

We noticed a general tendency to use unclear language. For example, "some data was collected up to <date>" (Anthropic), "may use <action> for filtering illegal content" (Microsoft). While this creates the impression of information, from a legal standpoint, this is dangerously incomplete and misleading. This means rightsholders and potentially impacted people are not sure whether these companies actually took an action or not, such as removal of content for a specific model which impacts them. Similar issues are also present regarding synthetic data despite the template providing explicit fields for this information.

Handling of rights / legal issues

Providers rarely describe (1) how they make sure that they are up-to-date with user rights requests, nor (2) how they select public data sources in a way which ensures that they do not contain copyrighted materials in the first place e.g. Common Crawl was reported by the Atlantic to ignore paywalls/blocks [Atlantic, 2025] , nor (3) what they do in the event that a public data source they use gets taken down through e.g. a DMCA request e.g. Books3 report from Gizmodo [Gizmodo, 2023]. We also noticed that none of the data sources currently being part of lawsuits regarding AI training have been mentioned in the public summaries even though they represent significant aspects such as 1) source of data if obtained from copyright holders or third-parties 2) removal of illegal content 3) use of specific crawlers. Of particular interest, xAI's summary for Grok does not mention CSAM removal [Wired, 2025], [Ars Technica, 2026] or any relevant measures in its section on removal of illegal content.

Summaries not for all models

We find that summaries for some of the most recent models are missing: full list is on our website. For example, GPT 5.6 Luna has a published summary, but GPT 5.6 Sol does not, and we are not sure as to why a provider would selectively publish information about only one of its models. If Luna is the base model for Sol, then in theory, OpenAI does not have to publish a separate summary for Sol, unless it meets the criteria for having significant differences so as to count as essentially a separate model. However, we should also consider that both Luna and Sol possess systemic risks under the AI Act given their significant technological advances, large size, and being leading frontier models. Based on this, we are aiming to treat such cases as being separate models requiring their own summaries in our future assessments.

Chinese providers do not provide summaries

We could find only one Chinese AI model provider's published summary (MiniMax). This should be of particular interest to the AI Office considering that recently models such as Moonshot's Kimi K3 have been reported to be comparable with the larger US models in terms of capabilities [Reuters, 2026] . Other notable Chinese companies like DeepSeek, Alibaba, Baidu, Tencent, Xiaomi, Huawei have all also recently released models, some of which were also in the news recently due to being open weights and having a perceived capability at par with US models [Bloomberg, 2026].