Running commentary · September 2026
Before We Train on 30 Million Objects

The paper
European Commission, DG EAC / DG CONNECT (2026). Artificial intelligence strategy for the cultural and creative sectors: Call for Evidence. Have your say portal, initiative 18592, feedback open until 11 September.
Paradata, the missing half of the European cultural data space. My feedback to the European Commission's Call for Evidence on an AI Strategy for the Cultural and Creative Sectors.
The Call for Evidence announces a planned AI Strategy for the Cultural and Creative Sectors, due Q1-2027. A deliverable of the Culture Compass, the strategy will set out a strategic vision for the use of artificial intelligence in the cultural and creative sectors and industries. It will benefit culture in general and the European cultural sector in particular, as well as citizens, by supporting innovation, unleashing AI for European cultural sovereignty and promoting the development and use of AI to uphold genuine creation, and cultural and linguistic diversity and inclusion.
Feedback is open till the 11th September at the Have your say portal.
My feedback
I am responding to one sentence in the Call for Evidence: that the common European data space for cultural heritage will make available over 30 million digital cultural objects for AI training. I support the ambition, and I want to record a caution about what those objects presently are. I write as someone who produces such objects, in a national heritage institution in a small Member State, and whose doctoral research asks what they actually contain.
A 3D model or a digitised photograph is not knowledge. It is a measurement made by a person, with an instrument, under constraints, for a purpose, and then interpreted. The record of those conditions is paradata. Without it an object can be described but not evaluated, and a model trained on it inherits an assertion it has no means of checking. Metadata tells us what a thing is called. Paradata tells us whether to believe it.
The gap is structural rather than a lapse by any institution, and that is why it is a policy matter. The Europeana Publishing Framework grades metadata through language, enabling elements and contextual classes, all criteria of discovery, while the Europeana Data Model offers no structure for documenting the digitisation process itself. The custodial provenance of a physical object is expressible. The making of its digital surrogate is not. Across the aggregated corpus, capture method, instrument, accuracy, operator and processing chain are therefore generally absent, because nothing asks for them. HERITALISE (Horizon Europe, GA 101158081), 423 respondents across 97 countries, documents fragmented practice and the near-systematic absence of process documentation in heritage 3D. VIGIE 2020/654, the Commission's own consolidated study, names the paradata gap but does not specify a capture structure for it. Objects offered for AI training on this basis are unsourced assertions, and no transparency obligation downstream can recover what was never recorded at capture. This will not be closed by guidance. It has to be specified.
My doctoral research proposes one response, which I call the Memory Twin. A digital twin mirrors the physical state of an asset. A Memory Twin carries, as machine-readable structure grounded in CIDOC-CRM, what a surface cannot hold: provenance, interpretive claims and their authors, degrees of uncertainty, contested readings, and community testimony. This bears on Pillar 2, because cultural diversity is not only a question of which languages and countries appear in the training data. It is a question of whether plurality survives inside a single record, which generative summarisation tends to resolve into one fluent narrative.
Four requests
The Call for Evidence sets out three pillars around which the strategy is expected to be structured. Since I refer to them repeatedly, they are worth stating plainly. Pillar 1 is about fostering innovation and competitiveness through collaboration between the cultural and creative sectors and European AI developers. Pillar 2 is about fair business models and the ethical use of AI, protecting European creation and original content, and safeguarding cultural and linguistic diversity. Pillar 3 is about equipping the sectors for the digital and AI transition, identifying the support and adaptations they need to seize the opportunities while protecting creativity, agency and cultural content.
My four requests sit mostly in the second and third, with one that cuts across both and one addressed to the monitoring arrangements described in Section B.
One. Make paradata a condition of licensing, not an optional enrichment (Pillar 2)
Pillar 2 asks how European creation and original content can be protected in an AI economy, and the Data Union strategy answers partly through fair and normalised licensing of the cultural heritage data space. My request is that the licence carry the record's own account of itself.
Concretely, a minimum paradata profile would state the capture method, the instrument and its configuration, the achieved resolution and accuracy, the operator and the date of capture, the processing chain from raw data to published derivative, and the interpretive decisions taken along the way, including what was reconstructed, what was cleaned, and what was left as recorded. None of this is exotic. Most of it exists transiently, in a project folder or a technician's head, at the moment the work is done, and evaporates within months. Capturing it at source costs very little. Recovering it later costs everything, because it usually cannot be recovered at all.
The reason this belongs in a licensing condition rather than in guidance is leverage. Guidance has been available for years and the corpus looks the way it looks. If access to a European data space of 30 million objects is being offered on normalised terms, those terms are the one place where a profile can actually be required, and where the absence of one becomes visible rather than invisible.
Two. Declare the boundary between what was measured and what was generated (cuts across Pillars 1 and 2)
This is the request I would most like the drafters to take seriously, because it is the one that gets harder every month.
Neural and generative representations, of which 3D Gaussian Splatting is the current example, do not record a surface. They optimise a set of primitives until the rendered views match the input photographs. The result can be extraordinarily convincing and genuinely useful, and it is not a measurement. Crucially, it is not reversible: you cannot recover from the representation which parts correspond to observed geometry and which are the optimiser's inference filling a gap it was never shown. A photogrammetric point cloud with a hole in it tells you there is a hole. A splat renders the hole away.
For heritage this matters more than for most domains, because our records outlive the things they record, and because a plausible surface is exactly what a future researcher will trust. My request is narrow and technically feasible: the distinction between measured record and generated reconstruction should be declared at file level, and that declaration should survive aggregation, so that a measured object and an inferred one are never silently equivalent downstream. Pillar 1's collaboration between cultural institutions and European AI developers is precisely where such a convention could be agreed, and it would give European tooling something to be distinctive about.
Three. Fund documentation capacity, not only tool adoption (Pillar 3)
Pillar 3 is framed around upskilling, reskilling and AI literacy, and the Call for Evidence rightly notes the risk of reinforced inequality for organisations with limited access to AI resources. From inside a small institution I would put the constraint differently. The binding limit is rarely access to AI, which is cheap and getting cheaper. It is person-hours. It is whether anyone is paid to describe what the institution already holds.
Funding instruments tend to reward the visible new thing: the new scan, the new platform, the new immersive experience. Documenting an existing collection to a standard that makes it usable by anything, human or machine, is unglamorous, slow, and almost never the headline of a successful application. The result is predictable and it is what the corpus shows: more objects, less about them.
So the request is that documentation and paradata capture be recognised as eligible, fundable activity in its own right under Creative Europe, Digital Europe and the successor programmes of the 2028-2034 MFF, and that documentation practice be counted explicitly as an AI-readiness skill in the upskilling measures the strategy proposes. A cataloguer who records how a model was made is doing AI preparedness work, whether or not the word AI appears in the job description.
Four. Measure paradata completeness, and publish the measurement (monitoring, Section B)
The Call for Evidence describes a monitoring architecture built on the forthcoming EU Cultural Data Hub and the AI Observatory announced in the Apply AI strategy. My request is that one of the things they measure is paradata completeness across the cultural heritage data space.
This is not a technically demanding ask. Europeana already computes quality tiers for every record it holds, automatically, in its ingestion pipeline, and publishes them back as part of the data. The mechanism exists. What is missing is a dimension that scores process documentation alongside the existing discovery-oriented criteria. Add it and the picture becomes visible for the first time: which collections can account for themselves, which cannot, and whether the position is improving.
I make this request last because it is the one that makes the other three enforceable. What is not measured will not be funded, and completeness is measurable.
Closing
The first guiding principle in Section B is technology at the service of culture. A working test: does an AI application in heritage preserve the traceability of human interpretation, or dissolve it? A Memory Twin is one attempt to keep the human in the record, and not merely in the loop.
This is a personal reading by Anthony Cassar, PhD research fellow on the Memory Twin Framework. The views expressed are my own.