Skip to main content
Have a personal or library account? Click to login
Dataset Context(ualisation) in Documentation: Best Practices, Recommendations and Open Questions Cover

Dataset Context(ualisation) in Documentation: Best Practices, Recommendations and Open Questions

Open Access
|Jun 2026

Full Article

(1) Context and motivation

The selection, acquisition, conservation, exploration and exhibition of cultural objects have always been core activities of institutions such as galleries, libraries, archives and museums (GLAM). Even in the 21st century, in the age of digitalisation and data, these tasks remain largely unchanged. Digitalisation, however, has led to metadata and representations of cultural objects being published as cultural heritage datasets. As high-quality, curated datasets that are permissively licensed continue to gain value, e.g. for research or the creation of machine learning (ML) models, the contribution of the GLAM sector is becoming increasingly important (Kandpal et al., 2025; Langlais et al., 2025).

Less visibly, datafication plays an important role in shaping what is included in these datasets and for what purpose. Datafication refers to the transformation of cultural practices, contexts, and meanings into structured and quantifiable data. As Edmond et al. (2022) point out, this process does not simply make culture available for computational use but can also shift attention away from cultural practices themselves toward what can be measured and standardised. As a result, cultural heritage datasets reflect both the diversity of cultural materials in terms of time, place, languages and social contexts, and the specific selection processes and use cases that have influenced how the datasets were created.

Cultural heritage institutions (CHIs) have long-standing expertise in documenting provenance, contextual relationships, and curatorial decision making around collections. At the same time, the notion of a dataset lifecycle and the associated logics of standardised, machine-readable documentation have only more recently begun to intersect with CHI practices. In particular, there is a lot of attention on how documentation can support making CH datasets more Findable, Accessible, Interoperable and Reusable, following the FAIR Principles (Wilkinson et al., 2016). In this context, instruments for dataset documentation function as a translational layer that supports the expression of curatorial choices, assumptions, and constraints in forms that are legible for computational reuse (Jo and Gebru, 2020). Consequently, templates for dataset documentation have been made available.

As a plurality of templates have emerged, this paper aims to consolidate approaches to dataset documentation in the CH sector, by articulating the key information categories, best practices and recommendations that can help (i) CH practitioners to create relevant dataset documentation and (ii) information technologists to support them in this endeavour.

(2) The templates: datasheets and data-envelopes

Templates for dataset documentation specific to digital cultural heritage have recently been established in the form of datasheets and data-envelopes (Alkemade et al., 2023; Eskevich & Luthra, 2024; see also Lee, 2023; Maemura & Byrne, 2024; for an overview, see Maemura, 2025). Moreover, there exist at least two journals that provide researchers with the possibility to publish concise descriptions of humanities and social science data in the form of peer-reviewed publications; both provide templates for these data papers (McGillivray et al., 2022).1 Such documentation can operate as a critical and reflexive device, making aspects of data collection and curation more visible that might otherwise remain implicit. The plurality of templates, approaches and target communities does not come by chance: it is largely rooted in the “Collections as Data” movement, which has deeply impacted how practitioners, institutions and researchers frame their interaction with heritage materials. This movement encourages CHIs to provide datasets “amenable to computation” (Padilla et al., 2022, p. 20), suitable for direct computer processing, or for developing machine learning applications. Datasheets and data-envelopes are aimed at supporting data consumers in finding and using datasets, and as such are highly structured and machine-readable, so that researchers or search tools can quickly pick out the relevant information. The differences between datasheets and data-envelopes reflect the preferences of their different user communities, e.g. for explanatory text or for selected terms from a vocabulary.

Researchers may also be interested in a transparent publication of the basis of their data-driven research, in order to improve the metrics associated with their publications as well as to promote data reuse (Roued-Cunliffe, 2020). Peer reviewed publications answer this need, using the more loosely structured, human-readable format of a scientific paper.

From a more technical point of view, templates for dataset documentation should incorporate highly structured metadata, for example, by using Schema.org2 or the W3C Data Catalog Vocabulary DCAT and the DCAT Application profile for data portals in Europe (DCAT-AP),3 which is used by most data portals in Europe and maintained by the Interoperable Europe SEMIC action. This facilitates machine-readability as well as interoperability. Documentation frameworks aim at integrating dataset documentation into repositories, databases and (meta-)search engines, dataset hubs or EU data spaces.4 DCAT and DCAT-AP are for instance employed as a basis for creating (machine-actionable) descriptions of datasets, data services and data offerings in catalogues for data spaces. We envision that in 10 years from now, datasets accompanied by documentation will be standard items in existing catalogues and finding aids provided by GLAM institutions, and the same will be the case for machine learning models that use such institutional data, accompanied by their model cards (Mitchell et al., 2019). To bring such a vision to life, dataset creators as well as data curators need to become accustomed to using dataset documentation templates. This article reflects on our experiences with creating and filling in datasheets and data-envelopes, provides best practices and recommendations, and raises open questions to stimulate discussion.

(3) Applications in practice

Since the publication of the first version of the datasheet template in September 2023, the Europeana Working Group dedicated to “datasheets for digital cultural heritage”5 has continued to work on improving and specifying the datasheet template. A structural update was carried out by rearranging sections and the sequence of individual items within them, rephrasing explanatory texts in order to emphasise the specifics of digital cultural heritage, and aligning all template items with DCAT-AP in order to enable machine-readability whenever possible. The latter was strongly motivated by the ongoing development of the common European data space for cultural heritage6 which, in line with data space efforts in other domains, has planned to develop a dataset catalogue. Following the approaches taken for the existing Europeana Data Model (EDM) and its extensions (Doerr et al., 2010), this calls for the adoption of linked data-compatible vocabularies as well as the establishment of more comprehensive, community-based documentation. This process led to the publication of version 2 of the “Template: Datasheets for Digital Cultural Heritage Datasets” in July 2025 (Alkemade et al., 2025).

Data-Envelopes, building on the ideas of datasheets, were developed out of the urgent need for a practical tool for the creation of both human- and machine-readable documentation, and as such are accompanied by their own editor. Since their introduction in 2024, the focus has been on testing the concept on a large number of diverse datasets from different institutions. Feedback from users has led to minor improvements in the template, simplifying and refining existing fields while adding requested new fields.

The datasheet and data-envelope templates broadly answer the same questions about datasets, shown in Table 1. There are some minor differences in levels of detail in different areas that reflect the differences in the domains they come from: research and cultural heritage. Both build on existing dataset documentation research (e.g., Gebru et al., 2021; Pushkarna et al., 2022), adding in essential elements for humanities and cultural heritage that were previously missing.

Table 1

Types of information about the dataset: content, context of production, use of the dataset.

CATEGORYQUESTIONTYPICAL DOCUMENTATION CONTENT
Content of the datasetWhat is in the dataset?Title, description, category, languages, geographical and temporal coverage, statistics etc.
How is the dataset structured?Structure, data fields, examples, compliance with standards, data splits etc.
Where do I find more information?Homepage, repository, papers and other references, contact persons etc.
Context of productionWhy was the dataset created and how?Motivation, description, collection process, sources, provenance, digitisation, preprocessing, cleaning, annotation etc.
Who was involved in the production of the dataset?Contributors, publishers, funders, annotators and their positionalities etc.
Use of datasetWhat can the dataset be used for?Supported tasks, existing uses, suitable uses, unsuitable uses, applications, use in datasets and models etc.
How do I use the data (responsibly)?Licence, access and citation information, version, maintenance, ethical considerations, sensitive information, biases, societal impact, limitations etc.

A good approach to creating data documentation is to look at examples. Within the project “Accessing Context” at the Huygens Institute,7 data-envelopes were filled in for a range of research datasets and archival materials. Members of the Europeana Working Group, by contrast, used the datasheet template for data publications from the cultural heritage domain. So far, about 70 data-envelopes have been created,8 and datasheets were written during several stages of template development: The first datasheets used the template established by Gebru et al. (2021, resulting in Gerber & Lehmann, 2023; Labusch & Lehmann, 2023), then the first version of the template “datasheets for digital cultural heritage datasets” was used (Federbusch et al., 2025; Lehmann & Schneider, 2024; Schneider & Lehmann, 2024), and a range of datasheets is available using version 2 (Alkemade, 2026; Baierer et al., 2025; Claeyssens, 2026; Krems et al., 2025; Rose, 2026; Schmideler et al., 2025; Schneider et al., 2025; Schneider & Lehmann, 2025). This paper results from the combined experiences the authors made while using data-envelopes as well as datasheets.

Experience with datasheets and data-envelopes shows that these two approaches should not be treated as mutually exclusive. Both datasheets and data-envelopes emphasise – to varying degrees – an interdisciplinary approach to creating as well as considering the positionality of the creators of the datasets (and of the documentation itself). During the process of drafting and writing the dataset documentation, the focus is clearly on the dialogue between domain and technical experts. As Alkemade et al. (2023) already mention, the “collaborative filling of datasheets helps to become aware of the particularities of the content”. For example, conducting some basic statistical analyses of the data and creating descriptive statistics turned out to provide a solid basis for further exchange between the technical staff and the cultural heritage practitioners and led in most cases to dynamic and fruitful discussions. Cultural heritage practitioners have an excellent knowledge of the collections in their subject areas as well as the analogue and digitised material, they are well aware of the biases in the datasets and are able to describe them narratively.9 These two kinds of expert knowledge provide the foundation for a translatability between statistical and ethical biases (see as an example the discussion of biases in Krems et al., 2025).

(4) Best practices and recommendations

With regard to the choice of dataset documentation, we recommend exploring your motivation and using templates and publication outlets that suit that motivation. It is better to have any kind of documentation than none. During workshops with dataset providers on data-envelopes, a distinct cultural difference was noted between archives and universities. Archivists tended to regard documentation as a core part of their role, and to be enthusiastic about the chance to transmit more of their knowledge. Universities – mainly represented by library staff and data stewards – were typically not responsible for documentation themselves but for guiding researchers, whose focus lay on research and not on documentation. As a result, they were sceptical about being able to persuade researchers to create more extensive documentation.

Data are not just electronic objects, they also have a social life. Platforms and repositories are only partly data silos; they are also places where digital humanities research, machine learning, and cultural heritage communities meet and exchange ideas. In principle, the choice of one or more repositories in which to publish a dataset should be based on the target community that is to be reached with the publication. Repositories that are reliably used by the relevant user communities increase the visibility of such datasets. Publicly accessible repositories like Zenodo, the Hugging Face Hub or the CLARIN10 and DARIAH11 research infrastructures, the upcoming European Collaborative Cloud for Cultural Heritage ECCCH,12 the common European Data Space for Cultural Heritage13 or the European Open Science Cloud EOSC14 are a recommended addition to the platforms provided by GLAM institutions. They may be more suitable for the publication of large datasets, they may not incur operating costs for the cultural heritage institution or library, and they obviate the necessity for the latter to operate a publicly accessible platform itself. The decision as to where a dataset should be published is left to each data provider; it is entirely possible and also sensible to publish a dataset in parallel in different repositories in order to reach a wide audience and thus ensure maximum impact with regard to the use of a dataset in different contexts. In addition, datasets published in different locations remain available at all times, because the LOCKSS principle applies: “lots of copies keep stuff safe”. However, publishing datasets in several places can cause issues, e.g., with multiple locations having different versions of the data. While the responsibility to solve these issues does not necessarily lie with the original publisher, using a unique (and persistent) identifier for the first publication – as supported by the templates – helps to solve this challenge. These identifiers can indeed be used to point back to the initial publication from the re-publications in different repositories.

We zoom in on best practices and recommendations for two aspects of the framework: the content of the documentation (4.1–4.3.) and the human component (4.4.–4.5).

(4.1) Process: Balancing narrative and interoperability

We recommend allowing free text and at the same time encouraging the use of vocabularies. Free text, in particular narrative text, is a natural form for data providers to fill in and for users to read, and it offers considerable scope and flexibility. It is indispensable for conveying contextual information such as provenance, historical background, collection practices, or nuanced descriptions of bias and positionality. However, free text is unhelpful to users who need to quickly find answers to specific questions, such as the licence of a dataset or its update frequency. As free text is not machine-readable, it is also less helpful for systems that aim to support users in searches or large-scale aggregation efforts (unless these exploit AI technology, which could be error-prone and/or judged not desirable for such ‘basic’ scenarios where a suitable alternative exists).

Using vocabulary terms and other structured value spaces is less familiar for many data providers and users. Data providers must search for suitable terms, while users must accustom themselves to definitions and hierarchies. Nevertheless, vocabularies support consistency, which in turn improves comprehensibility and interoperability (Smith, 2021). As structured values are machine-readable, they enable filtering and aggregation across datasets, e.g. identifying all datasets with a CC0 licence updated within the last year. Vocabularies can also assist in disambiguation, ensuring that an agricultural researcher finds datasets on apples while a computer scientist finds datasets on Apple computers. Hierarchies within vocabularies can further support discovery, such as retrieving datasets from all European countries or related to machine learning topics.

It is generally impossible to find a single vocabulary that covers all documentation needs. Moreover, the more comprehensive a vocabulary, the harder it can be to identify appropriate terms, and the higher the risk of ambiguous or irrelevant terms. Switching between several well-focused domain vocabularies is therefore an attractive option in practice if this is supported in a simple and transparent manner. Where no suitable vocabulary term exists, entering free text should remain possible, but as this reduces comparability and interoperability, it should be treated as a last resort, and users should be encouraged to use vocabularies wherever feasible.

Different documentation fields benefit from different levels of structure. Certain elements such as licensing, access conditions, identifiers, temporal and geographical coverage, or update periodicity strongly benefit from controlled vocabularies or standardised value spaces, as they are frequently used as decision criteria by dataset users and aggregators. Other elements including descriptions of data collection, selection criteria, historical context, or ethical considerations require narrative explanation and cannot meaningfully be reduced to predefined terms alone. For such fields, machine-readability can be supported indirectly through clear structuring, consistent sectioning, and explicit signalling of missing or inapplicable information. A specific case where free text can create confusion is that of empty fields (Di Pasquale et al., 2024). An empty field may indicate that information was forgotten, unavailable, deemed irrelevant, or omitted due to time constraints. Such ambiguity is problematic for both human users and automated processing. We therefore recommend explicitly indicating the reason for missing information, for example by using “not applicable” to signal that a field does not belong to the dataset, or “information not available” to explain its absence. This practice is particularly important for small or legacy datasets, where documentation templates may not be fully applicable.

Overall, balancing narrative flexibility with structured representation is essential for producing dataset documentation that is both usable and interoperable. Combining controlled vocabularies for key decision fields with narrative text where contextual richness is required allows documentation to serve diverse audiences while supporting aggregation, discovery, and reuse across platforms and communities.

(4.2) Process: Static vs. updated datasets

We recommend keeping a close connection between dataset documentation and the dataset itself. This can easily be implemented when depositing a static dataset in a repository, e.g. a publication of a dataset on Zenodo or the Hugging Face Hub accompanied by a datasheet or a data-envelope, or when a dataset harvested by an aggregator such as CLARIN VLO links back to the documentation stored at the data provider. In the case of datasets that are (automatically) updated, such as is the case with web archives, there should be at least a link to the dataset documentation present on the same page where the dataset is available. Such dataset documentation need not comprise information on specific file formats or their checksums. Rather, we recommend to provide information on what the dataset contains in a more generic way and to provide information on the update periodicity and the mechanism for updates, where available.

(4.3) Process: Link to other forms of documentation

We recommend linking to supplementary information or other forms of documentation. The dataset documentation does not have to contain absolutely every piece of information about the dataset. In fact, this is undesirable as it makes it harder for users to find the information that they need. The documentation should be compact, containing the key pieces of information that a user needs to know. Sometimes, it is enough to make a user aware of a potential issue and then point them towards more information. This can be done by linking to other forms of documentation, for example to a paper describing the production of the dataset, to a blog post going deeper into the historical context, to a README in a Github repository containing the details of processing, to a schema, an ODRL15 model, or the report of a FAIRness assessment. One should strive to provide such links using persistent identifiers where possible to ensure access in the future as it is frustrating for a user to be confronted with dead links.

As the majority of digitisation projects in the member states of the European Union are financed by public funds, data management plans (DMPs) are a regular (and often mandatory) component of the projects. We recommend that dataset documentation, preferably in a standardised form, becomes an integral part of data management plans, to ensure high quality data publications. Data management plans are also a potential source of information for documentation of legacy datasets.

(4.4) Human component: Skill development and collaboration in dataset documentation

We recommend collaborating, building and sharing expertise. New datasets can be documented by the team that creates them, whereas for legacy datasets that team is often long gone, so the process of documentation becomes a form of archaeology better suited to a data practitioner.

Collaboration is almost always essential as dataset documentation brings together domain knowledge, technical knowledge, legal knowledge and user and/or community knowledge. These rarely co-exist in one individual. For example, academic researchers are often unaware that open data should still have a licence specified, and creators may focus on their knowledge of the dataset while missing the information that dataset consumers need. Often the different types of knowledge required exist within the same institution, but not always. When filling in datasheets and data-envelopes it was noticeable how cultural heritage and humanities institutions are often lacking in knowledge of machine learning aspects of datasets and the needs of data consumers, in particular those coming from the field of AI.

Creating clear, well-structured, understandable documentation of a dataset is a skill. Users of the templates need to first become acquainted with them, given the sheer volume of information that has to be read, understood and filled in, and how potentially unfamiliar to the users it could be (e.g. licensing, use in machine learning, positionality). Routine use of the templates helps considerably, as it supports memorising the structure and aids consistency, which as a consequence improves readability and interoperability. Learning where to place specific information may be difficult, and template users need to learn to flexibly navigate between the sections. The variety in datasets and domains is also reflected in a variety of which information is required or relevant. Thus, template users who only document datasets occasionally may struggle. These challenges can be addressed with two concrete measures. First, it could be useful to develop a network of ‘key experts/consultants’, who are experienced in creating documentation and can assist their colleagues. Potentially, larger institutions could create a dedicated support team. Second, it is valuable to have high-quality, domain-specific examples to assist users in creating effective and reliable documentation.

Use of documentation templates in practice throws up all sorts of questions and also reveals ambiguities, gaps and redundancies. It is valuable for template users as well as template creators to receive feedback and work on improvements to both the template and its own documentation.

A broader question is whether the dataset provider should have the monopoly on dataset documentation, or whether other parties can produce their own version of the documentation to express their findings and opinions related to the dataset (see Smith, 2006). Potentially the versioning of the documentation and keeping track of its authorship should accommodate such cases.

(4.5) Human component: Documenting human contributions to dataset creation

In general, people seemed to be less enthusiastic about being named as a creator/contributor for dataset documentation than they typically would be about being included as a co-author for a paper. Quite often people undervalued their role, claiming that their intervention on the dataset was minimal. The reasons for this reluctance have not yet been investigated (Plantin, 2019; Scroggins & Pasquetto, 2020). One possibility is that authorship of a data paper published in, e.g., the Journal of Open Humanities Data or the Research Data Journal brings credits whereas contribution to a dataset does not; another is that they perhaps fear that association with a dataset will result in responsibility for that dataset.

The positionality aspects of datasheets and data-envelopes present dataset providers with some difficulties. Some dataset creators regard positionality information as irrelevant given that they regard themselves as neutral researchers, thus following a positivist paradigm. Where creators see the relevance of positionality, they may be reluctant to give information out of privacy reasons, in particular as positionality can include personal factors such as age, ethnicity and gender. This can be a particularly thorny issue when creators or contributors are no longer available to ask their permission. Where creators are willing to give information, they may be unsure what information to give. How do they identify what aspects of their background may influence their decision-making, when the most interesting influences may be factors they are unaware of? Some creators opt for adding short bios in the style often used in project descriptions, which tend to focus on positions held, interests and expertise. This latter positionality definition, even if somewhat formal, can help other researchers interested in working with the data to identify whether and how the lens of the data creator corresponds to their own. For example, whether a dataset put together by a historian could be of use for a linguist or an archaeologist.

(5) Open questions

(5.1) Open questions: documentation quality control

Dataset documentation can be incomplete, or it can be complete but of poor quality. For example: incorrect information, too much or too little detail, use of technical jargon, lack of user perspective, confusing structure. Reviews of dataset documentation are essential to ensure quality. Reviewers need to have domain expertise to check accuracy, data management expertise to ensure clear, complete, well-structured data descriptions, and user affinity to meet users’ needs. This suggests the need for multiple reviewers. In addition to this, some form of quality control is needed to ensure consistency across multiple datasets. Even a single documentation author can ‘drift’ over time, filling in the documentation in different ways. Multiple authors can make completely different choices, which may be perfectly valid but which may make it hard for users to understand and – crucially – compare and combine different datasets. A template for documentation is the essential first step towards consistency, but there are many other issues such as which vocabularies to use, what to include in free text sections, which licence to use under which circumstances, how to cite the dataset etc. Policies on these topics can help, but ultimately quality control is necessary to ensure consistency in practice.

(5.2) Open questions: documentation of bias

The creators of data documentation certainly must get better at the documentation of ethical, legal, and social aspects in datasets (Lehmann et al., 2025). A clear conception of how best to present these aspects in datasheets and data-envelopes may evolve over time; for now, the examples may provide guidance. Similarly, we have to learn how to deal with the biases which characterise the datasets published by CHIs and libraries. In this sense, it is certainly helpful to understand bias “as a productive category of analysis”, because bias is “relational and dynamic, actively shaping and being shaped by social, cultural, and historical contexts” (Luthra & Zijlma, 2025a, 2025b). The topic of bias in digital cultural heritage merits a paper on its own once a broad range of dataset documentation becomes available to provide the basis for a systematic analysis.

(5.3) Open questions: adaptable and interoperable templates

Some parts of the documentation templates are more relevant for certain domains, datasets and tasks than for others. For example, a dataset of weather observations probably has no issues of sensitivity, while a set of scanned 18th century maps will probably not have relevant information regarding data splits and sampling. A dataset user looking for ML training data may want to see licensing and rights information as a high priority, a history enthusiast may only care about the description and the link to the data. Being faced with irrelevant sections of documentation can discourage data providers from filling them in and dataset users from reading them. It may be helpful to design templates so that sections can be easily skipped or shown according to the domain, dataset type and user task; however this creates a new challenge of how to encourage data providers to document aspects that are less familiar or relevant to them but that may be important to others.

As discussed in Section 4, different users may find different templates more suitable for their work. To accommodate different needs but still promote interoperability, it is important to map between these templates. The underlying conceptual similarities in datasheets and data-envelopes, and in particular their mapping to DCAT-AP, offer good potential for a mapping, but this is yet to be implemented.

(5.4) Open questions: language

To make datasets as widely usable as possible, they should be described in a widely understood language. However, there is no universal language. Also, there are other considerations. For example, for cultural heritage datasets it is very important that they are described in the language of the culture and community they refer to, and for international portals, such as Europeana, serving users in their own language is part of their mission. Therefore it can be necessary to create dataset documentation in multiple languages, with all of them being equally reliable. This raises the questions of how to do this without multiplying the effort required and how to support the storage, search and selection of multiple language versions.

(5.5) Open questions: motivation

Creating good dataset documentation takes time: to discover and fill in the information, to review the documentation, to publish it and update it. Dataset providers need motivation to invest this time; they get the reward of being credited in the documentation. The question is what would motivate the very diverse types of dataset providers, and how this can be built into workflows and institutional culture to ensure documentation is created. Good tools and integration in existing systems can help by reducing the workload, as can ensuring that documentation becomes an integral part of the social life of datasets throughout their lifecycle.

(6) Conclusion and future work

As this paper presents, a lot of insight has now been gained on how datasets should be documented in the CH sector to support their well-informed reuse in the most recent application scenarios. Open questions remain, but we expect that some key ongoing developments will help us tackling them. First, we have embarked on an effort to implement the datasheet template as an extension of DCAT-AP. This will enable a wider use of datasheets in the context of the Common European Data Space for Cultural Heritage, as DCAT-AP is foreseen to be the backbone of that data space’s dataset catalogue. It will also formalise even more the alignment of the data-envelope and datasheet templates on the features they have in common.

Furthermore, within the SSHOC-NL Project,16 the development of a user-friendly tool designed to support the creation of datasheets, data-envelopes, and other dataset documentation templates is well underway. Sidedoc is a platform for describing datasets using structured metadata standards.17 It currently supports datasheets as the primary standard, with additional standards, such as data-envelopes, planned. The platform is based on JSON Schema, whereby forms, validation, and field behavior are defined through schema specifications. Data can be stored locally in the browser or exported as files (JSON(-LD), Markdown, PDF, and HTML), while cloud storage integration is in progress. Sidedoc further aims to support exporting workflows to facilitate the sharing and reuse of dataset documentation.

Notes

[1] These are the Journal of Open Humanities Data (JOHD) and the Research Data Journal (RDJ).

[2] https://schema.org/ (last accessed: 27 June 2026).

[3] https://semiceu.github.io/DCAT-AP/releases/3.0.0/ (last accessed: 27 June 2026).

[6] https://www.dataspace-culturalheritage.eu/en (last accessed: 27 June 2026).

[7] Details about the ‘Accessing Context’ project can be found at https://www.huygens.knaw.nl/en/projecten/accessing-context/ (last accessed: 27 June 2026), and broader about the data-envelopes initiative at https://knaw-huc.github.io/data-envelopes/ (last accessed: 27 June 2026).

[8] The list of these data-envelopes can be found at https://knaw-huc.github.io/data-envelopes/examples.html (last accessed: 27 June 2026).

[9] When researchers require further guidance in bias identification and documentation there are tools that support them on this reflective path – https://combattingbias.huygens.knaw.nl/ (last accessed: 27 June 2026).

[10] https://www.clarin.eu/ (last accessed: 27 June 2026).

[11] https://www.dariah.eu/ (last accessed: 27 June 2026).

[13] https://www.dataspace-culturalheritage.eu/ (last accessed: 27 June 2026).

[14] https://eosc-hub.eu/ (last accessed: 27 June 2026).

[15] ODRL Information Model 2.2 https://www.w3.org/TR/odrl-model/ (last accessed: 27 June 2026).

[16] https://sshoc.nl/ (last accessed: 27 June 2026).

[17] https://sidedoc.app/ (last accessed: 27 June 2026).

Acknowledgements

The authors thank all members of the Europeana Working Group, formed jointly from the Europeana Research Community and EuropeanaTech Community, who contributed intellectually to this article, even if they did not contribute to the text itself. The members of this Working Group are, in alphabetical order: Henk Alkemade, Gustavo Candela, José Eduardo Cejudo Grano de Oro, Steven Claeyssens, Giovanni Colavizza, Selda Eren, Nuno Freire, Lianne Heslinga, Alba Irollo, Antoine Isaac, Jörg Lehmann, Clemens Neudecker, Giulia Osti, Daniel van Strien, Andreas Weber, Melvin Wevers. They also thank the many dataset users and dataset providers who put their time and expertise into creating and evaluating dataset documentation or who provided feedback.

This publication was supported by the Europeana Network Association (ENA) Research and Tech communities and funded by the Victorine van Schaick prize money received by the Europeana Working Group for their work on datasheets.

Author Contributions

Conceptualisation – Maria Eskevich, Jörg Lehmann, Mari Wigham; Writing – original draft: Henk Alkemade, Gustavo Candela, Steven Claeyssens, Selda Eren, Maria Eskevich, Antoine Isaac, Jörg Lehmann, Giulia Osti, Mari Wigham; Writing – review and editing – all authors.

DOI: https://doi.org/10.5334/johd.571 | Journal eISSN: 2059-481X
Language: English
Page range: 79 - 79
Submitted on: Apr 18, 2026
Accepted on: May 22, 2026
Published on: Jun 22, 2026
Published by: Ubiquity Press
In partnership with: Paradigm Publishing Services

© 2026 Henk Alkemade, Gustavo Candela, Steven Claeyssens, Selda Eren, Maria Eskevich, Nuno Freire, Antoine Isaac, Jörg Lehmann, Giulia Osti, Mari Wigham, published by Ubiquity Press
This work is licensed under the Creative Commons Attribution 4.0 License.