
Dataset Context(ualisation) in Documentation: Best Practices, Recommendations and Open Questions
Abstract
This discussion paper reports on the ongoing work on two primary templates for dataset documentation, namely datasheets and data-envelopes. It explores the critical role of comprehensive and consistent dataset documentation in (computational) reuse of datasets. Based on the experiences of the authors while using these templates, the paper provides recommendations and discusses open questions. Key recommendations are balancing free text and interoperability, collaboratively establishing dataset documentation in an interdisciplinary manner, acknowledging the value of human contributions in the documentation process, dealing appropriately with dynamic datasets as opposed to static ones, and integrating other forms of documentation. As we ascertain that motivating dataset providers to maintain documentation presents a significant hurdle, we suggest integration into existing workflows, recognition of contributors, and normalisation within datasets’ social contexts as potential solutions. We also lay the groundwork for future research by identifying open questions related to quality assessment mechanisms, documentation of bias, developing adaptable and interoperable templates that cater to different domains and dataset types, describing datasets in the most appropriate languages, and integrating documentation into users’ workflows. Finally, we introduce a user-friendly dataset documentation tool, being designed to handle structured fields, controlled vocabularies, and format mapping.
© 2026 Henk Alkemade, Gustavo Candela, Steven Claeyssens, Selda Eren, Maria Eskevich, Nuno Freire, Antoine Isaac, Jörg Lehmann, Giulia Osti, Mari Wigham, published by Ubiquity Press
This work is licensed under the Creative Commons Attribution 4.0 License.