Abstract
This introduction to the special collection Benchmarking in Digital Humanities examines the emerging role of benchmarking within digital humanities research and argues that benchmarking in the humanities cannot be understood solely as a technical procedure for evaluating computational performance. While benchmarking is often associated with standardized metrics, model evaluation, and task performance inherited from scientific and computational traditions, the contributions collected in this special collection collectively demonstrate that benchmarking in digital humanities increasingly begins at the prior stage of entity construction.
Drawing together a range of data papers, methodological reflections, and evaluative frameworks, the collection highlights how benchmarking practices in digital humanities are shaped by domain-specific interpretive concerns, archival conditions, representational choices, and humanities-oriented research questions. Across the articles, benchmarking emerges as a layered and iterative process spanning schema design, evaluation logic, qualitative assessment, and epistemic reflection. The collection points toward a reflexive and humanities-oriented approach to benchmarking, in which alignment between entity design, representation, workflow, and evaluation becomes central to the research process itself.
© 2026 Jenny C. Y. Kwok, published by Ubiquity Press
This work is licensed under the Creative Commons Attribution 4.0 License.
