(1) Context and Motivation
This special collection emerged from a guest seminar in the Computational Humanities Research Group Seminar Series held at King’s College London in May 2025, where questions were raised about how accuracy might be meaningfully established in benchmarking practices within the digital humanities. The discussion, which involved both editors of this collection and the editor-in-chief, highlighted a broader gap in the field: how do we understand benchmarking accuracy in digital humanities, and how can it be measured when what is being evaluated is qualitative? Our inability to arrive at a fully satisfying answer became the starting point for this collection and led us to a more ontological question: what if humanities subjects and projects are not meant to be benchmarked in the same way as those in scientific disciplines – not only at the level of criteria and metrics, but from the design of entities through to the workflow of evaluation itself?
In attempting to answer the above question, I would first like to return to the origins and current perception of benchmarking within digital humanities. In general discussions, benchmarking has often been treated as a relatively stable and self-evident practice, typically associated with evaluating how well computational systems perform predefined tasks through standardised datasets and metrics, as exemplified by benchmark suites such as GLUE (Wang et al., 2019). Yet, this apparent coherence conceals a more complex genealogy that warrants closer attention. The central issue regarding benchmarking in digital humanities extends beyond whether the discipline should move from earlier benchmarking paradigms based on predefined tasks and annotated datasets towards more open-ended forms of large language model (LLM) evaluation or develop benchmarking approaches that are more suitable for humanities. The crux of the matter lies in the epistemic assumptions inherited across these different paradigms, under which the objects of evaluation are typically presumed to be clearly defined, measurable, and comparable before evaluation begins.
This understanding of evaluation objects persists despite the methodological development of benchmarking. Across successive generations of computational benchmarking, the logic of evaluation has evolved, while the underlying assumption has remained largely unchanged. Early computer science benchmarking focused on comparing performance across clearly defined tasks under controlled conditions (Rodríguez and Boyd-Graber, 2021).1 In natural language processing (NLP), this logic became especially visible in shared-task settings that linguistic tasks, annotated datasets, and standardised evaluation metrics (Cer et al., 2017). More recently, LLMs have introduced more open-ended forms of evaluation that increasingly rely on proxy measures and assessments of emergent behaviours instead of narrowly specified tasks (Liang et al., 2022; Wei et al., 2022). Despite their methodological differences, all three benchmarking traditions largely presuppose that the objects subjected to computational evaluation have already been established before formal benchmarking begins. It is precisely this epistemic assumption – not benchmarking methodology itself – that the present discussion seeks to reconsider and challenge: the epistemic work of benchmarking begins at the moment when the entity to be observed and measured is constructed.
As understood in this paper, the construction of entities does not refer simply to the technical specification of computational objects. It concerns the epistemological process through which a discipline determines what constitutes a meaningful analytical unit before computational representation or evaluation becomes possible. Every humanities discipline approaches its objects of inquiry through its own conceptual traditions, and these traditions already distinguish particular features, relationships, and categories as analytically significant. Literary studies, one of the fields in which digital and computational methods have been extensively applied, provides an instructive example for the present discussion. Early quantitative literary scholarship, including Josephine Miles’s punch-card-assisted studies, had already treated features of poetic language – especially diction, syntax, verbal patterns, and broader poetic modes – as meaningful dimensions through which poetry could be interpreted (Miles, 1964; Buurma and Heffernan, 2018). These categories did not become analytical entities because they were computationally tractable. They became computationally tractable because literary scholarships had already recognised these features as analytically significant. Computational methods therefore inherit and deploy these entities in Miles’s early study; the methods did not produce the entities.
This account should not be taken as proposing a universal ontology of humanities entities, since different humanities disciplines construct different analytical entities because they ask different questions of their materials and are grounded in different epistemological traditions. The individuality of these entities, nevertheless, does not undermine the broader and more universal idea of entity construction. It demonstrates that entity construction is necessarily disciplinary, as it emerges from how a field understands its objects of inquiry before those objects are passed through the representational requirements of computational systems. The challenge concerning benchmarking in the humanities is therefore multi-step. A discipline must first identify meaningful analytical entities through tracing its disciplinary traditions, then translate these entities into operationalisable forms that remain congruent with its epistemic design, and finally determine whether and how the resulting representations can be evaluated appropriately. The common digital humanities challenge of translating complex interpretive objects into computationally representable and operationalisable forms (whether as tokens, annotated features, metadata, or task-specific labels) arises only after these disciplinary entities have been established (Bowker, 1999; McCarty, 2005; Drucker, 2011; Edmond and Lehmann, 2021).2
(2) Description
Existing scholarship in both digital humanities and computational research has already made substantial contributions to questions of data modelling, knowledge representation, dataset construction and documentation, annotation practices, and evaluation (e.g. Bender and Friedman, 2018; Flanders and Jannidis, 2018). The present collection builds upon these discussions by directing attention to an earlier epistemic stage where discipline-specific analytical entities are constituted before they are represented, modelled, and evaluated. Benchmarking is therefore understood as more than a methodology for assessing computational performance; it is also as an epistemic practice through which various disciplines determine what becomes a legitimate object of evaluation. Collectively, the contributions reposition benchmark design as a site where disciplinary epistemology and computational methodology become mutually constitutive, foregrounding entity construction as the epistemic precondition of benchmarking. They suggest that benchmarking in digital humanities is emerging as a layered practice spanning epistemic design, evaluative logic, and critical reflection. While the articles can be broadly grouped into three clusters, many address overlapping concerns that cut across these categories.
The first cluster of articles is concerned less with datasets themselves than with the epistemological processes through which humanities materials become constituted as data. Bringing together data papers and discussion papers, these contributions demonstrate that datasets are not neutral repositories of information but epistemic constructions that determine what becomes available for computational analysis and, ultimately, benchmarking. Before any metric can evaluate performance, the object of evaluation must first be defined, delimited, and rendered analytically meaningful. For instance, “A Comprehensive Digital Corpus of Song Ci Poetry for Computational Analysis” (Song) does more than assemble a large-scale corpus; by aggregating and standardising texts from multiple public-domain sources and structuring each poem through its author, rhythmic title (cipai), and machine-readable text, it reconstitutes Song Ci poetry as a computationally analysable literary object. Similarly, “Modernism’s Vector: Dataset for a Chinese-Language Journal in Late 1950s Malaya” (Wong) demonstrates how data curation can reshape literary-historical inquiry by transforming a historically situated journal into a structured computational object, thereby expanding what may legitimately count as literary data for computational analysis. In a more linguistically focused context, “NoorGhateh: A Benchmark Dataset for Training and Evaluating Arabic Morphological Analysis Systems” (Minaei-Bidgoli and AlShuhayeb) does not simply provide annotated data for evaluation; it renders complex morphological interpretation into explicitly defined analytical units, illustrating how linguistic knowledge must first be constituted as computationally tractable data before meaningful benchmarking becomes possible.
The second cluster continues to engage with the question of entities, but does so through a more explicit focus on benchmarking frameworks and evaluation design. These papers address the qualitative, multi-dimensional character of humanities objects by developing evaluation frameworks that move beyond single metrics and respond to the specific demands of domain, interpretation, and explainability. To begin with, Hutchinson’s “Benchmarking as Source Criticism: From Recognition to Reasoning in LLM Assessment” questions whether existing forms of LLM evaluation are capable of assessing historical reasoning rather than mere recognition-based pattern retrieval. By examining the limits of benchmark designs across open-ended reasoning tasks and global historical contexts, the paper argues that evaluation itself must be treated as a form of source criticism, extending benchmarking toward a more explicitly interpretive and epistemological mode of assessment. The model-agnostic framework proposed in Audsley’s “‘aSimMatrix’ Dimensions: A Scalable Framework for Benchmarking Intertextual Similarity” echoes this stance by decomposing intertextual similarity into multiple interpretive dimensions, highlighting the gap between computational detection and human judgment. Similarly, “A Multi-Dimensional Evaluation Framework for Assessing LLM Performance in TEI Encoding” (Strutz) introduces a stratified evaluation system that aligns benchmarking with the structural and interpretive complexity of TEI, while distinguishing between aspects that can be automated and those requiring human interpretation. In the linguistic context, “Domain Sensitivity in Arabic Morphological Analysis: A Multi-Corpus Evaluation of Farasa, CAMeL, and ALP Across Modern, Classical Religious, and Classical Jurisprudential Domains” (Minaei-Bidgoli et al.) demonstrates how benchmarking outcomes vary significantly across corpora, underscoring the importance of domain-sensitive evaluation; whereas “The Linguistic Asymmetry Index (LAI): Benchmarking Equity in Multilingual Research Infrastructures” (Spence and Battaner) suggests a further reorientation of benchmarking: it is mobilized not only to measure performance, but also to diagnose structural inequities within research infrastructures, pushing benchmarking beyond an optimization tool toward a mode of critical inquiry shaped by humanities concerns.
A third cluster builds on the complexity and nuances of benchmarking design in the humanities by placing emphasis on the review, design, and reflection of the epistemic chain through which benchmarking becomes embedded in project workflows. Benchmarking is no longer treated purely as a discrete evaluative step, but as a process shaped from the outset by conceptual and interpretive decisions and requiring the integration of entity design, method, and evaluation. “A Classification Benchmark Based on the Literary Theme Ontology” (Visser Solissa et al.) demonstrates how abstract interpretive categories such as literary themes must be defined through ontology design before they can become measurable, thereby grounding benchmarking in an explicitly ontological framework. “From Character to Poem: Nested Contexts and Scalar Limits of Parallelism Detection in Classical Chinese Poetry” (Kurzynski) similarly shows that benchmarking outcomes depend on the alignment between the computational unit of analysis and the humanistic unit of inquiry, highlighting how different scales of representation produce different evaluative conditions. This design logic is also evident in “Evaluating English-Korean Literary Machine Translations: A Dataset Featuring the RULER and VERSE Annotation Methods” (Shafayat et al.), where literary translation quality is made benchmarkable through the prior design of annotation schemes and paragraph-level alignment. In this case, the representational categories through which literary quality becomes comparable are constructed before evaluation takes place. The two RISE Humanities Benchmark papers offer an alternative example of this workflow-oriented view. “From Experiments to Epistemic Practice: The RISE Humanities Data Benchmark” (Hindermann et al.) frames benchmarking as an epistemic practice derived from digital humanities research support, while “The RISE Humanities Data Benchmark: A Framework for Evaluating Large Language Models for Humanities Tasks” (Hindermann et al.) presents benchmarking as a structured configuration of datasets, tasks, model settings, and evaluation procedures. These contributions provide evidence that benchmarking is not the last mile of a project, but a process through which humanities epistemology can inform entity construction, methodological design, and evaluative criteria from the beginning.
Several workflow-oriented contributions drawn from ongoing humanities projects further illuminate how entity design shapes and anchors evaluative benchmarking practices. This workflow-oriented perspective is evident in “Digitising Death: Benchmarking Genealogical Data and Recovering Women’s Histories in Early Modern Ireland” (McShane et al.), where benchmarking is embedded within the process of digitising and recovering genealogical records, situating the evaluative method within broader humanities concerns surrounding archival visibility, recovery, and historiographical interpretation. Nguyen and Alvarado Rojas’s “Exploratory Computation in Digital Humanities: A Qualitative Evaluation Framework” proposes a more cautious evaluative framework based on the argument that exploratory computational work in the humanities requires forms of qualitative evaluation that cannot be reduced to conventional performance metrics alone. The paper serves as a reminder that benchmarking is an iterative and reflexive practice shaped by interpretive context, positionality, and research objectives. McCarthy’s “Are We There Yet? Notes Towards Benchmarking an Experimental AI-Assisted Workflow for Humanities Data Cleaning and Reconciliation” (McCarthy) testifies to the often retrospective nature of benchmarking design in digital humanities. By demonstrating how benchmarking often emerges through iterative experimentation, infrastructural limitations, and project management decisions, it offers a timely reflection and possible way forward on the practical constraints of AI-assisted reconciliation and cleaning within the STEMMA project.
(3) Discussion and Conclusion
Across the contributions to this special collection, a shared orientation becomes visible: benchmarking is increasingly understood as more than a framework for evaluating computational performance. It is an evaluative system built upon entity construction, which constitutes its epistemic precondition. In this regard, benchmarking does not begin with evaluation metrics alone. It depends on earlier disciplinary decisions concerning what constitutes a meaningful analytical entity, how that entity is represented, and, ultimately, how it should be evaluated with the rigour associated with computational benchmarking. Foregrounding carefully designed entity construction brings benchmarking forward into the early stages of research and workflow design. In this light, benchmarking is no longer a retrospective procedure applied after research design has been completed, but an iterative process intertwined with entity construction, interpretive framing, and workflow development. The design of the workflow itself demands clarification of how the identified entity can be effectively measured and by what means. Collectively, the contributions shift attention from benchmarking as a mechanism of measurement towards an epistemic design of humanities research that is congruent with the core concerns and disciplinary traditions of the respective fields.
The significance of this mutually reinforcing relationship between discipline-specific entity construction and benchmarking lies not only in reconciling what can be measured with how we know that the measure is epistemically valid and methodologically rigorous. Once benchmark design is understood as part of the epistemic process, it requires explicit reflection on the disciplinary assumptions that determine what becomes an object of evaluation in the first place. Digital humanities benchmarking, understood in this sense, addresses the common difficulty of feeling compelled either to reject scientific models or to adapt them without knowing where to begin or what should be changed, all the while attempting to preserve the rigorous scaffold of existing scientific benchmarking practice. It complements existing benchmarking methodologies by grounding benchmark design within the epistemological commitments of individual humanities disciplines. Scientific benchmarking continues to provide robust principles for systematic comparison, transparency, and reproducibility; benchmarking for digital humanities ensures that these methodological strengths remain accountable to the interpretive questions and disciplinary understandings from which humanities research begins.
Finally, this special collection has implications for the ways in which future datasets, benchmarks, and research workflows are conceived across the humanities. It also calls for sustained engagement with existing disciplinary traditions, since these traditions offer insights into what has historically been regarded as a meaningful unit of analysis within their respective fields. By offering a conceptual foundation for thinking about benchmark design as an epistemic practice, future work may build upon this discussion by examining how different humanities disciplines have historically constructed analytical entities, how these disciplinary differences can shape benchmark design, and how benchmarking may support more transparent, reusable, and theoretically grounded digital humanities methodologies.
Notes
[1] A useful point of reference here is the Cranfield paradigm in early information retrieval, often taken as the starting point for modern benchmarking practices. As Rodríguez and Boyd-Graber (2021) recount, the key move was deceptively simple: instead of evaluating systems through direct user interaction, researchers began to “build re-usable test collections and evaluate all systems by re-using the same collection,” allowing different approaches to be compared on shared, pre-defined tasks. In this sense, benchmarking, from its earliest formulation, is already grounded in the assumption that evaluation depends on standardised tasks and comparable performance conditions, an assumption that continues to shape later developments in computer science and NLP. See Cleverdon, Cyril. “The Cranfield Tests on Index Language Devices.” Aslib Proceedings 19.6 (1967): 173–194.
[2] This is not to suggest that scientific practice is unaware of the simplifications involved. As Bowker (1999) makes explicit, classification systems and standards make objects appear stable and comparable precisely by suppressing their material and contextual complexity.
Acknowledgements
This special collection grew out of a seminar hosted by the Computational Humanities Research Group at King’s College London. I am grateful to the organisers, participants, and fellow editors for the discussions that first brought its central questions into focus.
The broader epistemological framework developed here was informed by conversations during my Visiting Fellowship at Cambridge Digital Humanities. I am also grateful to Jesus College, University of Oxford, for meal provision, access to College facilities, and the collegiate environment that supported the development of this manuscript. Interdisciplinary conversations there, particularly with colleagues across several scientific disciplines, helped refine the articulation of the relationship between disciplinary epistemology, entity construction, and benchmarking methodology.
