Introduction
The landscape of participatory science within cultural-heritage institutions has undergone significant transformation, with Galleries, Libraries, Archives, and Museums (GLAM) increasingly adopting crowdsourcing methodologies to democratise knowledge production and enhance collection documentation (Owens 2013; Ridge 2013). While existing research has extensively examined engagement design—how institutions structure public-participation activities (Simon 2010; Oomen and Aroyo 2011; Owens 2012; Bonacchi et al. 2019)—and scholars have documented core integration challenges (Jansson 2017; Van Hyning 2019; Ridge et al. 2021b; Rinaldo et al. 2023), comprehensive analysis of how GLAMs actually process, validate, and absorb public contributions into institutional memory remains less fully developed.
This scholarly imbalance is particularly striking given that integration pathways represent one of the most pressing unresolved questions in participatory-heritage practice. Scholars have observed data-integration barriers since at least 2013, when Carletti et al. explicitly identified this as “the open challenge for crowdsourcing in the digital humanities,” noting that “the resources contributed by the crowd are not incorporated into the institutional collections” in many participatory initiatives. Subsequent research has documented these barriers’ many forms. Van Hyning and Jones (2021) examined the substantial labour required to ingest transcribed content into repositories, highlighting the prohibitively resource-intensive nature of data cleaning and assimilation. Most tellingly, Jansson’s (2017) analysis of crowdsourcing integration across Swedish GLAMs demonstrated that, even when professionals valued user contributions, collections management systems (CMSs) lacked appropriate fields for their incorporation, leading to systematic exclusion of community knowledge despite institutional goodwill. The obstacles are not merely technical, but architectural: without explicit attention to integration infrastructure, participatory projects risk perpetuating the very hierarchies they purport to dismantle—soliciting community contribution while systemic constraints preserve institutional control (Feinberg 2007; Yakel 2011; Thomer and Rayburn 2024). At stake in these integration processes is, therefore, epistemic authority itself—who holds the power to name what counts as knowledge, and until that question is confronted, whose contributions will continue to fall short of legitimacy.
To trace how these infrastructural barriers shape integration outcomes and authority distributions, this study adopts a comparative approach, examining how heritage institutions and citizen science projects attempt to negotiate the passage from crowd to collections and datasets. Heritage projects in the sample focus primarily on collection-data enhancement, transcription, and cataloguing; the citizen science initiatives focus on biodiversity observation and environmental monitoring. The contribution types they produce (transcriptions, metadata enrichment, and observational classification) are distinct enough to require different validation approaches. That said, while some institutions—particularly natural history museums—engage in both heritage documentation and scientific research, this study distinguishes practices by primary epistemological goals: heritage practice (cultural stewardship, preservation, and interpretation) versus scientific data generation and hypothesis testing via citizen science, which focuses on “the involvement of the public in scientific research” (Ridge et al. 2021a). This framing enables examination of transferable quality-assurance methods while recognising substantive operational differences. What follows interrogates both.
Literature Review
Heritage crowdsourcing and the integration gap
Heritage crowdsourcing emerged from converging forces: efforts to challenge traditional museum authority, practical needs for enhanced collection documentation, resource constraints, and expanding possibilities for public engagement (Simon 2010; Oomen and Aroyo 2011; Ridge 2014; Terras 2016). Early scholarship was largely celebratory—social tags surfacing community perspectives (Light and Hyry 2002; Van Hooland 2006; Anderson and Allen 2009; Yakel 2011), transcriptions enabling text access, initiatives like Steve.museum (formed 2005) and Your Paintings Tagger (launched 2011) demonstrating practical viability at considerable scale. Even so, as these projects proliferated, fundamental questions about how public contributions translate from platforms to collections remained stubbornly unresolved.
Technical and infrastructural constraints to user-generated-content integration
Heritage catalogues typically cannot accommodate the granularity or format of crowdsourced outputs—full-text transcriptions and free-text tags align poorly with systems architected for controlled vocabularies and item-level description (Ridge et al. 2021b). Crowe et al. (2021) identified what successful user-generated content (UGC) incorporation would require, emphasising the importance of attribution systems and systematic workflows, yet, practice has not always kept pace. When institutions employed external platforms like Flickr or Tumblr for crowdsourcing, lack of integration with internal CMSs created acute data management problems: updating collections required complete re-export of data, potentially erasing user voices not yet incorporated (Peccatte 2011; Jansson 2017). Validation mechanisms—whether peer-control or professional moderation—amplified these challenges by demanding sophisticated information architecture or continuous institutional surveillance that few organisations could sustain (Jansson 2017).
Institutional dynamics and resource realities
Owens (2014) offered critical perspectives on these structural limitations, cautioning against what might be termed “crowdsourcing without the crowd”: extractive models that solicit volunteer labour, while maintaining contributions peripheral to authoritative collections, instrumentalising participation without substantive integration. The question this raises echoes broader concerns. Does participatory infrastructure genuinely redistribute epistemic authority, or does it merely harvest volunteer labour while institutional gatekeeping remains intact (Yakel 2011)? Owens’ constructive alternative opens a different path: Successful projects create opportunities for meaningful engagement, not extraction, where data improvements emerge organically from participants’ purposeful exploration.
Still, the argument runs deeper. Decolonial scholarship frames epistemic authority as inherently political, demanding institutions confront how curatorial authority has historically silenced marginalised voices (Smith 2006; McCracken and Hogan-Stacey 2023). Eveleigh (2014) traces the resulting tensions through professional practice, revealing institutional fears that contributors might err, and anxieties about ceding control over authoritative descriptions—that unresolved pull between custodial instincts and the desire to share access (see also Liew 2016).
These tensions are compounded by resource constraints. Thomer and Rayburn (2024) reframed integration not as discrete implementation but as labour-intensive continuous practice, demanding staff time, technical expertise, and ongoing institutional investment that many organisations struggle to maintain. Underfunded collection staff often find themselves “patching together” database systems to accommodate evolving requirements, maintenance work that proves untenable without dedicated resources.
Nor are these challenges unique to heritage contexts. Participatory-science research has long documented concerns about power dynamics and knowledge validation across domains (Wynne 1992; Irwin 1995). More recently, Vessio et al. (2025) identify ongoing accuracy problems in public-generated species observations, including recurring inconsistencies in data collection and identification errors; Esguerra and van der Hel (2021) reveal how participatory-knowledge platforms often preserve existing power hierarchies by maintaining scientific autonomy and favouring elite institutional relationships. Seen in this light, integration barriers reflect fundamental questions about epistemic authority, not technical limitations alone.
Heritage-specific validation challenges
Transcription work exemplifies heritage context’s distinctive validation requirements: Volunteers must combine palaeographic accuracy, historical verification, and interpretive engagement with historical narratives—multi-dimensional validation extending beyond the accuracy-focused criteria characterising discrete scientific classification tasks. Community-sovereignty considerations complicate validation frameworks further: Indigenous scholarship emphasises knowledge systems operating under relational accountability principles and multiple ways of knowing that challenge Western institutional paradigms (Tuhiwai Smith 1999; Srinivasan et al. 2009; Christen 2015). Such considerations demand validation approaches that maintain institutional accountability while honouring democratic participation principles—requirements fundamentally different from consensus-based protocols predominating in scientific domains.
Citizen science quality assurance: lessons for heritage practice
Although integration challenges persist across participatory-science contexts, citizen science has developed sophisticated quality-assurance frameworks in response to scientific-rigour requirements, including consensus-based validation, expert review protocols, and tiered verification systems (Wiggins et al. 2011; Allahbakhsh et al. 2013; Kosmala et al. 2016; Kaldeli et al. 2021; Sharma et al. 2022), that could potentially inform heritage institutions grappling with integration questions. Projects like Galaxy Zoo and Snapshot Serengeti demonstrate that consensus-based vote aggregation, where non-expert classifications are statistically combined, can achieve accuracy levels rivalling expert assessment, while delivering educational benefits (Lintott et al. 2008; Swanson et al. 2016). These successes, then, raise a compelling question: Can such approaches effectively transfer to heritage contexts?
Not easily. Oomen and Aroyo (2011) highlight the difficulties heritage-volunteer collaboration faces in locating skilled participants and maintaining quality control, with project success contingent on careful design and implementation (Noordegraaf, Bartholomew, and Eveleigh 2014). More fundamentally, as discussed earlier, heritage validation operates through provenance assessment, contextual interpretation, and cultural appropriateness—criteria resisting the standardisation characterising scientific-observation tasks, and requiring attention to community sovereignty alongside accuracy (Tuhiwai Smith 1999; Christen 2015).
Platform-adaptation research demonstrates these incompatibilities concretely, while also documenting sustained efforts to bridge disciplinary divides. Blickhan et al. (2019) compared Zooniverse’s standard independent classification method—developed for STEM projects—against collaborative transcription where volunteers could view others’ work. Collaborative approaches produced significantly richer data, prompting platform redesign specifically for heritage projects (Van Hyning 2019). This finding carries important theoretical weight: Transcription tasks involving full-text processing and contextual interpretation operate under markedly different constraints than discrete classification tasks such as galaxy morphology or species identification, and STEM methodologies require considerable modification accordingly.
Even purpose-built solutions encounter such barriers. Zooniverse’s heritage-specific development—beginning around 2014—demonstrates both cross-domain adaptation complexity and the sustained technical investment it demands, with sophisticated aggregation algorithms (Krawczyk 2018) and ALICE (Aggregate Line Inspector and Collaborative Editor), which addressed heritage institutions’ limited capacity to process technical data formats by providing simplified exports for teams without specialist expertise. Yet, these tools still require substantial institutional capacity. Data cleaning, staff labour, ongoing system maintenance—none of these yielded to better software alone (Van Hyning and Jones 2021; Thomer and Rayburn 2024). Collectively, that even well-resourced platforms with dedicated technical teams required significant infrastructural adaptation is, in itself, telling—and points to integration mechanisms as critical yet underexplored research terrain.
Methodology
This investigation employed a convergent parallel mixed-methods design to examine how institutions incorporate public contributions. The approach collects quantitative and qualitative data simultaneously through targeted surveys (capturing institutional patterns) and semi-structured interviews (documenting practitioner experiences), then integrates findings to enable systematic comparison across methodological approaches and participant groups (Guest and Fleming 2015). This paper draws on data gathered as part of a broader investigation examining participatory methodologies across heritage and citizen science contexts, though it focuses specifically on integration approaches.
Data collection occurred between August 2020 and December 2022, encompassing three integrated phases: (1) targeted surveys across heritage and citizen science contexts, (2) semi-structured interviews with experienced practitioners, and (3) comparative analysis applying concepts from democratic-infrastructure theory. This timeframe captured both established practices and institutional adaptations during accelerated digitalisation.
Sampling strategy
Sampling prioritised theoretical saturation—the point where additional data yields no new theoretical insights—over statistical representativeness, following established principles for theory-building research in organisational studies (Eisenhardt 1989; Yin 2018). The multi-stage approach enabled systematic comparison across institutional contexts while respecting domain-specific vocabularies and operational constraints.
The survey component focuses on United Kingdom (UK) museums, examining crowdsourcing practices across diverse institutional types, ranging from national and art museums to local-authority museums and historic houses. This sector-specific focus reflects museums’ distinctive collections-management frameworks. Interview data provide international and cross-GLAM breadth, capturing integration mechanisms across museums, libraries, and archives in multiple national contexts. Given significant differences in descriptive paradigms across institutional types, and variations in national funding models and regulatory frameworks, comprehensive international comparative analysis would require separate investigation beyond this study’s scope. This sampling architecture enables in-depth museum-sector analysis while capturing broader integration patterns through comparative interview data.
Survey methodology
Three targeted surveys gathered perspectives from distinct professional communities, each designed to facilitate cross-domain comparison while respecting community-specific vocabularies. Streamlined instruments are provided in Supplemental files 1–3: Appendices A–C.
Heritage-professionals survey (n = 49, 22.6% response rate)
A 39-item questionnaire examined institutional characteristics, crowdsourcing practices, content validation approaches, and integration preferences (Supplemental file 1: Appendix A). Sampling drew on both accredited and non-accredited UK museums through the Museums Association directory. Respondents demonstrated substantial professional experience (56.5% with > 10 years), spanning curatorial (22%), collections management (11%), and documentation (11%) roles.
Collections management systems specialists survey (n = 7, 38.8% response rate)
A shorter 17-item survey (Supplemental file 2: Appendix B) explored technical-integration possibilities across database-support mechanisms and validation frameworks. Purposive selection targeted professionals affiliated with Collections Trust–validated companies (Collections Trust n.d.), ensuring expertise in and compliance with SPECTRUM, the UK’s widely adopted standard for museum collections management procedures (Collections Trust 2017).
Citizen science coordinators survey (n = 8, 6.1% response rate)
The most extensive instrument, a 42-item questionnaire (Supplemental file 3: Appendix C), addressed quality-assurance practices, community-management strategies, and data-processing methods. Active projects were identified through established platforms (SciStarter, Zooniverse) combined with direct outreach to coordinators located through academic literature.
Interview methodology
Semi-structured interviews (Supplemental file 4: Appendix D) elicited practitioner perspectives on institutional-integration approaches, drawing on organisational ethnography (Spradley 1979) and institutional analysis (Scott 2014).
Heritage professionals (n = 6)
Theoretical sampling prioritised practitioners with documented crowdsourcing experience, seeking diversity across institutional types, scales (five larger institutions, one small organisation), geographic contexts (UK, United States [US], Spain), and project types.
Citizen science coordinators (n = 6)
A parallel approach targeted established coordinators representing diverse scientific domains and project scales across four countries (US, UK, France, Austria), enabling comparison of quality-assurance methods.
Heritage professionals are identified as Interviewee 1.1–1.6, citizen science coordinators as Interviewee 2.1–2.6.
Interviews were conducted via Microsoft Teams between September and December 2022 (25–53 minutes), audio-recorded with explicit consent, and researcher-transcribed. The protocol employed a funnel structure progressing from broad contextual questions to specific technical details.
Data analysis
Quantitative analysis
Survey data underwent descriptive statistical analysis focused on pattern identification, employing cross-tabulation and correlation analysis to trace relationships between institutional characteristics and integration preferences.
Qualitative analysis
Interview transcripts were analysed using iterative-thematic coding via the Saturate App, a web-based qualitative analysis tool. The coding scheme was developed both deductively (from research questions and literature) and inductively (from emerging patterns in the data). Initial codes were applied to early transcripts and refined iteratively as analysis progressed to capture participant-specific vocabulary and emerging themes. Codes were systematically applied across all interviews, with exemplary quotes selected to illustrate key findings. The final coding scheme encompassed 110 codes organised into thematic categories including: quality-assurance methods, data-processing approaches, professional-development needs, resource constraints, and technical infrastructure barriers.
Integration and triangulation
Data integration followed established principles for convergent parallel design (Creswell and Plano Clark 2017), with quantitative and qualitative findings analysed independently before structured comparison and synthesis. The triangulation process operated through three phases: (1) survey-interview comparison mapping quantitative-preference patterns against qualitative themes; (2) cross-cohort analysis comparing heritage and citizen science approaches to identify shared challenges and sector-specific variations; (3) thematic integration synthesising findings across all data sources, tracing emergent institutional-integration approaches.
Triangulation occurred across methodological (surveys and interviews), participatory (heritage and citizen science domains), and temporal (2020–2022) dimensions, enhancing validity through multiple verification pathways while accommodating the exploratory nature of theory-building research.
Results
Heritage-professional survey: integration-preference patterns
Survey responses from heritage professionals (n = 49, 22.6% response rate) exhibited limited engagement with integration planning. Only 40.8% (n = 20) addressed integration-methodology questions, with fewer still (n = 11) providing specific implementation details. Analysis of these responses revealed three distinct patterns: recording into main database (46.1%, n = 6), validation from professional staff (30.7%, n = 3), and incorporation in places other than collections management databases (15.3%, n = 2).
Institutional-capacity correlations
Financial-resource analysis exposed consistent relationships between budgetary constraints and integration preferences. Of those reporting budget information (n = 19, 39% of total sample), organisations with annual budgets exceeding £250,000 constituted just 26% (n = 5) of respondents. The sample spanned diverse implementation contexts across independent museums (28%, n = 14), local-history museums (22%, n = 11), and local-authority museums (18%, n = 9).
Implementation confidence and barriers
Only 40% (n = 8) of heritage professionals addressing workflow-integration questions expressed confidence in their institution’s implementation capacity, yet most (88%, n = 22) recognised crowdsourcing’s transformative potential. Q7 responses (n = 25; see Supplemental file 1: Appendix A) in particular made clear that, while professionals overwhelmingly endorsed audience-created content’s capacity to increase engagement (88%), foster accessibility (84%), and diversify content (80%), they were equally emphatic (88%) that such content requires continuous professional maintenance, despite only 32% considering it potentially inaccurate. Practitioners nonetheless highlighted significant obstacles, including financial constraints, staffing limitations exacerbated by COVID-19 redundancies, and regulatory concerns.
Technical infrastructure: implementation-feasibility analysis
Among seven responding SPECTRUM-compliant CMS vendors (38.8% response rate), three reported current system support for UGC integration. These systems distinguish user contributions from professional documentation through separate data fields, temporal metadata (contributor name, date, time), and labelled public-comment sections.
All three require manual professional validation before integrating content into existing catalogue records. Of these vendors, two capture contribution provenance, while only one prominently displays contributor attribution alongside submission details.
Citizen science survey: quality-assurance approaches
Citizen science coordinators (n = 8, 6.1% response rate) reported rigorous quality-assurance approaches. Among initiatives implementing quality control (n = 4 of 5 responding to Q5; Supplemental file 3: Appendix C), multiple validation methods were employed simultaneously: staff moderation (n = 3), peer-review tools (n = 3), automatic checks (n = 2), and controlled vocabularies (n = 2). Seven of eight projects reported contributions meeting project objectives, with 80% of total respondents characterising UGC as “very valuable” for scientific research.
Interview findings: heritage-professional perspectives
The sample comprised three museums, two libraries, and one combined library/archive, spanning five larger institutions and one small organisation across the UK (n = 3), US (n = 2), and Spain (n = 1). Practitioners predominantly relied on existing external platforms for crowdsourcing activities, including Zooniverse, FromThePage, and proprietary software, with only one institution developing a bespoke in-house solution.
Quality-assurance strategies and scalability
For quality-assurance, heritage practitioners implemented various strategies: user training and support (33%, n = 2), meticulous review processes (17%, n = 1), and clear guidance provision (33%, n = 2). Scalability, however, proved a recurring pressure point, as Interviewee 1.2 (national museum, distributed newspaper research) described:
My responsibility was to review the material, quality control, and help the volunteers, and the team … over time, it just became unmanageable. So, we had a team of volunteers. We found that volunteer reviewers were helpful, but we really needed something even more reliable. So, we started to pay contractors.
This shift to paid contractors for quality control, while uncommon in heritage crowdsourcing, illustrates the resource pressures institutions face when volunteer participation exceeds validation capacity.
Collaborative-validation frameworks, wherein volunteers review and refine each other’s contributions, meanwhile, offered an alternative. Interviewee 1.4 (coordinating transcription projects across multiple libraries and archives) reported: “We don’t do double-blind or blind double keying; everything is keyed, reviewed, edited, and approved collaboratively. And all of that is also surfaced on the activity stream.”
Infrastructure barriers
Heritage respondents (83%, n = 5) flagged attribution concerns from both technical and ethical perspectives. Interviewee 1.4 articulated the infrastructural root:
One of the challenges that we discovered … was having a place to put that data right. Archival systems generally do not have places in their metadata for who worked on a record. If the digital-library systems don’t have a place to put contributors, there’s not much we can do, and this is I think an ongoing challenge.
Additional barriers encompassed copyright issues (33%, n = 2), contributor-crediting difficulties (67%, n = 4), and data-protection compliance (33%, n = 2). Interviewee 1.2 elaborated on attribution complexity:
For us, in particular, there’s this big question about newspaper copyright and do we have the right to display it. And then when the research is submitted, we give the first name and last initial credit to the user. But when we cite it like in historical work, who gets the sort of credit for that too? It’s a little bit ambiguous.
Workflow-integration requirements
Practitioners delineated comprehensive technical and institutional requirements for integration. Interviewee 1.3 (national library, archive transcription) noted:
I think at the end … the issue is how do you get that data? What do you do with that data? … Do you have means automating the way you put that back into the catalogue? And the other aspect is the technical expertise … Making crowdsourcing kind of business as usual, part of the plan, rather than just an ad-hoc bespoke initiative.
Institutional-capacity constraints proved equally significant. Interviewee 1.5 (national library, spatial-metadata project) observed:
When cataloguing large volumes of data, institutions face time constraints and may lack the necessary skills within the cataloguing community … And setting up a routine workflow that does not just rely on the goodwill of staff within institutions is harder than just doing this ad-hoc.
Timeline pressures were further emphasised by Interviewee 1.2:
We had less than two and a half years’ lead time, that wasn’t long enough … If the exhibition was opening now, seven years into it, we would have had a much better, bigger, richer data set to be able to contribute to.
Yet, advanced implementation necessitated inter-professional collaboration. Interviewee 1.6 (museum consortium, image description and metadata enrichment) stressed: “I think there has to be better feedback between cultural-heritage and technical professionals … The collaboration can be much more fluid and there could be co-creation at different levels.”
Interview findings: citizen science coordinator perspectives
Interviews with citizen science project managers (n = 6) revealed validation architectures and processing capabilities across diverse scientific domains. In contrast to heritage practitioners, most coordinators (5 of 6) built custom tools in-house, integrating purpose-built applications with existing recording infrastructure.
Hybrid validation and distributed expert networks
Most practitioners (67%, n = 4) endorsed combined automated and manual validation methods. In this context, Interviewee 2.1 (wildlife-monitoring project) outlined a layered approach:
For environmental monitoring, where the volunteers were really involved, we had personal contact, oversight, and quality checks. We would also get some of the volunteers to check through and make sure submissions looked correct. For records submitted by volunteers we had never met, the iRecord website provided built-in quality control by county recorders—these biological recorders were stationed all over the country.
This distributed validation architecture demonstrates scalability through structured delegation of validation responsibilities to trained community members while maintaining centralised quality standards.
Data-processing integration
All citizen science interviewees (n = 6) employed digital-processing tools. Quality-control approaches ranged from traditional statistical methods (Excel, SPSS) and automated validation using machine-learning processes (R, Python) to sophisticated interoperability frameworks integrating multiple external data sources. Citizen science coordinator 2.6, exemplifying the latter category with a project combining volunteer-submitted snow-depth measurements with satellite-remote sensing and automated telemetry data, explained:
Well, it’s set up with an API, right? We ingest the data from all these different platforms, and that data is just downloaded as a text file. … And then those [data] points are integrated into a model … It’s taking in various input parameters to come up with snow depth and then snow-water equivalent, and it will produce a snow-depth result.
Discussion
This study addresses a critical gap in participatory-science theory by examining institutional-integration mechanisms as sites where epistemic-authority is negotiated, contested, and reproduced. Through convergent analysis of survey and interview data, the research identifies three distinct integration approaches—peripheral contribution (15.3% preference), supervised collaboration (30.7% preference), and workflow integration (46.1% preference)—each reflecting different institutional strategies for managing epistemic authority in participatory contexts.
That only 40% of heritage professionals express confidence in workflow-integration capabilities, despite 88% recognising crowdsourcing’s transformative potential, reveals implementation gaps reflecting both technical and epistemic challenges. As Q7 responses demonstrate, professionals expressing little anxiety about inaccuracy still insist on continuous professional oversight; the operative concern, then, is institutional mediation, not data quality. The validation criteria heritage settings demand—provenance, contextual interpretation, and cultural appropriateness—complicate the picture further, resisting the standardisation that enables scalable integration regardless of institutional setting. What emerges extends theoretical frameworks around participatory-knowledge production (Simon 2010; Ridge 2014) by identifying integration infrastructure as the critical terrain wherein democratic-participation values encounter—and often yield to—institutional-authority structures. Technical incorporation and epistemic redistribution are, in other words, distinct challenges: one of capacity, the other of will. The sections that follow trace this distinction through the empirical evidence.
Epistemic authority and power dynamics
The three integration approaches listed above differ in how institutions exercise epistemic authority, not whether they retain it (Figure 1). While Haklay’s (2013) widely cited participation scale maps increasing citizen involvement in problem definition and data interpretation, heritage crowdsourcing integration mechanisms operate along an orthogonal dimension entirely: the locus of validation authority. Peripheral contribution externalises validation through spatial separation; supervised collaboration centralises it through hierarchical gatekeeping; workflow integration embeds contributions within institutional systems while retaining gatekeeping functions. Importantly, movement along this operational spectrum does not correlate with epistemic-authority redistribution. Even workflow integration, which would achieve full incorporation into authoritative records, preserves institutional control over what knowledge enters collections, how it is contextualised, and whose claims receive institutional legitimation.

Figure 1
Three integration approaches differentiated by the locus of validation authority: peripheral contribution, supervised collaboration, and workflow integration.
It is this decoupling of technical integration from epistemic parity—wherein contributor and institutional-knowledge claims would carry equivalent authoritative weight within collections—that resonates most strongly with Esguerra and van der Hel’s (2021) finding that participatory designs simultaneously open knowledge-creation processes while reinforcing institutional-power structures. All three approaches cluster at the same point on the epistemic spectrum: institutional retention of validation authority. True epistemic parity would require community involvement in validation-process design, shared publication decisions, and infrastructure enabling diverse epistemological frameworks (as exemplified by CARE standards for Indigenous data governance; Christen 2015)—none of which were observed in the institutions studied here.
Peripheral contribution operationalises what might be termed “institutional epistemic compartmentalisation”—building on existing scholarship (Srinivasan et al. 2009) to describe how heritage institutions separate public- and professional-knowledge contributions through technological design, extracting participatory content while preserving traditional-authority structures. One survey respondent’s reflection exemplifies this limitation:
The use of a UGC module that allows data to be held alongside core information in a separate ‘middleware’ product seems like a possible solution that allows us to add and share new and rich data without undermining core collections-management data and functionality.
It is this very architectural separation that enacts epistemic hierarchy: physical distance within database structures becomes the mechanism through which institutions assert authority over knowledge legitimacy.
Supervised collaboration increases participatory breadth while maintaining professional gatekeeping, integrating UGC into authority records but subsuming contributions within institutional voice, without crediting their source. As Jansson (2017) and Ridge et al. (2021b) demonstrate, this institutional oversight is reinforced by CMS design: Lacking dedicated fields for diverse user contributions, systems force professionals to selectively incorporate only information that aligns with existing metadata schemas and institutional-credibility requirements. Interview evidence uncovered the scalability costs—progressing from personal review to volunteers to paid contractors (Interviewee 1.2)—laying bare validation as a resource-intensive bottleneck wherever universal professional mediation is mandated.
Workflow integration distributes gatekeeping across institutional processes without eliminating it. The implementation barriers interview participants identified—attribution infrastructure limitations, legal frameworks, professional-training requirements—indicate that genuine workflow integration requires substantial process redesign: validation embedded throughout institutional systems and contributor attribution built-in from the outset. As Eveleigh (2014) argues, participation often functions pedagogically, with volunteers learning professional norms and terminology through engagement; the challenge, then, lies in designing scalable validation processes that acknowledge educational value while building the attribution infrastructure that would enable genuine integration.
Cross-disciplinary methodological transfer
Comparative analysis demonstrates both opportunities and constraints for methodological transfer. Although citizen science has developed sophisticated quality-assurance frameworks, recent scholarship reveals considerable variation, with projects “still in need of norms” around data-quality assessment (Borgman 2016; Bowser et al. 2020). Survey evidence highlights marked differences in validation conceptualisation: Citizen science coordinators expressed confidence in public contributions, rating crowdsourced content as highly valuable for scientific research. Heritage survey responses, however, struck a different note. One respondent, for instance, cautioned that “without moderation, the ramifications for a museum could be disastrous,” reflecting precisely those institutional anxieties around legal compliance, cultural protocols and provenance tracking that Eveleigh (2014) identified.
Nevertheless, certain citizen science approaches prove adaptable to heritage contexts. Distributed validation systems combining automated screening with trained volunteer networks demonstrate particular promise. Interviewee 2.1’s county-recorder systems—where trained biological recorders validate species observations across geographical regions—exemplify scalable quality control under institutional oversight. Heritage institutions show parallel evolution, if less systematically: Interviewee 1.2 described scaling from individual review to “a team of volunteers” and eventually “paid contractors” as volume became “unmanageable”; Interviewee 1.4 implemented collaborative validation where “everything is keyed, reviewed, edited, and approved collaboratively.” Whether consensus aggregation (Kosmala et al. 2016; Swanson et al. 2016) where multiple independent inputs are algorithmically combined, or the collaborative approaches described above, these models share a structural logic: They distribute validation labour while retaining institutional authority over final publication decisions, with raw volunteer data rarely published directly. Such approaches offer GLAMs pathways to maintain verification standards at scales where individual professional review proves untenable.
Interview evidence revealed universal digital-tool adoption among citizen science coordinators (R, Python, and machine learning), while heritage professionals showed more variable computational capability. Half employed sophisticated automated approaches (machine-learning methodologies, Python scripts, and natural-language-processing libraries), while others lacked such infrastructure, instead relying on manual, digitally mediated methods, such as page-by-page review of individual contributions (Interviewee 1.1) or manual back-end verification of submissions against source metadata (Interviewee 1.2). This variation reflects institutional resource constraints rather than inherent domain limitations, underscoring that successful integration requires sustained investment in technical capacity. But the gap is not only one of resourcing. Heritage practitioners’ reliance on external crowdsourcing platforms, which sit entirely outside institutional CMS architectures, points to an integration problem that is fundamentally infrastructural. Citizen science coordinators, by contrast, built ingestion pathways into their tools from the outset. Where that design work happens upstream, then, incorporation tends to follow naturally; where it is retrofitted, the friction risks compounding.
Implementation capacity: resources and professional development
Interview data unpack the confidence gap survey findings identified. Practitioners describe persistent resource barriers including “temporal and budgetary demands” that force institutions to rely on “trusted, pre-trained volunteers” rather than broader public engagement (Interviewee 1.5). Among surveyed heritage professionals who provided budget information, 47% operate on budgets of £30,000 or less, while roughly a quarter exceed £250,000. While these findings cannot be generalised beyond the study sample, they align with broader sector evidence of financial constraints, with recent research indicating that heritage organisations face significant resource pressures (Heritage Alliance 2024).
Beyond financial limitations, interview evidence exposes fundamental shortfalls in professional preparation, particularly around technical infrastructure and cross-disciplinary collaboration. Interviewee 1.3 argued that successful crowdsourcing requires methodical institutional integration rather than isolated initiatives, emphasising both the need for automated data-processing capabilities and the essential human expertise required to extract, process, and incorporate crowdsourced content into institutional catalogues. This challenge is sharpened by capacity constraints: As Interviewee 1.5 observed, “institutions face time constraints and may lack the necessary skills within the cataloguing community,” explaining that “setting up a routine workflow that does not just rely on the goodwill of staff within institutions is harder than just doing this ad-hoc.”
Interviewee 1.6, in turn, identified enhanced interdisciplinary collaboration as essential, arguing that more fluid partnerships between cultural-heritage and technical professionals could enable co-creative approaches leveraging methodologies not traditionally used within heritage contexts. These converging constraints—temporal, financial, infrastructural, and skill-based—indicate that successful workflow integration requires ongoing institutional support well beyond technical capability: inter-professional collaboration, structured staff development, technical infrastructure, and transformation of workflows from ad-hoc initiatives into embedded institutional practice.
Attribution and infrastructure barriers
Attribution surfaced as a concern for five out of six interviewees (83%), spanning both technical and ethical dimensions, and cutting across all three integration approaches. As Interviewee 1.4 poignantly noted, archival systems lack designated fields for contributor attribution—an infrastructural absence the CMS-specialist survey confirmed across the sector. Among systems supporting user-generated content, only one prominently displays contributor attribution alongside submission details, while all require manual integration of audience-created content with existing catalogue records, creating bottlenecks constraining routine workflow implementation. This infrastructural gap exposes a critical dimension of the operational-systemic divide: Even workflow integration, which achieves technical incorporation into authoritative records, risks architecturally obscuring contributor identity—rendering labour visible as knowledge while contributors remain effaced as knowledge producers.
These technical limitations compound across institutional processes. Interviewee 1.2 described uncertainty beginning with fundamental copyright questions, extending through inconsistent contributor-recognition practices, and culminating in ambiguous protocols for citing crowdsourced research in scholarly work. Interviewed practitioners responded through workarounds—content-sharing agreements, archive-location tracking, data-protection compliance—yet, such ad-hoc strategies expose rather than resolve systemic infrastructure gaps. Attribution infrastructure, seen in this light, exemplifies how capacity barriers (technical systems, institutional protocols) intersect with epistemic structures (whose knowledge receives formal recognition). Addressing it requires not merely database fields but a reconceptualisation of contributors as knowledge producers warranting epistemic standing within institutional records.
Methodological Contributions
Comparative analysis identified shared challenges and sector-specific variations across heritage and citizen science contexts. The convergent parallel design allowed quantitative and qualitative findings to be tested against each other before integration, strengthening interpretative rigour across a small but diverse sample.
Limitations and Future Research
Several methodological constraints affect the generalisability of these findings. The geographic concentration on UK museums in survey sampling and reliance on English-language contexts limit broader applicability, potentially obscuring significant cultural variations in crowdsourcing practices. While international interview sampling provides some comparative perspective, the institutional emphasis constrains direct evidence about participant experiences and how different integration approaches affect volunteer engagement. The low citizen science coordinator response rate (6.1%) presents a further constraint on statistical inference, though triangulation with qualitative data partly offsets this limitation.
The cross-sectional nature of the study provides a snapshot of crowdsourcing practices but may not capture longitudinal trends or evolving institutional perceptions. The timing of initial surveys during the COVID-19 pandemic introduces contextual variables, as documented studies have shown the pandemic’s impact on museum-workforce dynamics, digital-engagement patterns, and institutional capacity for public engagement (Agostino et al. 2020; Finnis and Kennedy 2020; Merritt 2021). Although subsequent 2022 interviews captured institutional learning from varied temporal relationships to the pandemic, specific project operational timelines relative to COVID-19 disruptions were not formally documented.
Additional constraints include potential interviewer bias and non-random participant selection affecting population representativeness, with the resource-intensive qualitative approach necessarily restricting the breadth of perspectives captured.
Future research should examine framework applicability across different national-heritage systems, institutional scales, and participatory-science domains. Significant differences between library, archive, and museum descriptive paradigms—including notably different CMS architectures and data-integration opportunities—warrant comparative investigation within heritage contexts. Mapping the specific relationships between crowdsourcing platforms and institutional CMS architectures, examining where platform-system overlaps facilitate or impede data-ingestion pathways, would further illuminate the infrastructural dimensions of the integration gap this research has identified. Longitudinal studies tracking institutional development and cross-cultural research would strengthen theoretical generalisability while incorporating participant perspectives to complement institutional findings.
Conclusion
This research establishes institutional integration capacity—encompassing technical infrastructure, inter-professional collaboration, staff development, and sustained resource investment—as a critical yet under-investigated factor constraining user-generated content integration in GLAM contexts.
The finding is, at one level, neither speculative nor distant from practice: Capacity limitations, not conceptual resistance, prove decisive. Only 40% of surveyed heritage professionals expressed confidence in workflow integration, despite 46% identifying it as preferable. Successful integration, interview analysis reveals, requires inter-professional collaboration between cultural-heritage and technical specialists, structured staff development addressing computational-skill gaps, and sustained institutional investment (Interviewees 1.3, 1.5, 1.6). Resource barriers compound these challenges: 47% operate on budgets ≤ £30,000, lacking infrastructure and expertise for systematic implementation.
Yet, resource constraints alone cannot explain why participatory initiatives so often fall short of their democratic promises. Analysis surfaces three institutional approaches—peripheral contribution, supervised collaboration, and workflow integration—which differ in how institutions exercise validation authority, not whether they retain it. What cross-sectoral comparison adds is that distributed-validation models operate across both heritage and scientific contexts; the differences lie in validation criteria—provenance, contextual interpretation, and cultural protocols—and the barriers remain systemic. In turn, this decoupling of operational advancement from epistemic redistribution exposes a fundamental tension: Technical sophistication does not automatically achieve democratic-knowledge production; rather, integration mechanisms determine whose knowledge receives institutional legitimation.
The analysis, then, yields three critical insights. First, successful integration demands inter-professional collaboration, staff development, and sustained investment beyond technical implementation. Second, the epistemic-authority spectrum demonstrates that operational integration does not necessarily entail epistemic redistribution. And third, integration barriers, including attribution infrastructure, reflect both capacity constraints and architectural choices that encode epistemic hierarchy: Systems lacking fields for user contributions exclude valuable knowledge on architectural grounds alone. Technical incorporation and epistemic redistribution constitute, therefore, distinct challenges—the former a matter of organisational capacity, the latter of institutional will to redistribute authority over whose knowledge counts as legitimate. Unless and until GLAMs rise to both, the passage from crowd to collection will continue on institutional terms alone.
Supplementary files
The Supplementary files for this article can be found as follows:
Supplemental File 1
Appendix A (Heritage-Professionals Survey Instrument). DOI: https://doi.org/10.5334/cstp.919.s1
Supplemental File 2
Appendix B (CMS-Specialists Survey Instrument). DOI: https://doi.org/10.5334/cstp.919.s2
Supplemental File 3
Appendix C (Citizen Science Coordinators Survey Instrument). DOI: https://doi.org/10.5334/cstp.919.s3
Ethics and Consent
This research was ethically reviewed and approved by the University Teaching and Research Ethics Committee (UTREC) and the University of St Andrews Research Ethics Committee in Summer 2020 (Code: AH14975).
Data Accessibility Statement
Anonymised interview transcripts and survey data are available upon request from the corresponding author. Interview recordings were anonymised and deleted; transcripts retain no identifiers. Streamlined survey instruments and interview protocols are provided in Supplemental files 1–4: Appendices A–D. Complete instruments available upon request.
