Skip to main content
Have a personal or library account? Click to login
GLAM Special Collection: A Generalized Framework for Participatory Transcription: Lessons, Challenges, and Interdisciplinary Impact from the Zooniverse Platform Cover

GLAM Special Collection: A Generalized Framework for Participatory Transcription: Lessons, Challenges, and Interdisciplinary Impact from the Zooniverse Platform

Open Access
|Jul 2026

Full Article

Introduction

In the early 1990s and 2000s, many galleries, libraries, archives, and museums (GLAMs) started to digitize their collections and place them online to better serve researchers, students, and the broader public. As Mika et al. (2017) of the large Biodiversity Heritage Library online consortium observe, these efforts “promised [to] dramatically enhance access but did not deliver” because without transcriptions, the data are not searchable and therefore unlikely to be discovered. To deliver on the unfulfilled promises of digitization, GLAMs turned to participatory transcription and tagging—often called crowdsourcing, citizen science, or citizen humanities—to create searchable transcriptions and metadata for documents to enhance search and discovery (Holley 2010; Oomen and Aroyo 2011; Terras 2016). Academics in the humanities, STEM, and social sciences were often key partners in these efforts, bringing research questions and requirements that shaped participatory projects beyond the GLAM aims of enhancing search and discovery (Ridge 2014). Creating transcriptions has long been an important but time-consuming activity required to create scholarly editions and encourage new research.

The first wave of GLAM digitization fostered the rise of the digital humanities as a sub or meta discipline (Terras 2016; Hedges and Dunn 2018), and expedited the creation of datasets for computational use in the humanities, STEM, and social sciences (Hill et al. 2012; Borgman 2015; Ridge et al. 2021). Case studies in the seminal volume Crowdsourcing our Cultural Heritage (Ridge 2014) described pioneering transcription projects like Transcribe Bentham on the MediaWiki platform (Causer and Terras 2014), and Old Weather on the Zooniverse platform (Blaser 2014), including information about how partnerships were formed, participant recruitment, transcription tasks, and, to a lesser extent, resulting data quality. Trevor Owens closed the volume with a piece arguing that “[w]hat crowdsourcing does (and most digital collection platforms fail to do) is to offer an opportunity for someone to do something more than consume information” by contributing to “public memory,” which “is the best way to engage our users in the fundamental reason that these digital collections exist” (pp. 278–9).

Over the past two decades, crowdsourced transcription projects featuring GLAM collections have attracted millions of volunteer participants, marking an unprecedented degree of collaboration and exchange between GLAMs, scholars, and the public (Causer and Terras 2014; Ferriter et al. 2019; Ridge et al. 2021). Simply by the measure of increasing engagement with collections—audience engagement being a key metric for most GLAMs—crowdsourcing has been a remarkable success. But engagement alone is an insufficient measure of impact when there are research questions driving data collection, and when enhancing collections search is a key aim. Crowdsourced transcription poses particular challenges for data quality due to the heterogeneity of contributors and their varying familiarity with source materials. Data quality assurance in crowdsourced transcription has been a topic of concern for decades (Brumfield 2012; Terras 2016; Van Hyning 2019; Prats López et al. 2020), and a concern in citizen science more broadly (Wiggins et al. 2011; Brabham 2013). Drawing on lessons from GLAM and academic crowdsourcing projects across many domains and platforms, Ridge et al. (2021) advise project creators to consider in advance how data quality will be measured, including facets like accuracy, completeness, or fidelity to the original sources.

While there have been hundreds of participatory transcription projects featuring different types of materials and transcription conventions (Terras 2016; Brumfield 2020) there are relatively few platforms to choose from if you want to build and launch your own project (Mika et al. 2017; Ferriter et al. 2019). Yet as Soacha-Godoy et al. (2025) observe, understanding of participatory science platforms and infrastructure “remains notably underdeveloped within the extant literature.” This is also true of participatory transcription across disciplines.

Zooniverse as a Case Study for Distilling a Participatory Transcription Framework

Herein, we use the Zooniverse platform (https://www.zooniverse.org) as a case study of participatory transcription methods. The purpose of this work is twofold. Our overarching purpose is to enhance understanding of the inherent complexities of text as a type of data, agnostic of discipline or transcription method, and illuminate how participatory transcription methods can either mitigate or exacerbate complexity. We argue that any project or platform design choice will interact in different ways with the baseline challenges of text as a data type and participatory transcription as a methodology, and without deeper insight into a given platform or project, the resulting data cannot be understood. Our second purpose is to shed light on the affordances and challenges of participatory transcription and data aggregation tools on the Zooniverse platform specifically, through references to multiple projects, publications, and hitherto unpublished information such as the full list of Zooniverse transcription projects, provided in Supplemental File 1.

Zooniverse is the largest and most used participatory research platform with transcription tools, but it did not start as a transcription platform. In 2007, the Galaxy Zoo project invited participants to contribute to real scientific research by identifying simple visual characteristics of images of galaxies captured by the Sloan Digital Sky Survey (Lintott et al. 2008). Tens of thousands of people answered the call, and the success of Galaxy Zoo led to the development of Zooniverse as a citizen science platform in 2009. Through external research partnerships over >15 years (including academic institutions and GLAMs), the platform has grown to include a variety of tools to facilitate different data extraction methods such as question tasks, marking, drawing, and transcription. Zooniverse is unusual in terms of the heterogeneity of disciplines and projects supported, the number of participants attracted, the geographic spread of project teams, and the depth of engagement inspired (Trouille et al. 2019).

Its size and scale make Zooniverse a valuable case study for exploring many of the innate challenges of text as a data type, and participatory transcription as an approach to creating textual datasets. Between 2010 and the time of writing, Zooniverse has hosted 174 transcription projects across disciplines in the humanities, STEM, and social sciences, many of which draw on GLAM holdings, are led by GLAMs, or have GLAM partners (see Supplemental File 1 for a full list of publicly launched transcription projects on Zooniverse). Sixteen bespoke transcription projects were launched between 2010 and 2015. In 2015, the Zooniverse launched the Do-It-Yourself Project Builder (PB), which allows anyone to build and run a crowdsourcing project for free. Transcription tools were incorporated into the PB in 2016, building on lessons learned from the early bespoke efforts. Early Zooniverse transcription projects such as Old Weather and Notes from Nature were cited frequently in foundational participatory research literature about the power and potential of crowdsourcing for GLAM and academic purposes (Oomen and Aroyo 2011; Ridge 2014; Hedges and Dunn 2018).

A Framework for Participatory Transcription

We have identified four common challenges in the creation, design, and execution of participatory transcription: variety, units of transcription, single-track versus multi-track methods (Brumfield 2012) with aggregation, and bias. For each challenge, we have distilled guiding questions based on existing participatory transcription literature, conversations with other practitioners, and our own professional experience. These guiding questions could be applied to other transcription projects and platforms not only at the point of project creation, but in data wrangling and analysis stages such as those outlined in Smith et al. (2023).

We are uniquely positioned to connect the broader theory and practice of participatory transcription with specific insights synthesized from Zooniverse projects to formulate the framework presented here. Blickhan has led transcription efforts on the Zooniverse platform from 2018 to the time of writing and has been Co-Director of the platform since 2021. Van Hyning led Zooniverse humanities and transcription efforts from 2014 to 2018, contributed to the design and running of the By the People project at the Library of Congress, and has used FromThePage to build GLAM transcription projects (Van Hyning and Jones 2021). The examples we share and lessons we distill about Zooniverse transcription projects are drawn from published scholarship, grey literature, internal documentation, public Zooniverse discussion boards and blogs, and institutional memory where no publications or documentation are available.

In the sections below, we use Zooniverse projects as a case study for distilling the framework and guiding questions in Table 1. Given the number of transcription projects hosted on Zooniverse, we can only touch on aspects of a few representative examples. We highlight projects that contributed significantly to Zooniverse transcription methodologies, and speak to the guiding questions, which have wider application to participatory transcription beyond the Zooniverse platform.

Table 1

A framework for participatory transcription with guiding questions.

CHALLENGEGUIDING QUESTION(S)
Variety
  • What are the properties (layout, form, genre, and any important features the transcription workflow needs to accommodate) of the texts to be transcribed?

  • What transcription conventions will capture the desired data? Is it important to capture text exactly as it appears in the original documents or standardize spellings, punctuation, etc. to align with disciplinary standards?

  • Is additional markup necessary, for example, to indicate supplied or deleted text or formatting like bold or italics?

  • What are the edge cases in which a particular approach that works for the bulk of the material might not work for a subset?

Units of transcription
  • What is an appropriate unit of transcription? A whole document, a page, a line, a word, a single character?

  • How does the unit of transcription translate to time spent on a task? Are there tradeoffs between unit size and task flow or immersion for participants?

  • What workflow design best supports participants in contributing the unit of transcription?

Single-track transcription or multi-track transcription with aggregation
  • Who will transcribe and how will their transcriptions be vetted, if at all?

    • Single-track: Can take many forms, for example, one or more participants transcribe, and other participants edit and/or an expert editor may finalize.

    • Multi-track: Can take many forms, for example, N-blind keyings combined algorithmically, with or without expert review.

  • What aggregation methods work best for combining multi-track transcriptions?

  • What resources (e.g., expertise, time) are needed to conduct expert review and/or aggregation? Do I/we have this expertise?

Bias
  • Is data quality higher when multiple independent transcriptions are compared and combined into a single reading, or when transcribers are allowed to collaborate?

  • Can collaborative methods be designed that retain the quality control metrics of independent transcription?

Variety and units of transcription

Variety is a challenge for any transcription platform because text is dense, heterogeneous, and multidimensional (Hedges and Dunn 2018). Documents can be written in multiple languages and scripts, text can be multidirectional, and layout can be structured (e.g., a form or table), unstructured, or semi-structured. As Ridge et al. (2021) observe, “transcription tasks may have very different user interface needs based on categories of texts.” Variety makes it difficult to design a single transcription approach that can be used for all types of text. At the very least, appropriate transcription conventions are needed to support different kinds of materials, while in some cases alternative transcription tools or platforms might be necessary.

We choose to discuss variety and units of transcription in the same section herein because they are very closely linked. Units of transcription refer to the way(s) that a text can be broken down into tasks that are suitable for the target audience of participants. Some transcription projects invite participants to transcribe a whole multi-page manuscript (Terras 2016; Van Hyning et al. 2023), while others structure tasks around a single page at a time, a section, line, word, or even character. The unit of transcription will vary based on the type of material being transcribed; the unit that works for transcribing a set of letters may not work for extracting weather data recorded in a ship’s logbook.

The variety and condition of original materials and differences in interpretation of handwriting are common sources of complexity and disagreement among transcription experts (Causer and Terras 2014; Hedges and Dunn 2018), but even clearly legible or printed text can be rendered in more than one way, depending on the goal(s) of transcription. For example, scholarly editors and those creating data for computational use must decide how to represent non-standard spelling and punctuation in original documents, and whether abbreviations such as N.Y. should have a space between N. and Y. or be expanded to New York: such seemingly minor differences can thwart traditional GLAM discovery systems, and limit data reuse (Matsunaga et al. 2016; Bowser et al. 2020).

Zooniverse has supported transcription of many document varieties and experimented with different units of transcription, adapting and updating methods over time based on findings about data quality and participant experience. For example, the platform has supported transcription of ancient papyri fragments as in the cases of Ancient Lives (launched 2011; Brusuelas 2016) and Scribes of the Cairo Geniza (launched 2017; Blickhan et al. 2021), modern military records as in Operation War Diary (launched 2014; Grayson 2016) and Measuring the ANZACS (launched 2016; Roberts 2022), and herbaria and specimen labels as in Notes from Nature (launched 2016; Thomer and Guralnick 2014; Matsunaga et al. 2016). This section highlights key Zooniverse projects that informed our understanding of the interactions between variety and units of transcription.

Though each project supported different text types and conventions for different GLAM and research purposes, a common goal was to allow for non-specialist participation (Van Hyning 2019). One way of doing this is through task design tailored to specific units of transcription and target audience. Old Weather (OW), the first Zooniverse transcription endeavor, engaged thousands of participants from 2010 to 2021, transcribing hundreds of thousands of 19th- and 20th-century ship logbooks of varying layouts (Brohan 2018). Iterations of OW presented different varieties of logbook, and asked participants to extract specific numerical data rather than transcribe whole pages (Blaser 2014). The OW team drew inspiration from crowdsourcing task design research (e.g. Haythornthwaite 2009) to design for participants who might contribute few classifications, but whose collective efforts add up (Eveleigh et al. 2014). They articulated design recommendations that would guide subsequent Zooniverse projects (not limited to transcription):

  1. Offer short tasks participants can complete quickly.

  2. Provide feedback about the value of contributions and the quality of the data.

  3. “Instead of trying to design projects that encourage all volunteers to become more committed [… design] projects that make dabbling easier and help dabblers to feel that their contribution is valuable and valued” (p. 2994).

These design choices were supported by concurrent research on other transcription projects and platforms, such as Transcribe Bentham (launched 2010), which invited participants to transcribe whole documents and apply specialized Text Encoding Initiative markup. Causer and Terras (2014) reported that combining transcription and markup increased task duration and complexity, posing significant barriers to participation: Subsequent builds reduced task complexity. Later efforts like the By the People project at the Library of Congress (launched 2018) opted not to include markup and provide feedback to participants through social media, email updates, and discussion fora, building on lessons from OW, Transcribe Bentham, and the Smithsonian Transcription Center (Ferriter et al. 2019).

Ancient Lives (AL) approached transcription very differently from OW. AL invited participants to transcribe ~500,000 thousand-year-old Graeco-Roman papyri fragments. The primary goals were to identify previously unknown works and authors and create scholarly editions. The fragments are written in multiple ancient languages and layouts vary. The transcription interface needed to enable non-specialists to contribute without pre-existing language skills or keyboards. The Zooniverse team created a transcription module in which participants first clicked an individual character within a fragment image, then selected a matching character from a pop-up keyboard (Figure 1).

Figure 1

Clickable onscreen keyboard within the Ancient Lives user interface.

This design relied on pattern recognition (Brusuelas 2016), and lowered barriers to engagement with difficult materials: More than 358,000 participants transcribed more than 1.16M characters over an 8-year period, and new textual discoveries were published (Brusuelas 2019).

In 2015, granular transcription or microtasking methods were applied to AnnoTate (AT) and Shakespeare’s World (SW). These methods were inspired by AL and other Zooniverse projects that have simple, click-based visual identification tasks. One such is Penguin Watch, which invites participants to count the number of penguins in a given image (Van Hyning 2019). The texts in AT and SW were heterogeneous, containing unusual letter forms, abbreviations, and multiple scripts and languages. AT participants transcribed by drawing a point at the start and end of a line of text, then entering text in the transcription box that appeared after the placement of the second point. SW used the same approach but modified the method so that participants could transcribe a line or as little as a single character. The goal was to build transcribers’ confidence by letting them make smaller contributions to a difficult text, though they could transcribe whole pages if they wished (Van Hyning and Wang 2024).

In each of these examples, the transcription interactions described succeeded in allowing non-expert participants to take part; the barrier to entry for volunteer transcribers is lower when they have the opportunity to contribute as much, or as little, as they like. However, the same methods that reduced barriers to engagement increased the complexity of the resulting data. Results are made yet more complex by the Zooniverse’s multi-track approach to transcription, wherein multiple transcriptions of the same data are collected as a way of incorporating quality control into the task (more on this in the following section); if participants are able to contribute partial transcriptions, but are unaware of what parts of the page other people have transcribed, there is no guarantee that an entire page of text will be completed. Sometimes giving people the option of choosing not to work on certain elements of a page of text can result in areas being under-transcribed.

Breaking a text into appropriate units of transcription can also be complicated by the need for context to accurately interpret text. Scribe, built in partnership with the New York Public Library (NYPL), was modeled on OW and intended for use by other organizations wishing to modify the codebase and run it on their own servers (Smith 2013). In Scribe projects, participants could select a region of a page of text (a “subject,” in Zooniverse parlance) using a bounding box tool. The regions within the bounding box became new, “secondary subjects” via post processing, to be transcribed in a later stage. A single image might produce multiple secondary subjects, but each with fewer words needing transcription than the original subject. Transcription of text in the secondary subjects resulted in “tertiary subjects.” Secondary and tertiary subjects could be labeled as accurate or inaccurate in yet another workflow, ideally resulting in “final subjects that […] represent singular assertions about the data contained in a document validated by between three and 25 people” (NYPL 2015).

While the Scribe approach helps to “reduce complex document transcription to a series of smaller decisions that can be tackled individually” (NYPL 2015), it can create scenarios where text is more difficult to interpret, as it is divorced from its original context. Rawson and Muñoz (2019) discuss this in the context of the NYPL’s What’s on the Menu? project through a framework of “scalable” versus “nonscalable” elements of digitization work; the infrastructure of that project allowed transcribers to zoom out from a smaller segment to view the area in its original context, noting that the text being transcribed would sometimes need “to be understood through the relation between the line of text in the bounding box and other nearby text, like a heading.” Retaining context in the transcription workflow can improve both data quality and participant experience.

Later Zooniverse projects addressed the need for contextualization through things like sequential page delivery, such as in The Davy Notebooks Project (launched 2019), which transcribed more than 75 lab notebooks hand-written by Sir Humphry Davy. While the unit of transcription was an individual line of text on a page, the notebook pages themselves were more likely to be successfully transcribed when viewed in chronological order. Participant feedback after a pilot version of the project—in which pages were served at random—confirmed this preference, and so Zooniverse developed the ability to transcribe sequentially (Blickhan et al. 2024). This feature was a divestment from the typical Zooniverse method of serving data at random, a choice originally intended to avoid bias in STEM classification efforts (Wiggins et al. 2011; Trouille et al. 2019) and enable the Zooniverse platform to scale well during periods of high participation (Smith 2013), but which proved less useful for transcribing certain varieties of text.

The lesson learned across Zooniverse projects, and confirmed by other projects and platforms, is that no single transcription method will fit all units or text varieties equally well. Considering variety and units of transcription will help teams in their design decisions no matter what participatory transcription platform they are using, but these are only two considerations when choosing the best approach or platform. There are further data quality considerations connected with how many people transcribe, and how to minimize bias when multiple transcribers are involved.

Single-track or multi-track methods with aggregation

In a blog post about crowdsourced transcription approaches and issues of quality control, Brumfield (2012) delineated potential participatory transcription methods in two main categories: “single-track,” in which one or more people create a single transcription that is edited by amateurs and/or experts; and “multi-track methods,” including various N-blind-keyings. In order for multi-track transcription methods to produce a single outcome, Brumfield describes the need for the keyings to be “compared programmatically.” Most participatory transcription platforms used by GLAMs deploy single-track methods in which one or more participants transcribe a text, and other project participants or experts edit it, as in Transcribe Bentham on MediaWiki (Terras 2016) and By the People on Concordia at the Library of Congress (Ferriter et al. 2019). Zooniverse uses multi-track methods, a notable exception among participatory transcription platforms.

Early Zooniverse platform designs prioritized participant independence and multi-track methods (Smith 2013), both foundational to early theories of distributed decision-making by heterogeneous groups of experts and non-experts (Surowiecki 2004; Brabham 2013). Participant independence and multi-keying are also used on platforms such as Amazon’s MTurk (Ipeirotis et al. 2010) and FamilySearch Indexing (Brumfield 2012), and to create transcriptions for qualitative interviews in the social sciences, legal transcripts, and GLAM and digital humanities projects such as the Early English Books Online-Text Creation Partnership (Van Hyning 2019).

In this section, we describe how Zooniverse transcription projects sought to identify appropriate N-values for multi-track transcription, how incorrect N-values can result in over- or under-transcription and low data quality, and how different experimental designs were deployed to find alternative methods to produce higher-quality data. For a project like the original Galaxy Zoo, where participants answered simple multiple-choice questions, aggregation was a relatively straightforward process of determining the majority response among N submissions, where N is the “retirement limit” or the point when an image has achieved sufficient input for the research needs and will not be shown to further participants (Lintott et al. 2008). Zooniverse project creators determine the N-value based on how many non-specialist classifications are needed on average to approximate expert classifiers (Swanson et al. 2016; Smith et al. 2023). In the case of textual data in any multi-track transcription platform, a high N can swiftly multiply into an unmanageable data processing problem, while a low N can lead to incomplete transcription (Singh et al. 2021). Early Zooniverse projects like OW, AL, and Operation War Diary (OWD) predetermined relatively high Ns—between 10 and 15—and then lowered the threshold after initial data analysis (Williams 2014; Grayson 2016; Brohan 2018).

For each project, a participant’s submission included positional data and a transcription. Positional data needs to be aggregated before the transcriptions to ensure that the same unit of text within an image is being compared. Different units of transcription (a page, a line, a character) require different aggregation approaches. In AL, N (retirement) was set at the fragment level, while the unit of transcription was an individual character. Positional data for each character had to be clustered before the transcriptions could be compared and combined into a single reading. A text string alignment method was adapted from an amino acid sequencing tool to derive the full text of each fragment (Williams 2014). As Williams et al. (2014) describe, the aggregated transcriptions required considerable intervention to be suitable for the project team’s goal of publishing scholarly editions.

As AL was harnessing genetic sequencing for papyri, Notes from Nature (NfN) researchers also drew on the fields of linguistics and genetic sequencing to understand the complexities of their resulting data and develop aggregation approaches (Matsunaga et al. 2016). Thomer and Guralnick (2014) explored how issues of variety and transcription conventions could interact with aggregation. They drew parallels with the scholarly editing of medieval manuscripts, which distinguishes between copy texts, which favor the conventions of an individual manuscript, and critical editions, which might amalgamate numerous textual witnesses into a single reading. In multi-track transcription methods, each participant is like a medieval scribe—capable of reproducing the text exactly as they find it or making interventions to improve clarity, supply missing words, or conversely, to introduce new errors. Whether an exact transcription (N.Y, NY, or New York) or an amalgamated one is appropriate depends on the intended use of the data. Project creators must consider in advance how data will be used and thus how quality will be measured.

For early Zooniverse projects, the question of identifying an appropriate N value was not dynamic: Aggregation was performed after data collection (Hines 2017). This resulted in delays between identifying and rectifying problems. For example, in OWD, participants transcribed and annotated WWI field diaries from the Western Front by clicking parts of pages which triggered a pop-up transcription box and a list of categories, such as date, troop activity, and weather (Barber 2018). In keeping with theories of crowdsourcing put forward by Surowiecki (2004) and Brabham (2013), the team assumed not every participant would transcribe everything, but enough would transcribe all parts to aggregate multiple responses per page. An early look at the data revealed that many participants focused on the least subjective data, like times, dates, places, and names, but tags that required interpretation were less often applied. N = 15 was too high for some parts of the page, and too low for others (Figure 2).

Figure 2

Operation War Diary (OWD) aggregation analysis of transcriptions on a War Diary page.

Drawing on lessons from AL, NfN, OW, and OWD, new real-time clustering and genetic-sequencing algorithm aggregation methods were developed for AT and SW (Hines 2017). Participants clicked on an image to place dots at the start and end of a section of text and then typed into a pop-up transcription field. Once three people transcribed a section of text identically it was deemed complete, and the surrounding dots turned grey, signaling to participants that no additional transcriptions were needed (Van Hyning 2016). Pages were considered complete when three participants confirmed that all lines on the page were surrounded by grey dots, essentially layering the count-based N-retirement technique over the new, agreement-based aggregation method. Early data analysis on both projects was promising, but heterogeneity of dot placement over the lifetime of each project led to overclassification, slow retirement, and lower data quality (Van Hyning and Wang 2024). This echoed findings from AL, where clustering algorithms’ tolerance for variation were found to be narrow, especially at the character level where data points are close together (Williams et al. 2014). Accurate positional data are key to successful aggregation, and imprecise positional data are frequently the root of failure for text string comparison.

Mutual Muses (launched 2017) tried a different approach by gathering N = 6 transcriptions of whole documents and measuring agreement among them as a percentage of text overall (Deines 2018) rather than through clustering and string alignment. Manual review was reserved for letters flagged as being particularly discordant. The team published the resulting data and processing scripts on GitHub (Lincoln and Gill 2018).

The complexities of variety and units of transcription mean that no single aggregation method can neatly compare different transcriptions and accurately render a majority rules transcription in every case. Zooniverse projects have experimented with a variety of transcription units, and multi-track methods and N-values to facilitate participation and quality data capture, but no method fits all units equally well. The challenges of text aggregation have thus been a major driver of new approaches to transcription on the Zooniverse platform. For many project teams, the question of single- or multi-track transcription should inform their choice of tools or platforms to use, due to the complexity of text aggregation methods.

Bias

Bias has received much attention in wider discussions of citizen science, and the extant literature is typically tied to established disciplinary practice (e.g., Griffiths-Lee, Nicholls, and Goulson 2023) or based on demographic information for distributed digital volunteer efforts (e.g. Blake, Rhanor and Pajic 2020). For the framework presented here, bias refers to the risk of influence from other transcribers leading to lower-quality data—an assumption that underpinned the creation of Zooniverse’s original infrastructure that prioritized independent, multi-keying methods (Smith 2013). This approach was celebrated for its potential to enable mass-participation by non-experts without sacrificing data quality due to the introduction of bias (Oomen and Aroyo 2011; Wiggins et al. 2011). However, multi-keying and task independence—especially when combined with the previously-discussed complexities of text as a data format—can make aggregation particularly difficult, leading to struggles with producing useful transcriptions even when high-quality individual transcriptions were provided by participants.

Between 2010 and 2017, the Zooniverse team tried many approaches to independent multi-keying in an effort to increase accuracy and efficiency, with varying results (Van Hyning and Wang 2024). In 2017, Zooniverse was awarded an Institute of Museum and Library Services National Leadership Grant to create tools and infrastructure to assist GLAMs in building crowdsourced transcription projects on the Zooniverse platform. The research goals included investigating whether the Zooniverse independent transcription method produced better results than “allowing volunteers to see previous transcriptions by others” (Van Hyning et al. 2017). This work produced Anti-Slavery Manuscripts (ASM; launched 2018), focused on a large collection of handwritten American Civil War-era letters (Blickhan 2020). ASM featured a custom frontend and bespoke transcription mechanism that supported an A/B test, comparing independent and collaborative transcription methods.

As described in Blickhan et al. (2019), rather than comparing single- and multi-track methods, the ASM experiment compared blind-keying multi-track methods with open multi-track methods in which participants could see and interact with data produced by previous contributors. In ASM, only the first person to transcribe a page would need to input positional data. Subsequent participants could essentially “re-use” existing positional data and add their transcription for the relevant area of text. As with SW, once three matching transcriptions were submitted for a given set of positional data, that unit was considered complete, and no new transcriptions were collected. This method improved on previous transcription advances in projects like SW by reducing the need for clustering positional data, which previous projects had demonstrated was a key barrier to successful data processing. The open multi-track methods allowed Zooniverse to retain some quality control elements of blind multi-track transcription, while reducing identified issues with data quality and use.

Blickhan et al. (2019) demonstrated that task independence in transcription can needlessly prolong projects, and lead to oversampling and lower data quality. Open multi-track transcription methods were subsequently adopted as standard into the Zooniverse platform and have since been used by major transcription efforts including the Davy Notebooks Project, which successfully transcribed more than 10,500 pages of handwritten scientific notebooks (Blickhan et al. 2024). The application of multi-track methods to transcription, alongside experiments with different units of transcription and bias mitigation, complicates the assertion in Prats López et al. (2020) that complex tasks like transcription are not “suitable for division into simpler tasks.” While complex task types are certainly more difficult to break into smaller tasks without impacting other factors like data quality and post-transcription analysis, the work described in this section demonstrates that design iteration over time can produce impactful tools and processes for breaking complex texts into manageable units without sacrificing quality or leading to a degraded user experience.

The results of the ASM experiment led to additional questions around whether collaborative efforts could be expanded to (at the time) novel ideas around human-computer interaction. With support from the American Council of Learned Societies, the Zooniverse team developed infrastructure for ingesting Handwritten Text Recognition (HTR) outputs, which could then be displayed for Zooniverse participants to review and edit. Though research to determine how best to optimize this human-in-the-loop approach to text transcription remains ongoing at the time of writing, the proof of concept succeeded. The Zooniverse infrastructure created to ingest and display machine-generated data for review and editing by participants is impacting projects beyond GLAM and humanities efforts; STEM research teams are using this approach in projects that include biomedical and astrophysical efforts, now colloquially referred to as Correct-A-Machine style projects (e.g., Sankar et al. 2023).

The challenge of bias is closely related to that of single- or multi-track transcription; the risk analysis for bias ultimately depends on the type of data needed and goals of the project. As with each challenge described, no single element of this framework can drive every decision when creating a participatory research project. Indeed, thinking across categories is often what leads to innovative solutions for challenging problems. For example, after the development of open multi-track methods for ASM, the Zooniverse team recognized that even with these new changes to the tools, many GLAM teams running transcription projects on Zooniverse would still need assistance in aggregating text. Even though the guiding question around bias was solved with open multi-track transcription, aggregation remained a challenge for the transcription data produced. To address this issue, Zooniverse, with support from the National Endowment for the Humanities, created the Aggregate Line Inspector and Collaborative Editor (ALICE; https://alice.zooniverse.org). ALICE is a browser-based application that research teams can use to review and edit Zooniverse transcriptions produced through the open multi-track transcription tools produced through ASM (Blickhan 2021). Since 2019, these new tools for transcription and aggregation have led to hundreds of thousands of pages of successfully transcribed text.

Discussion and Conclusion

By providing a framework and guiding questions for participatory transcription, and using Zooniverse as a case study, we have filled two gaps in the literature: the first about participatory transcription methodologies and the complexities of text as a type of data broadly; the second about how Zooniverse participatory transcription methods were iteratively developed in response to these complexities, illustrated through examples of challenges and successes from specific projects. This piece joins a small number of articles devoted to participatory transcription in this journal and the wider participatory research field, and is rare in its discussion of iterative tool development across distinct projects, disciplines, GLAMs, and time—a contribution that Zooniverse is uniquely placed to offer, given its status as a platform serving multiple disciplines, encompassing multiple modes of data collection. To our knowledge, no other volunteer-powered crowdsourcing platform has explored as many permutations of transcription methods as Zooniverse. It took many years of experimentation and iteration to understand the interactions between the four factors we identify and how to mitigate the impact if any of these factors are out of kilter.

Text is a ubiquitous data source across disciplines, and the need for textual data—driven in part by recent advances in AI—has never been clearer nor more fraught. Crowdsourcing, AI, and HTR models all have a role to play in expanding the availability of textual data (Nockels et al. 2022), and the framework and guiding questions articulated here can support those choosing amongst platforms and methods or indeed combining them. Recent projects like The Material Culture of Wills: England 1540–1790 (launched 2024) and Field Journal Fix-Up (launched 2025) use newly developed Zooniverse tools to ingest and display automated transcriptions of text from external sources for volunteer review and correction (Guralnick et al. 2024). The Material Culture of Wills displays HTR output from Transkribus, and Field Journal Fix-Up displays machine transcriptions produced by Amazon Textract.

Even as HTR and AI models become increasingly accurate (see, e.g., Cohen 2025), participatory methods play an important role in connecting people with primary sources, and with one another. GLAMs will continue to benefit from running crowdsourcing projects to engage people whom they might never welcome through the physical doors of their institutions, while academic partners can benefit from the generosity of volunteer participants, and the serendipity and fresh insights that a distributed contributor base brings (Oomen and Aroyo 2011).

Offering people opportunities to engage with primary sources and one another can foster literacy, community, and civic engagement, which are all vital to resisting the attacks on democracy, GLAMs, K–12, and higher education that we are witnessing around the world (Stanley 2026). More than a decade after Owens (2014) called on GLAMs to engage people in public memory work through crowdsourcing, inviting broad engagement through participatory projects is more important than ever, even in the face of the challenges described. As tools and infrastructure for participatory engagement have evolved over the past decade and a half, they can continue to be adapted into the future, rising to meet the needs of practitioners and publics alike, and providing opportunities for meaningful engagement with shared history and heritage materials, contributing to scientific research and discovery, and connecting people with one another.

Supplementary File

Supplemental File 1

List of Zooniverse Transcription Projects (2010–2026). DOI: https://doi.org/10.5334/cstp.934.s1

Acknowledgements

Many thanks are due to our Zooniverse colleagues past and present, fellow crowdsourcing practitioners and researchers, project teams, and the many volunteer participants who have contributed their time and energy to the often messy, often joyous, ongoing work of public memory.

Author Contributions

Both authors contributed equally to the conception and design of the work, as well as the writing and revision, and have provided final approval.

DOI: https://doi.org/10.5334/cstp.934 | Journal eISSN: 2057-4991
Language: English
Page range: 19 - 19
Submitted on: Oct 15, 2025
Accepted on: Jun 12, 2026
Published on: Jul 27, 2026
Published by: Ubiquity Press
In partnership with: Paradigm Publishing Services

© 2026 Samantha Blickhan, Victoria Van Hyning, published by Ubiquity Press
This work is licensed under the Creative Commons Attribution 4.0 License.