Skip to main content
Have a personal or library account? Click to login
Reusing Language Corpora for Token-Based Typology Cover

Reusing Language Corpora for Token-Based Typology

By:  and    
Open Access
|Jul 2026

Full Article

(1) Context and motivation

Linguistic typology is inherently empirical. As a discipline built upon the collation of data points from multiple languages, typological work lends itself to a culture of data reuse – information about properties of single languages is continuously combined with new insights from new descriptions to further improve what we know about linguistic diversity. While reuse studies are common for well-known typological databases such as the World Atlas of Language Structures (Dryer & Haspelmath, 2013), typological studies which reuse publicly available corpora are much rarer (Gawne & Berez-Kroeker, 2018, 27).

In this contribution to the special collection, we reflect on our praxis as usage-based typologists who frequently reuse data. In our work, we have encountered a range of different datasets, each with their strengths and weaknesses for reuse purposes, and each with their own requirements for data processing and analysis. Our purpose in writing this paper is not to wholesale criticise these datasets, but rather to make it easier for others to join us in their reuse. More generally, it is our hope that, by reflecting on our experiences, linguists will become more aware of issues around data reusability in all stages of the data management workflow and join us in the conversation about how to create a more data reuse-friendly ecosystem in linguistics. This conversation is also crucial for the flow-on effects of our decisions on the reusability of data outside of linguistics (cf. Desai et al., 2024).

(2) Description of the datasets

Our previous work has relied on a number of single-language and cross-linguistic corpora. The descriptions of the most recent versions of all datasets we make reference to can be found in the Supplementary Materials; we include the references for datasets we used in this paper and discuss their properties where relevant.

(3) Case study 1: Single-language documentation corpora

Our first case study comes from my (Peck) experience in preparing my dissertation (Peck, 2025). Like many other doctoral students, I was forced to change the topic of my dissertation due to COVID-19. While I had originally planned on writing a descriptive grammar of Kera’a (glottocode: idum1241), I made the practical decision to change topics to a comparative corpus-based study of serial verb constructions, for which I would use the data I collected in the field along with data from pre-existing language documentation corpora from the Himalaya.

My baseline selection criteria for corpora were as follows:

  1. The language should be spoken in the Himalaya.

  2. The language in the corpus should have a good linguistic description that I can access.

  3. The corpus should be based on naturalistic spoken data and be easily accessible.

  4. At least 30 minutes of audio should be accompanied by transcriptions, translations and glossing in English to match the size of the Kera’a corpus.

To find language resources, I first searched the aggregated metadata database run by the Open Languages Archives Community (OLAC).1 I supplemented this information with targeted searches of area-relevant archives and collections, such as the Pacific and Regional Archive for Digital Sources in Endangered Cultures (PARADISEC),2 Pangloss,3 and the Endangered Languages Archive (ELAR).4 Finally, I identified languages with a grammar or grammar sketch in the Himalaya via Glottoscope5 and tried to locate any corpora used for the descriptions. This resulted in the identification of 130 languages with datasets which I could potentially reuse for my project.

It quickly became clear to me that not all datasets fit my requirements. Datasets that I had identified via OLAC included written material scraped from Wikipedia or sourced from translations of religious texts to the tune of 42 languages. First perusal of archived deposits surfaced a decent number of corpora suitable for my research questions, but a surprising number of deposits were not suitable, as they were missing transcriptions, translations and/or glossing (cf. also Seifart, 2021, p. 127). At this point, I pivoted to make my selection criteria more stringent.

As I wished to conduct an in-depth comparative study, I decided to limit the genealogical and geographical variation for my choice of languages. As such, I limited myself to the East Himalaya and to one language per branch of Trans-Himalayan (cf. van Driem, 2014). I identified seven corpora for reuse, based upon the number of texts which could be worked with in ELAN (Wittenburg et al., 2006), my methodological tool of choice, as well as the amount of time and effort which it would take me to process the data in a manner similar to my Kera’a corpus. I ended up choosing a total of three language datasets for reuse in my dissertation and excluded one later during data processing for practical reasons.

The final two language corpora that I reused in Peck (2025) were from Duhumbi (Bodt, 2018a, 2018b, 2018c, 2018d, 2018e, 2018f) and Galo (Post & Modi, 2001).6 This is the first academic data reuse of these corpora by a third party, to my knowledge. Both corpora required me to go through a number of processing stages to create annotations comparable to what I had for my field data. The Duhumbi material archived by Bodt on Zenodo is one of the most reusable documentation corpora that I have encountered, due to his commitment to Open Science. While Bodt used Transcriber for transcription, the files created by this programme can easily be imported into ELAN. While the transcription files themselves were not glossed, each file was accompanied by a .pdf output from Toolbox with interlinear glossing. The inclusion of a Toolbox dictionary in .txt format allowed me to reproduce Bodt’s glossing semi-automatically via Interlinearization Mode in ELAN. In summary, it was very easy for me to reproduce the existing analysis using a different annotation schema built on the shared building blocks of transcriptions, glossing, and translation.

My experience working with the Galo materials was slightly different. When I first identified the dataset as reusable, I noticed that many of the files were not open access, although the data access details suggested they should have been.7 I reached out to Mark Post directly to let him know as well as ask him for access to the files, which he kindly provided me. From the beginning, the format of the Galo corpus was in .eaf, allowing me to jump straight into extra annotation of the data in ELAN. Furthermore, the quality of annotations of the Galo data was incredibly high, allowing me to approach further analysis of the data with ease. However, the structure of the ELAN files was different, as Post and I used different annotation schemas (cf. von Prince & Nordhoff, 2020). While I considered reproducing Post’s analysis in a new file using my own annotation schema, as I did for Duhumbi, the quickest way of doing this would have required me to manually create a lexicon from the pre-existing annotations on top of manual copy-and-pasting and editing of transcriptions. Given the high quality of pre-existing analysis, I decided to minimise my overall workload and avoid introducing human error on my part by instead simply adapting the schema that Post used by adding extra, independent levels of annotation to capture information that I otherwise would have had if I had reproduced the analysis using my original schema. The differences between the two schemas were negligible when carrying out qualitative work; for quantitative work, the differences were reconciled when importing and processing the data in R (R Core Team, 2024).8

The outcome of my data reuse was the creation of two smaller corpora, consisting of annotation files including enriched information about information structure and information flow for each language. In line with Open Data practices, I wished to make these downstream products available openly in a way that acknowledged the original work of Bodt and Post. However, there was no clear path as to how to do this. While I originally planned to credit both creators with authorship on my derivative datasets, automatic assignment of authorship without consultation is unethical. My sole authorship, however, would obscure their key role in making the reuse possible and the fact that my analysis builds upon their intellectual work. An inclusion of datasheets for datasets (Gebru et al., 2021) makes the provenance of the data clear but still does not give scholarly credit. After consultation with both Bodt and Post, the Duhumbi dataset is now co-authored by both me and Bodt, while I remain a solo author on the Galo dataset.

Choosing how to distribute the data was similarly difficult. When I submitted my thesis, I made the datasets available on OSF with a CC-BY license. In the process of preparing this article, I discovered that the PARADISEC Access Conditions9 precludes further sharing of derivative products, even if materials are Open Access, and requires derivative products to be deposited with PARADISEC. As such, I temporarily restricted access to the Galo dataset and reached out to the archive to fulfil my obligations. After discussing the issue with Post and PARADISEC, the conclusion was that my dataset should be separately distributed with a CC BY-NC-SA license and referenced in the metadata of the Galo archive on PARADISEC, rather than archived alongside the original collection. The Duhumbi materials clearly state the terms of use on each page, such as the CC BY-NC-SA license and how the material should be cited; the license was updated in my repository accordingly. No information is included about how reuse datasets should be handled beyond these license requirements. A DOI is now associated with all four datasets comprising my collection: the larger collection, my Kera’a subcorpus, the Duhumbi dataset co-authored with Bodt, and the Galo dataset.

My experiences throughout the process of reusing data for Peck (2025) highlight just a few of the many issues confronting the reuse of language documentation corpora. Finding suitable resources was a large problem at the beginning of my project (cf. Yi et al., 2022). While OLAC is a valuable resource, it does not capture all language resources, especially when depositors choose to archive on a platform outside traditional language archives.10 There is also limited information available more generally about the quality of corpora and their suitability for reuse (cf. Babinski et al., 2022). Different workflows in creating corpora must be reconciled for comparative work (Aznar & Seifart, 2020; von Prince & Nordhoff, 2020). Finally, the publication of downstream datasets has its own problems, in that the pathway to publication is not clear and must abide by restrictions on the use of the original data, which is not prominently displayed in many instances.

(4) Case study 2: Cross-linguistic spontaneous speech corpora

Our second case study stems from our ongoing typological work (Peck & Becker, 2024; Becker et al., 2025; Becker, n.d.) reusing data from the Multi-CAST (Haig & Schnell, 2021) and DoReCo (Seifart et al., 2024) corpora. In these studies, we examine different interactions between the spoken signal and grammar across languages, namely the association between pausing and (certain types of) clausal boundaries, as well as frequency effects in terms of the phonetic duration of verbal person marking. We outline a number of issues which we have encountered in our reuse journey, concerning data processing and curation and the generalisability of data.

The quality of data curation is crucial to the reusability of a dataset (Hardwicke et al., 2018). Multi-CAST and DoReCo are both centrally curated, although in different ways. One of the main purposes of the Multi-CAST corpora is to analyse referential choices in spontaneous speech across languages (e.g., Haig & Schnell, 2016; Schnell et al., 2021; Vollmer, 2023). The corpora therefore share common grammatical and referential annotations, the latter of which was developed specifically for Multi-CAST (Haig & Schnell, 2014). The annotations in this collection are, however, only time-aligned with the speech signal on the clause level, and do not include information on pausing. To analyse the distribution of silent pauses in different clausal contexts, we thus added pause annotations for Peck & Becker (2024).11

DoReCo, on the other hand, has time-aligned phonetic annotations, which reflects its main purpose to study phonetic properties of spontaneous speech that are relevant to the morphosyntactic structure of language (e.g., Blum et al., 2024; Paschen, 2025; Stave et al., 2021). Grammatical annotations in DoReCo are taken from the original single corpora and not standardized across the collection. This choice being more than justified, it can make quantitative corpus typology more difficult in certain cases when target constructions and contexts need to be defined and extracted for each language separately (if the use of glossing annotations is required) (cf. Levshina, 2022, p. 149). One such example is demonstratives. This comparative concept (Haspelmath, 2010) is reflected in language-specific analyses by specific glosses of “PROX” or “DIST” in Beja (Vanhove, 2024) or “RECOG” in Komnzo (Döhler, 2024), with the broader gloss “DEM” in Jejuan (Kim, 2024), or annotations on the part of speech level like “det” in Arapaho (Cowell, 2024) or “DEM” in Sümi (Teo, 2024). To work with DoReCo data for Becker et al. (2025) and Becker (n.d.), we unified the relevant grammatical annotations across languages for cross-linguistic comparison within R (R Core Team, 2024). Although the responsibility of defining comparative concepts for cross-linguistic studies lies with the typologist and not with the language experts and corpus creators, it is our impression that many typologists underestimate the time and work with the data needed to do so, assuming that curated corpus collections do not need many additional annotations and are ready to be used out of the box.

Another issue we ran into was the consistency of annotations within the different corpora within a collection, specifically when reusing DoReCo data for morphosyntactic questions. This is not to criticise the choices made by the curators of the DoReCo collection to carry over the original corpus-specific morpho-syntactic annotations. We believe this is simply an unintended consequence of prioritising adding more value to pre-existing corpora by creating additional phonetic data annotation layers over ensuring consistency in morphosyntactic annotations. To give one example, the Movima DoReCo corpus (Haude, 2024) is the result of merging separate, previously compiled corpora with different annotation conventions. As a consequence, morphosyntactic categories can sometimes have more than one label, which means that more careful (manual) data inspection is required before relevant contexts can be automatically extracted. In Movima, third person forms can refer to present or absent persons. The absent feature is glossed as either “a”, “AB” and “ABSN” in the corpus. This means that the reuser must first identify this many-to-one correspondence as a property of the dataset, before being able to further work with the data. Further inconsistencies for other DoReCo datasets are detailed in Paschen (2025, pp. 9–10). To what extent these “quirks” are known to the larger reuse community is an open question.

To the full credit of the curators of Multi-CAST and DoReCo, it is very easy to find information about the licenses of all datasets on their websites. The majority of materials in Multi-CAST are licensed under CC-BY, with the audio for English being in the public domain. The situation for DoReCo is more complex, with single language datasets regularly having different Creative Commons licenses for audio and annotations. This information is prominently displayed for each dataset in the list of all datasets in the collection as well as on the page of each individual dataset. A number of language datasets have licenses which restrict the distribution of derivative products.

We have sometimes run into the issue that the reuse of convenience samples, like these collections, receives pushback from a wider audience. This is not unique to reuse projects relying on language documentation corpora, such as Multi-CAST and DoReCo, but applies more widely to how we assess the robustness of findings for a given language. This assessment tends to favour well-resourced (Western) languages and disproportionately disprefers small datasets from legacy materials, such as Kamas (Gusev et al., 2024), a dataset of an extinct language from a single speaker. However, we believe the increased insight into linguistic diversity gained from the inclusion of a “suboptimal” dataset is worth the uncertainty of how well a language is captured by the data, which can be accounted for elsewhere in data analysis and/or the conclusion.

A last comment we would like to make concerns the platforms used for Multi-CAST and DoReCo.12 Multi-CAST has recently moved its collection interface from a university-hosted website13 to long-term repositories. While the shift reflects a move towards more sustainable data practices, the user experience in accessing the data is now much diminished, as useful sections such as descriptions of datasets and papers using the collection are now no longer prominent or available. DoReCo is also hosted via a university website with datasets available locally, rather than outsourced to a repository.14 However, at least in one case, we discovered that simple human error in link creation interfered with access to a dataset; this has since been remedied. Furthermore, products based on reuse of these data are not discoverable via the platforms themselves. Rather, the reuse of Multi-CAST and DoReCo is primarily trackable via citation practices. It is difficult to know the exact scale of reuse, though, as paper citations do not entail that the dataset associated with the paper was reused, nor does it capture data reuse outside of the scholarly ecosystem (Piwowar & Vision, 2013).

(5) Case study 3: Cross-linguistic corpora and databases

This case study reports on my (Becker) experiences concerning reusability when using the Universal Dependency (UD) treebanks (Zeman et al., 2021) and UniMorph (McCarthy et al., 2020) for quantitative typology in Guzmán Naranjo & Becker (2018; 2021) and Becker (2024). The studies deal with syntactic and morphological questions, namely cross-linguistic tendencies regarding word order and zero marking in inflectional morphology. UD is a large and ever-growing collection of dependency treebanks, i.e., syntactically parsed corpora. It provides a language-independent framework for syntactic analysis and annotation and, despite being mainly used for natural language processing (NLP; e.g., syntactic and semantic parsing), it has also become important for quantitative corpus typology in recent years (de Marneffe et al., 2021, p. 256). UniMorph is a cross-linguistic database of inflectional paradigms for individual lemmas including nouns, verbs and adjectives. It is compiled and curated by computational linguists with the objective to create more cross-linguistic resources especially for NLP-related tasks in computational linguistics (cf. McCarthy et al., 2020, p. 3922). Despite not being primarily designed for typological research, both resources are immensely valuable for quantitative typology (e.g., Berdicevskis et al., 2018; Gerdes et al., 2021; Levshina, 2016, 2019). That said, UD and UniMorph have certain properties that researchers should be aware of when using the resources for their projects. In the remainder of this section, I will focus on four issues that I have experienced when working with these resources, which are likely representative of the issues that typologists could encounter when working with UD and UniMorph (and similar resources).

Both resources include data from a large number of different languages: UD currently has data for 150+ languages and UniMorph contains inflectional paradigms from 169 languages. While this is a substantial number of languages, these datasets are not fully representative of the diversity found in the world’s languages, given their bias towards European (and Asian) written, standard languages, which is acknowledged by the curators themselves (Nivre et al., 2020, p. 4040) and likely a consequence of our data bias overall (Levshina, 2022, p. 148). However, it should serve as a reminder to typologists who wish to work with these data that the results of working with these datasets may be biased towards certain language families and areas (Schnell & Schiborr, 2022, p. 182). Such bias can be mitigated to an extent in statistical modeling (cf. Guzmán Naranjo & Becker, 2022), but it cannot make up for a lack of data, e.g., from Australia or North and South America.15

Related to this issue is the fact that language varieties are not always clearly identifiable (cf. Tables 15 and 18 in the Supplementary Materials). Although UniMorph provides language names together with ISO 639-3 codes, both are sometimes too broad for a language variety to be classified properly. For instance, the ISO 639-3 code used for Quechua is que, although this denotes a macrolanguage (language family) and more specific language-level codes are available. UD faces similar issues, as it provides language names and families as sourced from WALS (Dryer & Haspelmath, 2013). For typological studies, however, it is important to know exactly which varieties (in which locations) we are working with. The usability of both UD and UniMorph for linguistic typology could be substantially improved by adding glottocodes for clearer identifiability of the languages included. Glottocodes correspond to a curated, precise system of identifiers for language varieties with constant updates as new varieties are documented (Forkel & Hammarström, 2022).

Since both UD and UniMorph aim at broad coverage of cross-linguistic diversity (cf. de Marneffe et al., 2021; Batsuren et al., 2022), the sources of data are quite heterogeneous in both datasets. This can make typological research more difficult, especially for corpus-based work with UD. The text types and genres used in UD vary considerably from language to language and include spoken and written data from fiction, non-fiction, legal texts, blogs, news, medical texts and Wikipedia entries. This is not a problem per se, as the text types are clearly and transparently documented.16 However, some languages are only represented in UD with examples from grammars and language textbooks.17 Such data may be useful for some questions, but they are not for most typological research questions that use corpus data to take into account the distribution of certain linguistic features in actual language use. This means that linguists working with the data need to carefully consider the final selection of languages in the dataset, use appropriate controls in statistical models and potentially exclude certain languages (which is not always done in quantitative typological studies). This concern is shared by others working in corpus-based typology: Levshina et al. (2023, p. 858) argue, for example, that the genres and modalities of data must be carefully curated to ensure the comparability of cross-linguistic samples.

The push for cross-linguistic coverage in both UD and UniMorph has also resulted in language datasets of differing quality. This is not a priori an issue. However, the linguist working with the data needs to evaluate the quality of a dataset (most likely manually) to decide what to use. In other words, data that seem ready for reuse in quantitative typology are not necessarily usable without a substantial amount of additional (automatic and manual) work. In the case of UD, lower quality data may involve data with missing annotations or automatically annotated data with errors. In the case of UniMorph, the issue of data quality is potentially somewhat more severe.18 Many of the datasets appear to be automatically extracted without any manual checks or further data curation. This leads to problems for reusers on several levels. First, datasets with a large number of lemmas such as English, Finnish or Latin tend to contain a lot of entries (up to 15,000 observations) that are not real words and have to be removed semi-manually. Second, most datasets are taken from Wiktionary where sources are not indicated consistently. It is praiseworthy that UniMorph provides full inflection paradigms for large numbers of lexemes that would otherwise not be available for many languages, but the data quality of crowdsourced dictionaries needs careful manual evaluation by the linguist working with these data. An example of annotation mistakes that had to be eliminated semi-automatically by language is a final “!” in imperative forms for Gujarati and Turkish. Third, UniMorph largely retains original annotation labels, which leads to inconsistencies across languages, e.g., “1+INCL” and “1;INCL” for first person inclusive, or “NDEF” and “INDF” for indefinite, or “3;SG” and “SG;3” for third person singular. Another consequence is that linguistic labels are not necessarily comparable across languages, since they are not designed as cross-linguistic but as language-specific categories (see Haspelmath (2010) for an in-depth discussion of this issue).19 While some language-specific labels are retained, others are converted to “LGSPEC” (language-specific), which obscures the function of a given form and makes it difficult to reconstruct the original value. Both issues were resolved in Becker (2024) by careful manual corrections and adaptations of the UniMorph labels to make them more consistent and typologically informative. For more details about this manual step of data cleaning, see “preprocessing.txt” in the supplementary files of Becker (2024) at https://osf.io/e48qc/files/98wuv (Accessed 2026-06-26).

All of the issues brought up in this section are not meant to criticise UD or UniMorph as resources; we are aware of how much effort and time needs to be put in to build such corpora and databases. This section rather serves as a reminder for the linguistic community that reusing corpora and cross-linguistic databases for new projects still requires much care and manual evaluation in order to understand the data and to decide which parts are useful for a given study and which parts are better to exclude.

(6) Recommendations and good practices

If our work has taught us anything, it is that there is still a considerable gap between published language resources and reusable language resources (cf. Weber, 2021; Yi et al., 2022; Babinski et al., 2022). This gap may be due to differing motivations for sharing data (Pronk, 2019) – are we making data available so that others can verify our analyses (cf. Berez-Kroeker et al., 2018), or are we actively making data available for others to use (Woodbury, 2014)? When the work put into building datasets is already undervalued (Thieberger et al., 2016), the work needed to curate the data to make them reusable can be argued to be even less so (Weber, 2021). Both of these activities need to be equally recognised, incorporated into data management workflows, and – crucially – funded, if we are to have a thriving data reuse ecosystem in linguistic typology and beyond.

Corpus collections such as Multi-CAST and DoReCo are great examples of how data collected for language documentation purposes can be reused for typological studies. However, the amount of work needed to prepare a corpus for each of these projects is massive. This work tends to primarily fall on the documenter (or requires their active input, at the very least). At the other end of the spectrum, the work of curating the entire collection is not to be underestimated either. Curators must try to ensure comparability across individual language datasets and different versions of the whole collection as well as tackle different conditions of use between corpora. A considerable amount of work is also invested into making the collections reusable, such as providing documentation about annotation guidelines (e.g., Haig & Schnell, 2014). This labour is not necessarily reflected in an uptake of these resources outside of the original research teams (cf. Pronk, 2019). This is in part due to a time lag associated with data reuse in general (Piwowar & Vision, 2013) but also reflects the still considerable amount of labour that is needed to reuse these data. Lowering the barrier to reuse via measures such as providing detailed documentation of the datasets allows for a more efficient and wider use of data, especially for uses not originally anticipated by the creators (e.g., the use of DoReCo for machine learning benchmarks in Zhu et al. (2024) and Samir et al. (2024)). This documentation does not always have to rest on the shoulders of the curators – the reusers can help here too. This could be as simple as providing a reproducible workflow as supplementary material with any article which reuses data.

Everyone who contributes to building a data reuse ecosystem should receive recognition for their work. The main currency for contribution in linguistics is via scholarly citation. This is reflected by the importance laid upon correct citation of data in linguistics (Berez-Kroeker et al., 2018; Andreassen et al., 2019), a prerequisite for any kind of data reuse. The thorny issue here is authorship. Practices of authorship in philological traditions tend to centre around more traditional intellectual outputs, such as articles, and little thought has been paid to data as a research product (Wallis & Borgmann, 2011).20 This often results in valuable labour by student research assistants or community members being undervalued and/or underrecognised. There are also no clear workflows for how we should handle the authorship of derivative products and how we can ensure that credit is paid to the originators of the reused dataset, although paper-based solutions such as recognition in acknowledgement sections or CRediT statements are gaining popularity. Here we find the example of UD to be of particular inspiration: all contributors of language datasets are included as authors of the database as a whole.

A reuse ecosystem similarly requires active buy-in from the community. Best Open Data practices advocate for explicit statements of the creator’s wishes for data reuse as well as the use of non-restrictive licenses (Molloy, 2011). However, there is currently little choice of licenses to best accommodate possible reuse practices by third parties, especially reusers outside for-profit enterprise, despite the push for Open Data on the part of funding agencies. If anything, it is the opposite – linguists have recently started to reach more and more for restrictive reuse licenses, in view of the data scraping practices by companies training Large Language Models. To what extent license choice reflects an informed (and community-driven) decision remains an open question (Hagedorn et al., 2011; Holton et al., 2022; see also Barwick et al., 2019).

Our main proposal here is simple: we need to extend the conversation around data use to data reuse in typology and linguistics more broadly. The incredible uptake of data citation practices over the past 10 years is a testament to how quickly practices in the field can change (Berez-Kroeker et al., 2018; Andreassen et al., 2019). To embrace the potentials of data reuse, we require similar concerted community efforts to share know-how, with the ultimate goal of developing a field-wide ethical and sustainable culture of reuse. This may be as simple as including documentation about a dataset when publishing it, to give a better idea of the research context the data were collected in (cf. Himmelmann, 1998, pp. 170-171; see also Dobrin & Schwartz (2021)). Transparency around data processing can increase the likelihood of data reuse, especially if reusable code is provided or protocol for a reproducible workflow is accessible (cf. Appendix D of Becker & Guzmán Naranjo, 2025). Landing pages for datasets should prominently include information about licenses and citations, including for potential derivative products, even if this leads to redundant information on a site. These sorts of moves towards making data more reusable do involve an extra investment of effort and time, both in the short-term and in the long-term, and may even involve a modicum of risk for those who do not have stable employment, if problems are raised with their dataset. These are very legitimate worries and must be addressed in the conversation around data reuse. What we would like to reiterate is simply that every step towards a culture of reuse counts, whether it be the publication of a dataset or the inclusion of more details about how a dataset was created or annotated in an article.

More ambitious community initiatives can also be taken. Teachers can involve their students in assessing existing language resources à la Babinski et al. (2022) or improving the quality of archives by helping in their curation (cf. Weber, 2021). Educators can recommend standard annotation conventions or reuse Open Educational Materials with predefined templates (e.g., The ELAR Team, 2022). Hackathons or shared tasks could be organised to create or improve documentation or coverage of datasets, like how UniMorph improvements are driven by regular shared tasks. We can write more (third-party) reviews of datasets, especially for documentation corpora, which explicitly address their reusability potential (cf. Haspelmath & Michaelis, 2014; Thieberger et al., 2016; Thieberger, 2017). Initiatives for helping dataset-internal consistency such as RefCo (Aznar & Seifart, 2020, 2022) or validation scripts can be supported and expanded upon. The list is endless and we encourage our colleagues to consider incorporating these kinds of actions the next time they plan a course or write a grant.

There is a lot of unused and underused data out there which are perfect for typological studies. With buy-in from the larger community and some elbow grease, we believe that it is possible to make much of this data (more) reusable. It will simply take some conversation, some creativity and a commitment to making collective progress one step at a time.

Additional File

The additional file for this article can be found as follows:

Supplementary Materials

Detailed descriptions of the datasets referenced in the article following the standards set out by the Journal of Open Humanities Data. DOI: https://doi.org/10.5334/johd.530.s1

Notes

[1] https://www.language-archives.org/ (Accessed 2026-06-26).

[2] https://catalog.paradisec.org.au/ (Accessed 2026-06-26).

[3] https://pangloss.cnrs.fr/?lang=en (Accessed 2026-06-26).

[4] https://www.elararchive.org/ (Accessed 2026-06-26).

[5] https://glottolog.org/langdoc/status (Accessed 2026-06-26).

[6] See the Supplementary Materials for more details on the contents of each corpus.

[7] The relevant files have now been made open access by PARADISEC. The message I saw is still present in the data access details of the corpus today: https://catalog.paradisec.org.au/collections/TANI (Accessed 2026-06-26).

[8] The script “network-analysis.R” from the supplementary materials for Peck (2025) shows how I practically reconciled the difference in annotation schema for comparative work. This script can be found at: https://osf.io/45ge7/ (Accessed 2026-06-26).

[10] Not all archives are properly captured by the search, either. At the time of searching, the OLAC records for ELAR were out-of-date by two years, presumably due to the move of the archive from the UK to Germany. The status of data harvesting for each archive on OLAC can be viewed on their participating archives page: https://www.language-archives.org/archives (Accessed 2026-06-26).

[11] Supplementary materials for Peck & Becker (2024), including a document detailing our data annotation and processing workflow in Praat, ELAN and R, can be found at: https://osf.io/juz9r/. (Accessed 2026-06-26).

[12] Details of where to access the datasets can be found in the references and supplementary materials.

[14] Note that not all audio is hosted on DoReCo; in many instances, the audio is hosted in linguistic archives.

[15] UD currently only includes one indigenous language from Australia (Warlpiri) and three from North America (Gwichin, Highland Puebla Nahuatl and Kiche), which is not representative of the linguistic diversity in these areas. Its coverage of South American indigenous languages is slightly better at 16 included languages. The bias towards standard languages from Eurasia is also present in UniMorph, although it has a slightly better cross-linguistic coverage, which could also be due to the fact that it does not include corpus data but inflectional paradigms. Australia is severely underrepresented in both resources, though; UniMorph includes only two Australian languages (Kunwinjku, Murrinhpatha).

[16] The issue of heterogeneous text types and genres is also discussed by the curators of UD, e.g., here: https://universaldependencies.org/contributing/genres.html (Accessed 2026-06-26).

[17] Examples of languages in UD whose corpora are based on examples from grammars and textbooks include: Azerbaijani, Bengali, Cebuano, Chintang, Macedonian, Malayalam and Warlpiri. This may be justified for some lesser-studied languages, but this list also includes standardized, written, national languages, for which more naturalistic data should and may be available.

[18] This refers to UniMorph versions 2.0 (Kirov et al., 2018) and 3.0 (McCarthy et al., 2020), which are the versions used for our work. UniMorph 4.0 (Batsuren et al., 2022) has since then been released and it is possible that the issues we mention here have since been addressed. Note that we describe UniMorph 4.0 in our Supplementary Materials.

[19] The ongoing PARALEX project (Beniamine et al., 2023) aims to address this issue by proposing a typologically informed standard for cross-linguistic morphological datasets.

[20] This seems to be slowly changing, though, with more linguistic (and typological) journals offering shorter publication formats for new language data, datasets and other resources such as software.

Acknowledgements

Thank you to Laura Schleicher, Jessica K. Ivani and two anonymous reviewers for their comments, which have helped us improve the paper immensely.

Author Contributions

Naomi Peck – conceptualisation, writing – original draft, writing – review & editing.

Laura Becker – conceptualisation, writing – original draft, writing – review & editing.

DOI: https://doi.org/10.5334/johd.530 | Journal eISSN: 2059-481X
Language: English
Page range: 96 - 96
Submitted on: Feb 27, 2026
Accepted on: Jun 22, 2026
Published on: Jul 20, 2026
Published by: Ubiquity Press
In partnership with: Paradigm Publishing Services

© 2026 Naomi Peck, Laura Becker, published by Ubiquity Press
This work is licensed under the Creative Commons Attribution 4.0 License.