Skip to main content
Have a personal or library account? Click to login
Reusing Language Corpora for Token-Based Typology Cover

Reusing Language Corpora for Token-Based Typology

By:  and    
Open Access
|Jul 2026

References

  1. Andreassen, H. N., Berez-Kroeker, A. L., Collister, L., Conzett, P., Cox, C., Smedt, K. D., McDonnell, B., & Research Data Alliance Linguistic Data Interest Group. (2019). Tromsø recommendations for citation of research data in linguistics. 10.15497/RDA00040
  2. Aznar, J., & Seifart, F. (2020). RefCo: An initiative to develop a set of quality criteria for fieldwork corpora. In T. Poibeau, Y. Parmentier, & E. Schang (Eds.), Actes des 2èmes journées scientifiques du Groupement de Recherche Linguistique Informatique Formelle et de Terrain (LIFT) (pp. 95101). CNRS. https://hal.science/hal-03047143
  3. Aznar, J., & Seifart, F. (2022). The RefCo Toolkit. Zenodo. 10.5281/zenodo.6470807
  4. Babinski, S., Jewell, J., Haakman, K., Kim, J., Lake, A., Yi, I., & Bowern, C. (2022). How usable are digital collections for endangered languages? A review. Proceedings of the Linguistic Society of America, 7(1), 5219. 10.3765/plsa.v7i1.5219
  5. Barwick, L., Green, J., Vaarzon-Morel, P., & Zissermann, K. (2019). Conundrums and consequences: Doing digital archival returns in Australia. In L. Barwick, J. Green, & P. Vaarzon-Morel (Eds.), Archival Returns: Central Australia and Beyond (pp. 127). University of Hawai’i Press. https://hdl.handle.net/10125/24875
  6. Batsuren, K., Goldman, O., Khalifa, S., Habash, N., Kieraś, W., Bella, G., Leonard, B., Nicolai, G., Gorman, K., Ate, Y. G., Ryskina, M., Mielke, S., Budianskaya, E., El-Khaissi, C., Pimentel, T., Gasser, M., Lane, W. A., Raj, M., Coler, M., … Vylomova, E. (2022). UniMorph 4.0: Universal Morphology. In N. Calzolari, F. Béchet, P. Blache, K. Choukri, C. Cieri, T. Declerck, S. Goggi, H. Isahara, B. Maegaard, J. Mariani, H. Mazo, J. Odijk, & S. Piperidis (Eds.), Proceedings of the Thirteenth Language Resources and Evaluation Conference (pp. 840855). European Language Resources Association. https://aclanthology.org/2022.lrec-1.89/
  7. Becker, L. (2024). Zero marking in inflection: A token-based approach. Journal of Language Modelling, 12(2), 349413. 10.15398/jlm.v12i2.361
  8. Becker, L. (n.d.). Frequency effects in verbal argument indexing: A spoken typology approach. https://laurabecker.gitlab.io/papers/becker_person_indexes.pdf [In revisions].
  9. Becker, L., & Guzmán Naranjo, M. (2025). Replication and methodological robustness in quantitative typology. Linguistic Typology, 29(3), 463505. 10.1515/lingty-2023-0076
  10. Becker, L., Peck, N., Asvarova, S., Kaminski, M., Klöckner, C., Postawa, R., & Spasovski, A. (2025). Pausing and prosody-syntax mismatches: A spoken typology approach. Talk presented at Syntax of the World’s Languages X, University of Potsdam. https://laurabecker.gitlab.io/presentations/pausing_swl_potsdam.pdf
  11. Beniamine, S., Anderson, C., Carroll, M., Guzmán Naranjo, M., Herce, B., Pellegrini, M., Round, E., Sims-Williams H., & Tresoldi, T. (2023). Paralex: a DeAR standard for rich lexicons of inflected forms. Talk presented at Presentation at International Symposium of Morphology, Université de Lorraine, Nancy.
  12. Berdicevskis, A., Çöltekin, Ç., Ehret, K., von Prince, K., Ross, D., Thompson, B., Yan, C., Demberg, V., Lupyan, G., Rama, T., & Bentz, C. (2018). Using Universal Dependencies in cross-linguistic complexity research. In M.-C. de Marneffe, T. Lynn, & S. Schuster (Eds.), Proceedings of the Second Workshop on Universal Dependencies (UDW 2018) (pp. 817). Association for Computational Linguistics. 10.18653/v1/W18-6002
  13. Berez-Kroeker, A. L., Gawne, L., Kung, S. S., Kelly, B. F., Heston, T., Holton, G., Pulsifer, P., Beaver, D. I., Chelliah, S., Dubinsky, S., Meier, R. P., Thieberger, N., Rice, K., & Woodbury, A. C. (2018). Reproducible research in linguistics: A position statement on data citation and attribution in our field. Linguistics, 56(1), 118. 10.1515/ling-2017-0032
  14. Blum, F., Paschen, L., Forkel, R., Fuchs, S., & Seifart, F. (2024). Consonant lengthening marks the beginning of words across a diverse sample of languages. Nature Human Behaviour, 112. 10.1038/s41562-024-01988-4
  15. Bodt, T. A. (2018a). Duhumbi Personal Narratives–Sound files. 10.5281/ZENODO.1406179
  16. Bodt, T. A. (2018b). Duhumbi Personal Narratives–Transcribed, parsed, glossed, translated text files [Dataset]. Zenodo. 10.5281/ZENODO.1406176
  17. Bodt, T. A. (2018c). Duhumbi Procedural Texts–Sound files. 10.5281/ZENODO.1406157
  18. Bodt, T. A. (2018d). Duhumbi Procedural Texts–Transcribed, parsed, glossed, translated text files [Dataset]. Zenodo. 10.5281/ZENODO.1406154
  19. Bodt, T. A. (2018e). Duhumbi Stories–Sound files. 10.5281/ZENODO.1400495
  20. Bodt, T. A. (2018f). Duhumbi Stories–Transcribed, parsed, glossed, translated text files (Version 1) [Dataset]. Zenodo. 10.5281/ZENODO.1400505
  21. Cowell, A. (2024). Arapaho DoReCo dataset. In F. Seifart, L. Paschen, & M. Stave (Eds.), Language documentation reference corpus (DoReCo) 2.0. Laboratoire Dynamique Du Langage (UMR5596, CNRS & Université Lyon 2). 10.34847/nkl.36f5r1b6
  22. de Marneffe, M.-C., Manning, C. D., Nivre, J., & Zeman, D. (2021). Universal Dependencies. Computational Linguistics, 47(2), 255308. 10.1162/coli_a_00402
  23. Desai, M. A., Pasquetto, I. V., Jacobs, A. Z., & Card, D. (2024). An archival perspective on pretraining data. Patterns, 5(4), 100966. 10.1016/j.patter.2024.100966
  24. Dobrin, L. M., & Schwartz, S. (2021). The social lives of linguistic legacy materials. Language Documentation and Description, 21, 136.
  25. Döhler, C. (2024). Komnzo DoReCo dataset. In F. Seifart, L. Paschen, & M. Stave (Eds.), Language documentation reference corpus (DoReCo) 2.0. Laboratoire Dynamique Du Langage (UMR5596, CNRS & Université Lyon 2). 10.34847/nkl.c5e6dudv
  26. Dryer, M., & Haspelmath, M. (Eds.). (2013). WALS Online (v2020.4). Zenodo. 10.5281/zenodo.13950591
  27. Forkel, R., & Hammarström, H. (2022). Glottocodes: Identifiers linking families, languages and dialects to comprehensive reference information. Semantic Web, 13(6), 917924. 10.3233/SW-212843
  28. Gawne, L., & Berez-Kroeker, A. L. (2018). Reflections on reproducible research. In B. McDonnell, A. L. Berez-Kroeker, & G. Holton (Eds.), Reflections on Language Documentation 20 Years after Himmelmann 1998 (pp. 2232). https://hdl.handle.net/10125/24805
  29. Gebru, T., Morgenstern, J., Vecchione, B., Vaughan, J. W., Wallach, H., Daumé III, H., & Crawford, K. (2021). Datasheets for Datasets (No. arXiv:1803.09010). arXiv. https://arxiv.org/abs/1803.09010
  30. Gerdes, K., Kahane, S., & Chen, X. (2021). Typometrics: From implicational to quantitative universals in word order typology. Glossa, 6(1), 117. 10.5334/gjgl.764
  31. Gusev, V., Klooster, T., Wagner-Nagy, B., & Arkhipov, A. (2024). Kamas DoReCo dataset. In F. Seifart, L. Paschen, & M. Stave (Eds.), Language Documentation Reference Corpus (DoReCo) 2.0. Lyon: Laboratoire Dynamique Du Langage (UMR5596, CNRS & Université Lyon 2). Laboratoire Dynamique Du Langage (UMR5596, CNRS & Université Lyon 2). 10.34847/nkl.cdd8177b
  32. Guzmán Naranjo, M., & Becker, L. (2018). Quantitative word order typology with UD. Proceedings of the 17th International Workshop on Treebanks and Linguistic Theories (TLT 2018), December 13–14, 2018, Oslo University, Norway, 91104.
  33. Guzmán Naranjo, M., & Becker, L. (2021). Coding efficiency in nominal inflection: Expectedness and type frequency effects. Linguistics Vanguard, 7(s3), 20190075. 10.1515/lingvan-2019-0075
  34. Guzmán Naranjo, M., & Becker, L. (2022). Statistical bias control in typology. Linguistic Typology, 26(3), 605670. 10.1515/lingty-2021-0002
  35. Hagedorn, G., Mietchen, D., Morris, R. A., Agosti, D., Penev, L., Berendsohn, W. G., & Hobern, D. (2011). Creative Commons licenses and the non-commercial condition: Implications for the re-use of biodiversity information. ZooKeys, 150, 127149. 10.3897/zookeys.150.2189
  36. Haig, G., & Schnell, S. (2014). Annotations using GRAID (Grammatical Relations and Animacy in Discourse): Introduction and guidelines for annotators, Version 7.0.
  37. Haig, G., & Schnell, S. (2016). The discourse basis of ergativity revisited. Language, 92(3), 591618. 10.1353/lan.2016.0049
  38. Haig, G., & Schnell, S. (2021). Multi-CAST: Multilingual corpus of annotated spoken texts. https://Multi-CAST.aspra.uni-bamberg.de/
  39. Hardwicke, T. E., Mathur, M. B., MacDonald, K., Nilsonne, G., Banks, G. C., Kidwell, M. C., Hofelich Mohr, A., Clayton, E., Yoon, E. J., Henry Tessler, M., Lenne, R. L., Altman, S., Long, B., & Frank, M. C. (2018). Data availability, reusability, and analytic reproducibility: Evaluating the impact of a mandatory open data policy at the journal Cognition. Royal Society Open Science, 5(8), 180448. 10.1098/rsos.180448
  40. Haspelmath, M. (2010). Comparative concepts and descriptive categories in crosslinguistic studies. Language, 86(3), 663687. 10.1353/lan.2010.0021
  41. Haspelmath, M., & Michaelis, S. M. (2014, March 11). Annotated corpora of small languages as refereed publications: A vision [Billet]. Diversity Linguistics Comment. 10.58079/nst3
  42. Haude, K. (2024). Movima DoReCo dataset. In F. Seifart, L. Paschen, & M. Stave (Eds.), Language documentation reference corpus (DoReCo) 2.0. Laboratoire Dynamique Du Langage (UMR5596, CNRS & Université Lyon 2). 10.34847/nkl.da42xf67
  43. Himmelmann, N. P. (1998). Documentary and descriptive linguistics. Linguistics, 36(1), 161195. 10.1515/ling.1998.36.1.161
  44. Holton, G., Leonard, W. Y., & Pulsifer, P. L. (2022). Indigenous Peoples, Ethics, and Linguistic Data. In A. L. Berez-Kroeker, B. McDonnell, E. Koller & L. B. Collister (eds.), The Open Handbook of Linguistic Data Management, pp. 4960. The MIT Press. 10.7551/mitpress/12200.001.0001.
  45. Kim, S.-U. (2024). Jejuan DoReCo dataset. In F. Seifart, L. Paschen, & M. Stave (Eds.), Language documentation reference corpus (DoReCo) 2.0. Laboratoire Dynamique Du Langage (UMR5596, CNRS & Université Lyon 2). 10.34847/nkl.06ebrk38
  46. Kirov, C., Cotterell, R., Sylak-Glassman, J., Walther, G., Vylomova, E., Xia, P., Faruqui, M., Mielke, S. J., McCarthy, A., Kübler, S., Yarowsky, D., Eisner, J., & Hulden, M. (2018). UniMorph 2.0: Universal Morphology. In N. Calzolari, K. Choukri, C. Cieri, T. Declerck, S. Goggi, K. Hasida, H. Isahara, B. Maegaard, J. Mariani, H. Mazo, A. Moreno, J. Odijk, S. Piperidis, & T. Tokunaga (Eds.), Proceedings of the Eleventh International Conference on Language Resources and Evaluation (LREC 2018). European Language Resources Association (ELRA) https://aclanthology.org/L18-1293/
  47. Levshina, N. (2016). Why we need a token-based typology: A case study of analytic and lexical causatives in fifteen European languages. Folia Linguistica, 50(2), 507542. 10.1515/flin-2016-0019
  48. Levshina, N. (2019). Token-based typology and word order entropy: A study based on Universal Dependencies. Linguistic Typology, 23(3), 533572. 10.1515/lingty-2019-0025
  49. Levshina, N. (2022). Corpus-based typology: Applications, challenges and some solutions. Linguistic Typology, 26(1), 129160. 10.1515/lingty-2020-0118
  50. Levshina, N., Namboodiripad, S., Allassonnière-Tang, M., Kramer, M., Talamo, L., Verkerk, A., Wilmoth, S., Rodriguez, G. G., Gupton, T. M., Kidd, E., Liu, Z., Naccarato, C., Nordlinger, R., Panova, A., & Stoynova, N. (2023). Why we need a gradient approach to word order. Linguistics, 61(4), 825883. 10.1515/ling-2021-0098
  51. McCarthy, A. D., Kirov, C., Grella, M., Nidhi, A., Xia, P., Gorman, K., Vylomova, E., Mielke, S. J., Nicolai, G., Silfverberg, M., Arkhangelskiy, T., Krizhanovsky, N., Krizhanovsky, A., Klyachko, E., Sorokin, A., Mansfield, J., Ernštreits, V., Pinter, Y., Jacobs, C. L., … Yarowsky, D. (2020). UniMorph 3.0: Universal Morphology. In N. Calzolari, F. Béchet, P. Blache, K. Choukri, C. Cieri, T. Declerck, S. Goggi, H. Isahara, B. Maegaard, J. Mariani, H. Mazo, A. Moreno, J. Odijk, & S. Piperidis (Eds.), Proceedings of the 12th Language Resources and Evaluation Conference (pp. 39223931). European Language Resources Association. https://aclanthology.org/2020.lrec-1.483
  52. Molloy, J. C. (2011). The Open Knowledge Foundation: Open Data Means Better Science. PLoS Biology, 9(12), e1001195. 10.1371/journal.pbio.1001195
  53. Nivre, J., de Marneffe, M. -C., Ginter, F., Hajič, J., Manning, C., Pyysalo, S., Schuster, S., Tyers, F., & Zeman, D. (2020). Universal Dependencies v2: An evergrowing multilingual treebank collection. In N. Calzolari, F. Béchet, P. Blache, K. Choukri, C. Cieri, T. Declerck, S. Goggi, H. Isahara, B. Maegaard, J. Mariani, H. Mazo, A. Moreno, J. Odijk, & S. Piperidis (Eds.), Proceedings of the Twelfth Language Resources and Evaluation Conference (pp. 40344043). European Language Resources Association (ELRA).
  54. Paschen, L. (2025). Acoustic disambiguation of homophonous morphs is exceptional. Journal of Linguistics, 133. 10.1017/S0022226725100777
  55. Peck, N. (2025). (Re)contextualising serialisation: Multiverbal constructions in the East Himalaya [PhD Thesis, University of Freiburg]. 10.6094/UNIFR/274498
  56. Peck, N., & Becker, L. (2024). Syntactic Pausing? Re-examining the associations. Linguistics Vanguard, 10(1), 223237. 10.1515/lingvan-2022-0156
  57. Piwowar, H. A., & Vision, T. J. (2013). Data reuse and the open data citation advantage. PeerJ, 1, e175. 10.7717/peerj.175
  58. Post, M., & Modi, Y. (2001). The Tani Languages. PARADISEC. 10.4225/72/56E979C0C510A
  59. Pronk, T. E. (2019). The Time Efficiency Gain in Sharing and Reuse of Research Data. Data Science Journal, 18(10), 18. 10.5334/dsj-2019-010
  60. R Core Team. (2024). R: A language and environment for statistical computing [Manual]. https://www.R-project.org/
  61. Samir, F., Ahn, E. P., Prakash, S., Soskuthy, M., Shwartz, V., & Zhu, J. (2024). Efficiently Identifying Low-Quality Language Subsets in Multilingual Datasets: A Case Study on a Large-Scale Multilingual Audio Dataset (No. arXiv:2410.04292). arXiv. 10.48550/arXiv.2410.04292
  62. Schnell, S., Schiborr, N. N., & Haig, G. (2021). Efficiency in discourse processing: Does morphosyntax adapt to accommodate new referents? Linguistics Vanguard, 7(s3). 10.1515/lingvan-2019-0064
  63. Schnell, S., & Schiborr, N. N. (2022). Crosslinguistic Corpus Studies in Linguistic Typology. Annual Review of Linguistics, 8(1), 171191. 10.1146/annurev-linguistics-031120-104629
  64. Seifart, F. (2021). Combining documentary linguistics and corpus phonetics to advance corpus-based typology. Language Documentation & Conservation, SP25, 115139.
  65. Seifart, F., Paschen, L., & Stave, M. (Eds.). (2024). Language Documentation Reference Corpus (DoReCo) 2.0. Laboratoire Dynamique Du Langage (UMR5596, CNRS & Université Lyon 2). 10.34847/nkl.7cbfq779
  66. Stave, M., Paschen, L., Pellegrino, F., & Seifart, F. (2021). Optimization of morpheme length: A cross-linguistic assessment of Zipf’s and Menzerath’s laws. Linguistics Vanguard, 7(s3), 20190076. 10.1515/lingvan-2019-0076
  67. Teo, A. (2024). Sümi DoReCo dataset. In F. Seifart, L. Paschen, & M. Stave (Eds.), Language documentation reference corpus (DoReCo) 2.0. Laboratoire Dynamique Du Langage (UMR5596, CNRS & Université Lyon 2). 10.34847/nkl.5ad4t01p
  68. The ELAR Team. (2022). Resources for Language Documentation and Archiving [Dataset]. https://hdl.handle.net/2196/985b23d3-d948-4566-a6cd-84f545926d16
  69. Thieberger, N. (2017). LD&C possibilities for the next decade. Language Documentation & Conservation, 11, 14.
  70. Thieberger, N., Margetts, A., Morey, S., & Musgrave, S. (2016). Assessing Annotated Corpora as Research Output. Australian Journal of Linguistics, 36(1), 121. 10.1080/07268602.2016.1109428
  71. van Driem, G. (2014). Trans-Himalayan. In T. Owen-Smith & N. Hill (Eds.), Trans-Himalayan Linguistics: Historical and Descriptive Linguistics of the Himalayan Area (pp. 1140). Walter de Gruyter. 10.1515/9783110310832.11
  72. Vanhove, M. (2024). Beja DoReCo dataset. In F. Seifart, L. Paschen, & M. Stave (Eds.), Language documentation reference corpus (DoReCo) 2.0. Laboratoire Dynamique Du Langage (UMR5596, CNRS & Université Lyon 2). 10.34847/nkl.edd011t1
  73. Vollmer, M. (2023). Comparing zero and referential choice in eight languages with a focus on Mandarin Chinese. 10.1075/sl.21072.vol
  74. von Prince, K., & Nordhoff, S. (2020). An Empirical Evaluation of Annotation Practices in Corpora from Language Documentation. In N. Calzolari, F. Béchet, P. Blache, K. Choukri, C. Cieri, T. Declerck, S. Goggi, H. Isahara, B. Maegaard, J. Mariani, H. Mazo, A. Moreno, J. Odijk, & S. Piperidis (Eds.), Proceedings of the Twelfth Language Resources and Evaluation Conference (pp. 27782787). European Language Resources Association. https://aclanthology.org/2020.lrec-1.338/
  75. Wallis, J. C., & Borgman, C. L. (2011). Who is responsible for data? An exploratory study of data authorship, ownership, and responsibility. Proceedings of the American Society for Information Science and Technology, 48(1), 110. 10.1002/meet.2011.14504801188
  76. Weber, T. (2021). The Curation of Language Data as a Distinct Academic Activity: A Call to Action for Researchers, Educators, Funders, and Policymakers. Journal of Open Humanities Data, 7(28), 110. 10.5334/johd.51
  77. Wittenburg, P., Brugman, H., Russel, A., Klassmann, A., & Sloetjes, H. (2006). ELAN : a professional framework for multimodality research. In 5th International Conference on Language Resources and Evaluation (LREC 2006), 15561559. (5 January, 2026). 10.63317/5pwa5zpssv4z
  78. Woodbury, A. C. (2014). Archives and audiences: Toward making endangered language documentations people can read, use, understand, and admire. Language Documentation and Description, 12, 1936.
  79. Yi, I., Lake, A., Kim, J., Haakman, K., Jewell, J., Babinski, S., & Bowern, C. (2022). Accessibility, Discoverability, and Functionality: An Audit of and Recommendations for Digital Language Archives. Journal of Open Humanities Data, 8(0), 10. 10.5334/johd.59
  80. Zeman, D., Nivre, J., Abrams, M., Ackermann, E., Aepli, N., Aghaei, H., Agić, Ž., Ahmadi, A., Ahrenberg, L., Ajede, C. K., Aleksandravičiūtė, G., Alfina, I., Antonsen, L., Aplonova, K., Aquino, A., Aragon, C., Aranzabe, M. J., Arıcan, B. N., Arnardóttir, órunn, … Ziane, R. (2021). Universal dependencies 2.9. https://hdl.handle.net/11234/1-4611
  81. Zhu, J., Yang, C., Samir, F., & Islam, J. (2024). The taste of IPA: Towards open-vocabulary keyword spotting and forced alignment in any language. In K. Duh, H. Gomez, & S. Bethard (Eds.), Proceedings of the 2024 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers) (pp. 750772). Association for Computational Linguistics. 10.18653/v1/2024.naacl-long.43
DOI: https://doi.org/10.5334/johd.530 | Journal eISSN: 2059-481X
Language: English
Page range: 96 - 96
Submitted on: Feb 27, 2026
Accepted on: Jun 22, 2026
Published on: Jul 20, 2026
Published by: Ubiquity Press
In partnership with: Paradigm Publishing Services

© 2026 Naomi Peck, Laura Becker, published by Ubiquity Press
This work is licensed under the Creative Commons Attribution 4.0 License.