Skip to main content
Have a personal or library account? Click to login
Reusing Language Corpora for Token-Based Typology Cover

Reusing Language Corpora for Token-Based Typology

By:  and    
Open Access
|Jul 2026

Abstract

Linguistic typology as a discipline is built upon a culture of data reuse. Despite the uptake of large-scale typological databases such as WALS for reuse studies, fewer studies have leveraged publicly available corpora or lesser-known database resources. In this contribution to the special collection, we share our experiences as usage-based typologists who frequently reuse data for token-based typology. Through three case studies involving different types of language datasets, we critically reflect upon the reusability of data in linguistic typology. By reflecting on our experiences, we wish to engage more linguists in the conversation about how to create a more data reuse-friendly ecosystem in linguistic typology, and linguistics more broadly.

DOI: https://doi.org/10.5334/johd.530 | Journal eISSN: 2059-481X
Language: English
Page range: 96 - 96
Submitted on: Feb 27, 2026
Accepted on: Jun 22, 2026
Published on: Jul 20, 2026
Published by: Ubiquity Press
In partnership with: Paradigm Publishing Services

© 2026 Naomi Peck, Laura Becker, published by Ubiquity Press
This work is licensed under the Creative Commons Attribution 4.0 License.