Abstract
This paper looks at opportunities and challenges of language dataset reuse in methodological research, using the development of down-sampling designs for corpus-linguistic work as a case study. This research is concerned with the construction and evaluation of strategies for down-sizing data that are too large given the resource constraints confronting a study. Work in this line of inquiry is heavily dependent on real datasets to field-test and compare techniques, which makes it a prime example of how open data can serve as a catalyst for methodological progress. In the reuse case discussed here, a methodological study made use of an informally published dataset. The original data were revised and modified, and the derived dataset archived in the Tromsø Repository of Language and Linguistics (TROLLing). The current paper draws attention to a number of questions that may arise when repurposing language data, including issues related to licensing and the question of how to determine authorship credit. It also demonstrates how depositors and reusers of data can profit from a careful review process, which is ideally assisted by trained data curators. The case study points to TROLLing as an exemplary repository that offers quality control and substantive feedback on a host of aspects including dataset documentation as well as legal and research-ethical matters.
© 2026 Lukas Sönning, published by Ubiquity Press
This work is licensed under the Creative Commons Attribution 4.0 License.
