
A CLDF-Standardized Lexical Dataset for Quantitative Diachronic Research of South American Languages
Abstract
This paper presents a lexical dataset for 299 South American languages, including 15 reconstructed proto-languages. The dataset comprises 25,203 lexical forms compiled from multiple publicly available sources and standardized according to the Cross-Linguistic Data Formats (CLDF) Wordlist specification. Language varieties and concepts were harmonized using Glottolog and Concepticon identifiers to improve interoperability across heterogeneous lexical resources. The dataset was designed to support quantitative diachronic research, including automatic cognate detection and phylogenetic inference. By addressing inconsistencies in transcription, naming conventions, and glossing practices, the dataset facilitates data reuse and large-scale comparative research for one of the world’s most linguistically diverse regions.
© 2026 Álvaro Arozarena Gómez, Annemarie Verkerk, Marc Tanti, published by Ubiquity Press
This work is licensed under the Creative Commons Attribution 4.0 License.