Skip to main content
Have a personal or library account? Click to login
A CLDF-Standardized Lexical Dataset for Quantitative Diachronic Research of South American Languages Cover

A CLDF-Standardized Lexical Dataset for Quantitative Diachronic Research of South American Languages

Open Access
|Aug 2026

Abstract

This paper presents a lexical dataset for 299 South American languages, including 15 reconstructed proto-languages. The dataset comprises 25,203 lexical forms compiled from multiple publicly available sources and standardized according to the Cross-Linguistic Data Formats (CLDF) Wordlist specification. Language varieties and concepts were harmonized using Glottolog and Concepticon identifiers to improve interoperability across heterogeneous lexical resources. The dataset was designed to support quantitative diachronic research, including automatic cognate detection and phylogenetic inference. By addressing inconsistencies in transcription, naming conventions, and glossing practices, the dataset facilitates data reuse and large-scale comparative research for one of the world’s most linguistically diverse regions.

DOI: https://doi.org/10.5334/johd.555 | Journal eISSN: 2059-481X
Language: English
Page range: 104 - 104
Submitted on: Apr 9, 2026
Accepted on: Jul 1, 2026
Published on: Aug 3, 2026
Published by: Ubiquity Press
In partnership with: Paradigm Publishing Services

© 2026 Álvaro Arozarena Gómez, Annemarie Verkerk, Marc Tanti, published by Ubiquity Press
This work is licensed under the Creative Commons Attribution 4.0 License.