1 Context and motivation
Historical and etymological datasets are typically compiled for qualitative rather than for quantitative research. This focus creates specific structural problems for querying and analysis. Among these problems are data redundancy, label inconsistency, and multi-values fields (fields containing multiple values per cell). To address these issues, the data can be normalized, so that each value is stored in one place only, which can then be referred to with an identification code (id). This approach reduces errors, ensures consistency and facilitates analyses and updates. This paper describes the ongoing normalization of a historico-etymological dataset (García-Covelo, 2026b) originally designed to study the development of Obstruent + Lateral (OL) clusters (and their palatalization) in Ibero-Romance, with particular focus on Galician, Portuguese, and Spanish.
OL palatalization mainly targeted five clusters (/pl fl bl kl gl/), which can be classified as either primary or secondary. Primary clusters were etymological in origin, whereas secondary O(V)L clusters originally contained an unstressed vowel between the two consonants (i.e., C1(V)C2) that was subsequently lost through syncope, e.g., Latin clāvus ‘nail (metal)’ vs. Late Latin oclus (<ŏcŭlus) ‘eye’.1 These clusters appeared in three environments: word-initially, postconsonantally, and postvocalically.2 One example of the Romance outcome variety of OL palatalization is Latin clāvis ‘key’ > Galician []ave, Portuguese [S]ave, Spanish [L] or [J]ave, Romanian [kj] or [c]eie, and Italian [kj]ave (cf. Recasens, 2026; Zampaulo, 2019: 63–5).3
Ibero-Romance OL palatalization stands out within Romance for three reasons: it operated irregularly, even among words sharing similar phonological contexts; it did not affect all clusters equally; and its outcomes varied according to obstruent voicing and word position. Despite this complexity, previous research had relied on a relatively narrow body of evidence, frequently limited to a single language. The lack of a comprehensive historical foundation of this change hindered the study of this unusual evolution, both quantitatively and qualitatively.
To address this documentation gap, a dataset was compiled, comprising primarily Galician, Portuguese, and Spanish inherited words derived from etyma with OL clusters.4 The dataset’s central aim was to provide the historical foundation needed to shed light on the development of OL clusters and their palatalization in Ibero-Romance. More specifically, it was compiled to answer the following questions:
Is there new historical evidence of OL palatalization in words and clusters for which no evidence had previously been attested?
How regular was Ibero-Romance OL palatalization?
What sound changes may have blocked or interfered with OL palatalization?
The analysis of the resulting wordlist, comprising 659 entries, provided new and deeper insights into these questions and their answers (García-Covelo, 2025). However, its current design poses challenges for long-term sustainability and future expansion since the underlying structural issues that complicated or limited that analysis, such as multi-valued fields, remain unresolved. The following sections outline the compilation process and the design choices that informed the original dataset, as well as limitations derived from those decisions. Furthermore, they also discuss the challenges in the normalization process, from over-packed fields to column merger or split and conventions for linguistic abbreviations, especially of historical language stages and dialectal or regional varieties.
2 Dataset description
2.1 Design and compilation
This wordlist was originally designed to support the qualitative, rather than quantitative, analysis of OL palatalization in Ibero-Romance, as the outlined research questions show. Accordingly, design decisions prioritized readability and compilation practicality over structural optimization for querying. As the dataset and research questions evolved, the limitations of certain design decisions became clear; the normalization steps described in Section 3 represent an effort to bring the dataset’s structure in line with those expanded analytical goals.
The original dataset was stored in a single spreadsheet; this format proved practical during data collection because it facilitated data entry and readability. However, this design choice introduced data redundancy, label inconsistency, and multi-valued fields, which complicated the quantitative analysis of the wordlist and, ultimately, its expansion and maintenance. The normalized version addresses these issues by distributing information across multiple tables, namely main, lookup, and junction tables, where each piece of information is stored only once and reorganized to facilitate efficient querying. For instance, phonological derivations, which in the original wordlist were embedded within the dense sources_and_notes column, are intended to be accessed directly through a dedicated column in the root table (see 3.2).
The dataset describes the etymological relationship between an etymon and an attested inherited form. Both inherited words with and without palatalization were included to enable the analysis of the regularity of the change. Given the clusters with traces of palatalization, only /pl fl kl bl gl tl dl sl/5 were included, both primary and secondary; secondary clusters were only included if they contained the vowel /u/ in intertonic medial position, as this was the environment most likely to undergo syncope and produce a cluster (see García-Covelo, 2025 for further details).
As a guideline for identifying words with OL clusters, Corominas & Pascual (2012) was used. The dictionary data was extracted from Rich Text Format (RTF) files, structured into tabular format, and searched via regular expressions on the etymology column. The results were cleaned to exclude clear borrowings6 and forms without etymological value (see García-Covelo, 2026a), then compared and extended with further etymological, historical, and philological sources (e.g., González Seoane, Álvarez de la Granja and Boullón Agrelo, 2006; Machado Filho, 2013; Mariño Paz, 2017; Meyer-Lübke, 1935; Varela Barreiro, 2004).
The original dataset contains 15 columns, which describe three main elements:
Etymon
etymon: also in italics: source word undergoing the relevant sound changes, mostly Latin.
ol_cluster: cluster within the etymon, i.e., /pl fl kl bl gl tl dl sl/; clusters with long obstruents are marked with double consonants, e.g., /kkl/.
ol_type: type of cluster, which could be primary (plŭvĭa ‘rain’) or secondary (pŏpŭlus > poplus ‘people’).
ol_position: position of the cluster within the word, i.e., word-initial, postconsonantal, postvocalic.
root: root or stem of the etymon.
etymon_lang: language variety of the etymon.
Attested inherited word
inherited_form: attested inherited word form, showing the output of the original OL cluster.
other_forms: spelling variants or parallel developments of the inherited word.
lang_variety: language variety of the inherited word.
date: century of the 1st attestation of the inherited word.
Etymology
sound_change: sound changes undergone by a given cluster in the etymon.
outcome: outcome of a given OL cluster in the inherited word.
status: status of the etymology, e.g., uncertain etymology or (possible) borrowing.
palat: presence, absence or uncertainty of OL palatalization.
sources_and_notes: references, discussion, and further information.
The initial design was relatively simple, describing the etymon, its language, its clusters, the inherited word and its variety, and a notes column. Further columns were added as compilation progressed, to capture additional information about the inherited words and their etymologies. Several of these later additions, particularly other_forms, sound_change, and outcome, reflect convenience-driven decisions that created structural challenges for quantitative analysis. That said, the preservation of a single representative form per etymon was motivated by the intended statistical analysis, specifically to avoid misrepresenting the regularity of OL palatalization: since descriptive statistics do not offer mechanisms for fixed or random effects at the etymon level, retaining multiple forms per etymon would have skewed any quantification of palatalization patterns (see Section 4). Consequently, where OL palatalization is present, the palatalized form was selected as the primary entry; where it is absent, the form reflecting the oldest attested sound changes was preferred. As a result, alternative forms, included for their distinct evolutions, were retained in a dedicated other_forms column rather than as independent entries. However, this decision introduced structural ambiguity into other_forms, sound_change, and outcome. This ambiguity is addressed in the normalization steps described in Section 3.
A related convenience decision shaped sources_and_notes. This column originally contained all information pertaining to the etymon, inherited word, and etymology; however, as new layers of information were added during compilation, it became increasingly difficult to parse.
2.2 Metadata
Repository location: 10.5281/zenodo.18818049
Repository name: Zenodo
Object name: Historico-etymological dataset on the development of obstruent-lateral clusters in Ibero-Romance
Format: comma-separated values (CSV)
Version: 1.0
Size: 659 rows x 15 columns
Creation dates: Approximately between 2022-01-01 and 2026-02-28
Dataset creators: García-Covelo, Andrea
Language of the dataset: English (for the structure and discussion)
Languages included: (Late) Latin, Galician-Portuguese, Galician, Portuguese, (Old) Spanish, Asturian, (Astur-)Leonese, (Ribagorçan) Aragonese, Old French, Mozarabic, Germanic, Catalan (further languages and varieties may occasionally occur).
License: CC BY 4.0
Publication date: 2026-02-28
2.3 How to query
The original dataset (García-Covelo, 2026b) can be explored and analysed in R; an example script is found in the Zenodo repository (see example_script.R). This script can serve as a provisional reference while normalization is ongoing. A querying script for the normalized dataset will be made available, once the normalization process is complete, in a dedicated GitHub repository,7 where the process is currently being documented.
3 Normalization methods
This section describes the reorganization of the single-table wordlist into multiple tables, and the challenges arising from this process. After the analysis in García-Covelo (2025), it was decided that the wordlist would benefit from restructuring, both to enable more in-depth quantitative analysis and to improve sustainability and expansion. The tabular format was retained but the information was distributed across several tables optimized for distinct structural functions. The different table types used are defined as follows:
Main tables store the core entities of the dataset. Each row represents a distinct object of study, in this case, an etymon, an inherited word, or an etymological relationship.
Lookup tables store controlled lists of values and their descriptions, which are referenced in the main tables via identification codes. This eliminates redundancy by ensuring each piece of information is stored only once. Lookup tables are used here for categorical fields such as language variety or cluster type.
Junction tables link two or more tables in a many-to-many relationship, where a single object in one table may be associated with multiple objects in another and vice versa. Junction tables are used here for the bibliography, linking bibliographic references to the main tables.
As its structure dictates, the dataset should be reorganized into three main tables: one for etyma, one for inherited words, and one for etymologies. The etymology table functions as a junction table, capturing the many-to-many relationship between etyma and inherited words: a single etymon can give rise to several inherited words, and a single inherited word may derive from more than one etymon. The main tables are supported by several lookup tables for categorical fields and a junction table for bibliographic references. The lookup and junction tables are discussed first, as the main tables depend on them.
The normalization process is carried out in Python (Python Core Team, 2022), using the pandas library for data manipulation (McKinney, 2010). The ongoing process can be followed on GitHub (see footnote 7), where the tables in their current state of reorganization are progressively made available. It should be noted that not all tables are yet available, and that some of the decisions outlined in this paper may be subject to revision as the normalization progresses.
3.1 Look-up tables
3.1.1 Languages
The original wordlist encodes language variety as a free-text abbreviation, which means that the same variety can appear in different forms. Misspellings or trailing spaces were the main cause of contradictions during data analysis and the abbreviations had to be manually changed. A lookup table with the linguistic information solves this issue. Both etymon_lang and lang_variety contained linguistic information and were consequently merged.
For the linguistic abbreviations, the conventions in Corominas & Pascual (2012) and Meyer-Lübke (1935) were originally used. Given that this is a historical dataset targeting both linguistic and computational audiences, International Organization for Standardization (ISO) 6398 codes were also considered and are included alongside the traditional abbreviations where available (see Table 1).
Table 1
Example of abbreviation conventions for linguistic varieties.
| LANG_NAME | LANG_CODE | ISO_639_2_3 |
|---|---|---|
| Asturian | Ast. | ast |
| Galician | Gal. | glg |
| Galician-Portuguese | GP | |
| Portuguese | Pt. | por |
A further decision concerns which linguistic varieties receive their own entry in the table. Several regional varieties — on the one side, Alavese, Andalusian Spanish, and the regional varieties of Burgos, Cantabria, Navarra, Salamanca, and Santander, and on the other, Trás-os-Montes and Minhoto — appear only once or twice in the wordlist and have been consolidated under the umbrella terms Spanish and Portuguese respectively. Their geographic information is preserved at the level of the inherited words table (see 3.3.2) rather than the languages table. Other varieties, specifically Asturian, Leonese, and Ribagorçan Aragonese, are retained as distinct entries because their OL palatalization outcomes may diverge from the standard developments and are therefore of independent analytical interest. Section 4 discusses how representing geolinguistic information as both defined dialectal varieties and geographic information impacts future dialectometric analysis of the data.
3.1.2 Roots
Latin vowel length is phonemically distinctive: two etyma that differ only in vowel quantity are semantically independent words, e.g., pŏpŭlus (short /o/) ‘people’ vs. pōpŭlus (long /o/) ‘poplar-tree’. Diacritics marking vowel length are therefore sometimes necessary to distinguish etyma, but they are not always required to determine word stress, and different reference works are inconsistent in their use.9 Consequently, the same etymon may appear with or without diacritics depending on the source, producing inconsistent representations of the same etymon, which had to be manually removed. Converting the etymon column to a lookup table reduces this risk since the number of unique etyma can be defined explicitly.
In addition to a unique id (root_id) and the root form itself (root_form), the lookup table is planned to include a language field (lang_id) and a phonological derivation column (phonological_derivation) (see 3.2). Assigning values to lang_id at the root level, however, presents a difficulty: in the original dataset, language information was assigned to etyma rather than roots. This means that the root itself carries no language label when etyma from different language stages (e.g., Latin and Late Latin) share the same root, which creates ambiguity that cannot be resolved automatically.
3.1.3 OL categorization
The columns describing OL clusters have very restricted values. Nevertheless, without explicit encoding, these values are constantly repeated along the dataset, increasing the risk of misspellings and errors. By turning these columns into lookup tables, this issue is effectively solved. OL clusters are described with four categories, which affect the presence or absence of palatalization and the specific outcomes: ol_cluster_type (primary OL or secondary O(V)L), ol_cluster (/pl fl kl bl gl tl dl sl/), ol_length (long or short obstruent) and ol_pos (word-initial, postconsonantal, and postvocalic).
This categorization upgrades the original wordlist, where the information for ol_cluster and ol_length was condensed in the column ol_cluster and length was indicated with double obstruents, e.g. /ffl/. Splitting this column improves querying OL clusters by their attributes.
3.1.4 Etymological status
Like the columns described in Section 3.1.3, the status column was straightforwardly converted into a lookup table. While label inconsistencies in the original data were manually resolved prior to the analysis in García-Covelo (2025), the lookup table ensures that such inconsistencies cannot enter the dataset in future expansions.
3.1.5 Parallel evolutions and spelling variants
The other_forms column conflates three structurally distinct phenomena: spelling variants, parallel developments, and conservative evolutions or possible borrowings (see 2.1). This made the column difficult to systematically parse and analyse, as illustrated in Table 2.
Table 2
Types of alternative forms in the other_forms column.
| LATIN | GALICIAN-PORTUGUESE (MAIN) | GALICIAN-PORTUGUESE (OTHER) | TYPE |
|---|---|---|---|
| gĕnŭcŭlum ‘(little) knee’ | gẽollos | geollos, geonllos | Spelling variants |
| artĭcŭlus ‘joint’ | artillos | artigoos | Parallel developments |
| macŭla ‘stain’ | malha | mãchas, mágoa, mangra | Parallel developments |
| plantāre ‘to plant’ | chantou | prantar | Conservative evolution or possible borrowing |
To address this, the column should be dissolved: spelling variants should be moved to a dedicated table linked to the inherited word table via word ids, while parallel developments and conservative evolutions should be promoted to independent entries where appropriate. These cases and their theoretical implications are further discussed in Section 4.
Since the values in outcome refer primarily to the main inherited word, but were sometimes extended to cover alternative forms as well, these columns inherit the structural ambiguity of other_forms: dissolving that column resolves the issue. Subsequently, outcome should also be converted into a lookup table.
3.1.6 Sound changes
The sound_change and palat columns are both converted into lookup tables. These columns overlap, as the information captured in palat is also contained within sound_change: palatalization and uncertain palatalization are recorded there alongside other sound changes, including lenition, rhotacism, lack of syncope, and metathesis. The dedicated palat column was retained because it simplifies querying for this specific sound change, accepting only three values: yes, no, and uncertain. Whether this column is preserved alongside sound_change remains under consideration.
Similar to the outcome column, the structural ambiguity introduced by other_forms is partially resolved by its dissolution (see 3.1.5). However, OL clusters may have undergone multiple sound changes (especially in the absence of palatalization), identifiable through their outcomes in the inherited words. As a result, a junction table linking inherited words to the sound changes lookup table may be necessary.
3.2 Sources and notes
The reorganization of sources_and_notes represents a particularly significant undertaking due to its informational density: during compilation, all information was entered into a single cell, which kept the data entry process straightforward (see 2.1) but complicated parsing and interpretation. As a result, most of the information must be manually reviewed and redistributed into the appropriate tables and columns. Nevertheless, this redistribution will substantially improve the limitations discussed above. As the following examples illustrate, this column not only contains bibliographic references, but also notes on meaning and spelling, etymological discussion, and further attestations and word variants (see García-Covelo, 2026b; the searched inherited word and its variety are given at the end of each quote):
(1) CRC (290); GCLI (235); Torreblanca (1990: 322); REW (1155); DEEH (510); DCECH (blasfemar/lastimar) set the first attestation already in the 11th century (cf. CORDE). blasphemare > blastemare through dissimilation but it is unclear whether the dissimilation happened in Greek or in Latin (cf. DCECH and REW). Attempted root or stem reconstruction for words like Lat. blasphēmātĭo, blasphēmĭa, or blasphēmus (see LS for words containing the root -*blasph*-): *bllas-, *llas-, jas-*, *chas- (CORDE). DRAE (blasmar) argued that Sp. *blasmar* is a French borrowing (against this etymology, TDHLE: blasmar). (Searched term: OSp. lastima)
The first entry (1) contains several distinct types of information: plain bibliographic citations (the opening references), attestation data (“DCECH […] in the 11th century (cf. CORDE)”), evolution of the etymon (“blasphemare […] (cf. DCECH and REW)”), an attempted phonological derivation of the root undertaken to search for evidence of OL palatalization in corpora (see García-Covelo, 2025), and a disputed etymology (“DRAE (blasmar) […] TDHLE: blasmar)”).
(2) Georges and LS (carbunculus); DCECH (carbunco); REW (1677). With the meaning of ‘reddish precious stone’, there is one attestation of GP carbũche in the Miragres de Santiago (14th-15th centuries) (see DDD and TMILG: MS). With the meaning ‘tumor, abscess, boil’, see DCECH (carbunco), REW (1677), DELG (carbunco), and DXL (carbuncho). Gal. carafuncho may be a mixture of Lat. carbunculus and Lat. fūrŭncŭlus (cf. DELG: carafuncho; DCECH: carbunco/hurto; REW: 3607; and TLPGP: carafuncho). (Searched term: GP carbũche)
Similarly, the second entry (2) combines bibliographic references with the discussion of an additional inherited word (Gal. carafuncho). In addition, the attestation and etymology of the main inherited word are treated separately according to its two distinct meanings: ‘reddish precious stone’ (with OL palatalization) and ‘tumor, abscess, boil’ (without OL palatalization).
Together, these examples illustrate the following types of information: bibliographic references, attestation data, evolutionary pathways, phonological derivation of the etymon root, additional inherited words, and meaning. Several of these can be redistributed into columns in existing tables, e.g., meaning and attestation information can be incorporated into the inherited words table. Similarly, evolutionary pathways can be included in the etymologies table, either within the discussion or notes column, or as an independent column. Phonological derivations should be linked to the root table since it is the root rather than the etyma that were derived (cf. García-Covelo, 2025; see 3.1.2); whether a single column suffices or a separate table is needed remains under consideration. Additional inherited forms that represent independent developments, rather than misspellings, should receive independent entries in the main tables (cf. 3.1.5).
Bibliographic references present a further structural challenge: a given source may discuss one or more etymologies, etyma, or inherited words, and conversely, any entry may be supported by multiple sources. This many-to-many relationship requires a dedicated junction table for each of the three main tables. The bibliography table stores the metadata for each source: type, author, title, pages, URL, DOI, etc. All three junction tables should share the same structure, differing only in the main table they refer to (see Table 3). In this way, a bridge is established between the bibliographic references and the specific etymologies, etyma, or inherited words they relate to.
3.3 Main tables
3.3.1 Etyma
The etyma table contains all information concerning the source words, including their language variety, root and meaning, and the characterization of the OL cluster within each word according to four parameters: type, specific cluster, obstruent length, and word position (see 3.1.3). Accordingly, the table currently comprises the following columns: etymon_id, etymon, ol_cluster_id, ol_type_id, ol_length_id, ol_pos_id, root_id, lang_id, meaning, and notes.
3.3.2 Inherited words
Information pertaining to the attested inherited words, such as meaning, language variety, and first attestation, is contained in the inherited words table. Its current structure includes the following columns: inherited_word_id, inherited_word, meaning, lang_id, first_attestation, is_regional, variety_or_region, and notes.
In the original dataset, regional or geographical varieties were indicated directly in lang_variety via a specific abbreviation. However, as no general conventions exist for such abbreviations, most of them have been recategorized as either Portuguese or Spanish (see 3.1.1). Whether a given word belongs to a regional or geographical variety is now captured separately in is_regional and variety_or_region.
3.3.3 Etymologies
The etymological relationship between a given etymon and its reflex constitutes the central link in the dataset. This table is accordingly the most information-dense, bringing together data on the etyma and inherited words themselves, the sound changes that occurred between them, the outcome of the original OL cluster, the status of the etymology, and any accompanying discussion. The table is planned to include the following columns: etymology_id, etymon_id, inherited_word_id, outcome_ol_cluster, ol_pal_id, sound_change_id, status_id, and discussion.
4 Results and discussion
The normalization described in the preceding section is an ongoing process. The restructured dataset is progressively made available on GitHub so that readers can follow its development and, in time, draw on its outputs directly (see Section 3). The following discussion addresses the theoretical questions the restructuring brings into focus, which stand independently of the work’s ongoing state.
The classification of linguistic varieties involves a tension that needs to be acknowledged. The decision to retain Asturian, Leonese, and Ribagorçan Aragonese as distinct entries is principally motivated by their potentially divergent phonological outcomes, but also by the difficulty of delimiting regional and dialectal varieties. Linguistic varieties rarely have clear-cut boundaries: dialects typically exist along a continuum, and the line between a regional variety, a dialect, and a language is often as much a matter of convention or political history as linguistics. That said, some dialectal areas are sufficiently well-defined to warrant distinct treatment; for instance, classifying Astur-Leonese varieties under Spanish, when they represent historically distinct Ibero-Romance varieties (cf. Hammarström et al., 2025), would risk misrepresenting their status and obscuring potentially divergent developments.
The decision taken here is not necessarily the only defensible one, and users working on dialectal or phylogenetic questions should be aware of its implications. In particular, it creates a circularity worth keeping in mind: retaining a variety as a distinct category partly because of its phonological outcomes, and then using the same dataset to study those outcomes, risks building the conclusion into the classification. The present normalization does not solve this limitation, but it does offer greater transparency in the form of a controlled lookup table where the reasoning behind each linguistic entry is made explicit.
This difficulty in establishing clear, theory-neutral categories may not be unique to the present dataset. To this author’s knowledge, no standardized system exists in Romance or historical linguistics for representing regional varieties and dialects, and abbreviation conventions may vary considerably across reference works and research traditions (cf. Corominas & Pascual, 2012; Mariño Paz, 2017; Meyer-Lübke, 1935; Recasens, 2020; Weiss, 2009; Zampaulo, 2019). This motivated the decision to deprecate several of the regional abbreviations previously used: without a shared standard, retaining low-frequency entries of uncertain status compromises the interpretability of the dataset. ISO 639, a commonly used classification system in computational linguistics, was considered as a reference, but it does not cover several varieties that are central to this historical wordlist (Galician-Portuguese and Late Latin being the most notable gaps). This gap is widened by the fact that historical and computational linguistics have largely developed independent conventions (with different focuses) for representing language varieties. For a dataset that aims to be useful to both communities, this is a genuine constraint, and the dual-convention approach adopted here (retaining traditional abbreviations alongside ISO codes where available) is a pragmatic rather than a theoretically grounded solution.
Glottolog (Hammarström et al., 2025) and the Cross-Linguistic Data Formats (CLDF) initiative (Forkel & List, 2020) offer a more comprehensive reference point, covering a wider range of historical varieties (like Late Latin) and providing stable identifiers that could help bridge the two traditions. Integrating these standards is a promising direction for future work. However, they do not resolve the underlying difficulty: Glottolog operates at the level of genealogically defined varieties and, accordingly, varieties defined by geographic attestation rather than genealogical distinctness fall outside the scope of this system. Classifying such forms as dialects of a given language would presuppose boundaries that may be historically or linguistically contested.
A related set of questions concerns lexical identity across parallel evolutions. Reorganizing other_forms raises a deeper question about when two reflexes of the same etymon constitute distinct inherited words rather than variants of one. The dataset’s use of OL palatalization as the primary selection principle reflects a chronological assumption, i.e., palatalized forms represent the earliest layer of evolution. However, this boundary is not always clear. Non-palatalized forms such as artigoos and mágoa (see Table 2) may be native reflexes, exhibiting lack of vowel syncope and regular intervocalic /l/ loss in Galician-Portuguese; therefore, treating them as secondary variants may be historically inaccurate. Forms like prantar present a different problem: whether they constitute conservative evolutions or early borrowings reflects how the dataset draws the boundary between inheritance and contact (see footnote 6).
More generally, analysing OL palatalization per surface form rather than per etymon risks skewing measures of its regularity. If any inherited form of an etymon underwent the change, palatalization must be considered to have affected that etymon regardless of whether parallel forms also palatalized. Descriptive statistics cannot capture this distinction because they do not offer mechanisms for fixed or random effects at the etymon level. Therefore, the restructuring described in this paper is not merely a data cleaning step, but a prerequisite for more comprehensive statistical modelling. This reflects a broader challenge in humanities data work: the structure of a dataset is not independent of the research questions it was designed to answer, and may need to be revisited as those questions evolve.
5 Implications/Applications
This dataset provides the first comprehensive cross-linguistic compilation of historical evidence to study the development of OL clusters in Ibero-Romance. While its original qualitative focus limited its quantitative potential, its maintenance and expansion, these limitations are directly addressed by the present normalization process.
The restructured wordlist should be easier to query and parse: over-packed fields are intended to be dissolved, and information is being distributed across clearly defined main, lookup, and junction tables. Lookup tables are designed to eliminate redundancy and enforce label consistency, making the dataset easier to maintain as it grows and easier to extend without introducing the kind of inconsistency that complicated earlier analyses. Together, these improvements are a prerequisite for the analytical gains described below.
The normalization also opens up concrete possibilities for expansion. The dissolution of other_forms and sources_and_notes alone is expected to add entries that were previously embedded in multi-valued fields. Beyond that, the dataset can be enriched with further words from Galician, Portuguese, and Spanish; regional variety coverage can be deepened, with dialect groups such as Aragonese and Astur-Leonese represented more comprehensively to give a fuller Ibero-Romance picture; and in principle, other Romance varieties could be incorporated, extending the dataset’s scope further still. Another direction for expansion involves moving from types to tokens, i.e., from unique inherited words with only their first attestation documented, to multiple attestations with associated frequency data. This would add a valuable chronological layer to the study of OL palatalization, illustrating not only how early a word palatalized, but also how frequent that word was at the time of the attestation. Such expansions would naturally broaden the dataset’s reuse potential, supporting a wider range of research questions.
The dataset was conceived to address three questions about the development of OL clusters: whether there is new historical evidence of OL palatalization; how regular that palatalization was across Ibero-Romance; and what sound changes may have interfered with it. The normalization improves the dataset’s ability to answer all three, most directly by enabling more robust quantitative analyses of the irregularity of OL palatalization and the outcome variety of OL clusters.
Reuse potential beyond these original questions is real, while bounded by the dataset’s specific focus and the relatively small pool of words it covers. The dissolution of other_forms and the outcome column allows the study of sound changes coexisting with OL palatalization, as well as the analysis of parallel inherited forms and spelling variants. Previously, these questions were structurally difficult or impossible to pursue. Beyond these specific questions, the structured representation of sound change data may also be of value for computational phonological work in Ibero-Romance.
More fundamentally, the process documented here illustrates how a dataset originally designed for qualitative analysis can be progressively restructured to support quantitative modelling without abandoning its original descriptive depth. This kind of development is likely familiar to many researchers working with historical linguistic data, and making the process explicit, including its evolving priorities and direction shifts, contributes to emerging best practices for corpus design in the field.
Notes
[1] Unless otherwise specified, all Latin meanings and words were taken from Lewis and Short (1879).
[2] Postvocalically, OL clusters could have a short (dēclārāre ‘to make clear’) or long obstruent (acclāmāre ‘to shout at’).
[3] These realizations are illustrative: they vary across speakers and speech situations, and their transcription may vary across sources.
[4] Throughout this paper, etymon (plural etyma) is used to refer to the source forms in the dataset (from which the Ibero-Romance words derive), whether Latin, Arabic, or Germanic, and whether attested or reconstructed. For the Romance forms derived through oral transmission, inherited word is used (Spanish palabras patrimoniales, German Erbwörter; reflex is a common alternative). These inherited words are attested historical forms that are not necessarily preserved in the modern language, e.g. Old Spanish llor (from Latin flōs, -ris ‘flower’) but Spanish flor (García-Covelo, 2026b).
[5] While OL palatalization primarily affected /pl fl kl bl gl/, /tl dl sl/ occasionally underwent the change, though not necessarily directly, e.g. Latin vĕtŭlus ‘old’ > veclus > Galician vello (García-Covelo, 2025: 39-40).
[6] In this step, only unambiguous borrowings were excluded. Borrowings are words introduced through written rather than oral transmission that did not undergo regular sound changes; for this dataset, words first attested after the 15th century and lacking these changes were classified as such, e.g., Spanish declarar and claudicar (from Latin claudĭcāre ‘to limp’, dēclārāre ‘to make clear’) show no intervocalic stop voicing, marking them as learned forms, and tomate (from Nahuatl tomatl ‘tomato’) entered Spanish after OL palatalization had concluded (cf. Corominas & Pascual, 2012). Inter-Romance loans were excluded unless they show evidence of early changes like lenition (García-Covelo, 2025). Words whose borrowed status required further investigation were not discarded at this stage.
[7] GitHub repository for the normalization of the OL clusters dataset: https://github.com/agcovelo/normalization_corpus_olclusters.
[8] View ISO 639 language code page: https://www.iso.org/iso-639-language-code.
Acknowledgements
The author would like to thank the three anonymous reviewers for their insightful comments and suggestions, particularly regarding the discussion and broader implications of the findings, which greatly improved this paper. The author is also grateful to Christoph Draxler for his valuable advice on the restructuring and normalization of the dataset.
AI Declaration
Generative AI tools (Claude Sonnet and Opus 4.6) were used for manuscript refinement, for instance to improve writing clarity and style, and for writing assistance. AI assistance was also used to develop code for the normalization of the dataset. All intellectual judgements and revisions remain the author’s own.
Author Contributions
Andrea García-Covelo (ORCID: 0009-0001-6677-2463): Conceptualization; Methodology; Data curation; Investigation; Formal analysis; Software; Visualization; Writing - original draft; Writing - review & editing.
Author Note
The author was affiliated with the LMU Munich at the time this research was conducted, but is no longer based there.
