(1) Overview
Repository location
Context
The Anatolian languages form the earliest attested branch of the Indo-European family and therefore occupy an important place in historical and comparative linguistics (Kloekhorst, 2008). Within this branch, emotional vocabulary is a useful domain for historical-semantic study: it demonstrates lexical derivation, semantic development from concrete to abstract meanings, and preserves evidence for affect-related concepts. Emotional vocabulary for a set of closely related languages helps reconstruct affect-related concepts (for terminology, see Jackson et al., 2019) in a proto-language, in this case, Proto-Anatolian (PA) and Proto-Indo-European (PIE).
At the same time, this material is difficult to study: relevant lexemes are dispersed across several corpora of Anatolian languages, etymological dictionaries, critical text editions, and specialized publications (see especially Sonik & Steinert, 2023; Galhano, 2023), and some items are not included in any single corpus or dictionary. As a result, comparison often has to be reconstructed case by case from heterogeneous sources. No single resource has so far brought together the emotional lexicon of the Anatolian branch in a compact comparative form organized around etymological relations.
The dataset presented here was produced within the framework of the AHEC project (Akkadian and Hittite Emotions in Context). Its compilation draws on two major digital resources for Anatolian philology: eDiAna (Miller et al., 2017) and HPM (Hethitologie-Portal Mainz, Müller & Schwemer, 2018). These resources are essential for access to primary textual and lexicographic material, but they do not in themselves provide a way to search systematically for affect-related concepts. This dataset addresses the gap by assembling emotional lexemes from the Anatolian languages into a single structured resource organized comparatively and etymologically. In this form, it is intended for independent use in historical semantics, emotion lexicon research, Anatolian philology, and Indo-European comparative linguistics.
The references cited in this paper document the context for the dataset and are used for citation within the data paper itself. The dataset deposit on Zenodo does not carry an independent reference list.
(2) Method
Steps
The dataset was compiled through manual philological and comparative analysis, using the eDiAna corpus for the non-Hittite Anatolian languages, the TLHdig (Thesaurus Linguarum Hethaeorum digitalis) corpus of transliterations at the Hethitologie-Portal Mainz (HPM, 2001) for Hittite, and critical text editions and etymological dictionaries as principal sources of lexemes with emotional semantics (Kloekhorst, 2008; Melchert, 1993, 2004; Miller et al., 2017; Müller & Schwemer, 2018; Puhvel, 1984–2011; Thesaurus Linguarum Hethaeorum digitalis; Tischler, 1983). Compilation proceeded in six stages:
identifying relevant lexemes in the sources (dictionaries, texts, lexical corpora, specialized papers),
grouping forms attested in different languages of the Anatolian branch and their cognates in core Indo-European languages into cognate sets,
defining their common Proto-Anatolian and/or Proto-Indo-European root,
reconstructing semantic derivation from concrete to abstract meaning,
placing the cognate set into a certain emotional domain,
entering the results into a structured comparative table.
Candidate lexemes were selected on semantic grounds. Forms were considered for inclusion when their lexical meaning in context justified their treatment as part of the emotional lexicon. Attestations were checked in the source corpora and against available lexicographic documentation. Cases in which emotional relevance was uncertain, highly context-dependent, or philologically insecure were treated conservatively.
The selected forms were then organized into entries: attested lexemes were grouped into cognate sets where the etymological evidence supported such alignment, and each entry was indexed by lemma. Every entry was linked to a PA and/or PIE reconstruction. Only attestations that can be reconstructed on either the PA or PIE level were included in the dataset.
The final dataset was organized as a structured table, with one row per cognate set and dedicated columns for reconstructions and language-specific material. Editorial normalization was applied in order to maintain consistency of lemma assignment and comparative alignment. Empty cells indicate lack of attested material within the scope of the dataset, not proof of lexical absence in the language concerned.
Sampling strategy
The dataset is a criterion-based scholarly selection. Inclusion depended on two factors: relevance to the emotional lexicon and the availability of sufficient etymological evidence to support reconstruction on PA or PIE level. The resource is therefore deliberately bounded by PA/PIE reconstruction, so it is not an exhaustive inventory of Anatolian emotional vocabulary.
Quality control
Quality control consisted of repeated manual checking of forms against the source corpora and lexicographic resources, together with iterative review of cognate grouping and various approaches to PA/PIE reconstruction. Derivational nests and lemma assignments were checked for internal consistency during compilation, and doubtful cases were handled conservatively. Both authors carried out the checking process cooperatively throughout compilation. Approximately one hundred candidate roots were excluded on the grounds of philological uncertainty or because they could not be reconstructed at the PA or PIE level; the precise number is not reported here, as establishing it would require a systematic review of excluded material that lies outside the scope of this deposit.
(3) Dataset Description
Repository name
The latest version of the dataset (v2.1.0) is deposited in Zenodo (Uvarova & Molina, 2026) and is available under the stable identifier 10.5281/zenodo.19477826.
Object name
The deposited object consists of a main tabular dataset containing the emotional lexicon of the Anatolian languages, its semantic derivation and reconstructions, Anatolian Emotions.csv. The deposit also includes a separate PDF file of etymological commentaries, “Comments for the lexicon of Anatolian emotions.pdf”.
The PDF commentary covers approximately three-quarters of the 72 entries; commentary is provided for those entries where the treatment in available etymological dictionaries and specialized literature required clarification or supplementation. Entries for which the existing lexicographic record is unambiguous do not have dedicated commentary. Navigation between the CSV and the PDF is by lemma: users locate the relevant lemma in the CSV and search for the corresponding lemma in the PDF file. Automated cross-referencing between the two files is planned for a future version of the dataset.
Detailed data structure
The non-lexical columns of the table include the following information:
lemma;
semantic derivation: a reconstruction of the semantic development from the Proto-Indo-European or Proto-Anatolian root to the attested emotional meaning;
domain: the assigned emotional macrodomain;
semantics: a concise keyword for the more specific emotional concept;
PA_root: the reconstructed Proto-Anatolian root;
PIE: the reconstructed Proto-Indo-European root, where identifiable;
IE_cognates: non-Anatolian Indo-European cognates available;
translation: the English rendering of the reconstructed basic emotional meaning.
Language-specific material is recorded further in dedicated columns:
luw: Luwian
hitt: Hittite
pal: Palaic
lyd: Lydian
lyc_A: Lycian A
lyc_B: Lycian B
car: Carian
Each populated language column contains the derivational nest attested for that root in the language concerned. Empty cells in a language column indicate that no material for the relevant cognate set has been found for that language in the dataset. They do not constitute proof that the root or lexeme was historically absent from that language. More detailed etymological discussion is provided in the accompanying PDF commentary file.
Dataset statistics
The dataset contains 72 etymological entries covering 150 language-specific derivational nests. Attested entries by language are as follows: Hittite (65), Luwian (44), Palaic (12), Lycian A (11), Lycian B (7), Lydian (7), and Carian (4).
Format names and versions
The main dataset is distributed in CSV format. The etymological commentary is provided as a PDF file. Version 2.1.0 was published on 9 April 2026.
Creation dates
The deposited version was finalized on 9 April 2026.
Dataset creators
The dataset was compiled by Maria Molina (principal investigator, supervisor, and data compiler, Tel Aviv University) and Ksenia Uvarova (lexical compiler and semantic annotator, Russian State University for the Humanities) within the framework of the DFG-funded joint German-Israeli project AHEC (Akkadian and Hittite Emotions in Context; https://hittite-emotions.net).
Language
The documentation and metadata are in English. The lexical material covers seven Anatolian languages: Hittite, Luwian, Lycian A, Lycian B, Lydian, Palaic, and Carian. Lexical forms are given in scholarly transliteration (broad transcription for cuneiform data).
License
The dataset is published under the Creative Commons Attribution 4.0 International (CC BY 4.0) license.
Publication date
2026-04-09
(4) Reuse Potential
Philological and comparative reuse
The dataset’s primary reuse value lies in Anatolian philology and Indo-European comparative linguistics. By assembling emotional lexemes from seven Anatolian languages in a single comparative table linked to Proto-Anatolian and Proto-Indo-European reconstructions, it makes available material that would otherwise have to be assembled from multiple corpora, dictionaries, editions, and specialized papers (for example, from Galhano, 2023 and various etymological resources cited above). It can be used to trace semantic developments from reconstructed to attested reflexes, to compare Anatolian emotional vocabulary with cognate material from core Indo-European languages, and to reconstruct PIE emotional lexicon. The combination of lexemes available for comparison within a single table organized for emotional semantics is what makes future work on PIE and PA emotional lexicon significantly more manageable.
Semantic and typological reuse
Beyond Anatolian studies, the dataset can support work in historical semantics, emotion lexicon research, and lexical typology. Its structure makes it possible to examine patterns of semantic shift and the lexical encoding of affect (cf. Jackson et al., 2019), and to compare these patterns with analogous datasets from other languages of the world. The Jackson et al. (2019) cross-linguistic dataset of emotion semantics, which establishes both universal structure and cultural variation in emotional vocabulary across a large typological sample, provides an immediate point of comparison: the present dataset can be used to position the Anatolian evidence within that broader cross-linguistic picture and to contribute diachronic depth to typological generalizations about emotion lexicon.
Digital and educational reuse
The machine-readable CSV format makes the dataset suitable for integration into broader lexical and philological databases. Concretely, the tabular structure is compatible with the Cross-Linguistic Data Formats (CLDF) standard (Forkel et al., 2018), which provides a widely used framework for cross-linguistic datasets and is the basis for repositories such as Lexibank; conversion to CLDF would make the dataset directly interoperable with existing cross-linguistic lexical resources. The dataset could also be linked to existing etymological and philological databases that cover Indo-European and Anatolian material. Beyond computational reuse, the dataset is suitable for teaching in historical linguistics, Indo-European linguistics, Anatolian philology, and digital philology.
Data Accessibility Statement
The dataset described in this paper is openly available on Zenodo at https://doi.org/10.5281/zenodo.19477826. The Zenodo record includes the main CSV dataset and the accompanying PDF file, Comments for the lexicon of Anatolian emotions.pdf. The Zenodo deposit (v2.1.0) is the stable archival record and should be cited for reproducibility. The most recent revisions of the dataset, together with additional explanations and supplementary documentation, are available at the GitHub repository Etymology of Anatolian emotional lexicon (Molina & Uvarova, 2026); note that this repository is a living working space and may differ from the deposited version.
Supplementary Files
A related GitHub repository, “Etymology of Anatolian emotional lexicon” (Molina & Uvarova, 2026, https://github.com/mashenkeisraeli/hittite_emotions/tree/etymology) provides the most recent revisions of the dataset together with additional explanations and supplementary documentation. Users are advised that the GitHub repository is a living working space and may diverge from the deposited version over time. For citation and reproducibility purposes, the Zenodo deposit (v2.1.0) is the stable archival record and should be cited in preference to the GitHub repository.
Acknowledgements
The AHEC project rests on the sustained work of Prof. Doris Prechel, Dr. Ulrike Steinert (Johannes Gutenberg-Universität Mainz), and Prof. Amir Gilan (Tel Aviv University). We thank Dr. Ilya Yakubovich (Philipps-Universität Marburg), Prof. David Sasseville (École Pratique des Hautes Études), and Dr. Vladimir Shelestin (Institute of Oriental Studies, Russian Academy of Sciences) for consultations on specific etymological questions. Access to eDiAna and the Hethitologie Portal Mainz (HPM) is gratefully acknowledged.
Author Contributions
Maria Molina was responsible for conceptualization, methodology, formal analysis, project administration, resources, validation, visualization, data curation, investigation, and writing. Ksenia Uvarova contributed to data curation and investigation.
