1 Overview
Repository location
Context
This dataset was originally developed and annotated for the quantitative analysis presented in Ürker and Boleda (2026). The primary objective of that study was to systematically examine the prevalence and distribution of various types of lexical semantic change across languages.
For the analysis reported in Ürker and Boleda (2026), only completed semantic changes from DatSemShift were annotated and analyzed. The dataset presented in this paper extends that work by also including cases of synchronic polysemy documented in the database. Furthermore, while the analysis in Ürker and Boleda (2026) was conducted at the pattern level, the present dataset provides annotations at the level of individual realizations.
Overall, this dataset provides cross-linguistic examples of widely recognized types of semantic change, covering both completed semantic changes and cases of synchronic polysemy. Furthermore, we provide annotation guidelines for classifying different types of semantic change, developed on the basis of widely accepted semantic change typologies (Geeraerts, 2020; Traugott, 2000, 2017). These guidelines can facilitate future efforts to extend this dataset and support the creation of new semantic change datasets.
2 Method
Steps
The initial step of our dataset construction was to extract data from the original database. The Database of Semantic Shifts in Languages of the World (DatSemShift) (Zalizniak et al., 2000–2026) contains records of semantic change across 2,248 languages from 20 language families. DatSemShift groups realizations of the same semantic shift patterns across different languages (e.g. the change from suffering to work is attested both in French and Irish, so these changes are grouped together) and for each realization within a semantic shift pattern, it specifies the source and target meanings using an arrow (when the direction is known; otherwise, it leaves it unspecified). It also indicates how each change has taken place, such as through polysemy, derivation, or semantic evolution. However, the database does not include information on what type of change has taken place, e.g., metaphor, broadening etc.
We extracted all available data from DatSemShift1 in October 2025 and filtered the realizations labeled as semantic evolution or microevolution. Changes categorized as semantic evolution in the database represent shifts from a parent language to a daughter language (e.g., from Latin to Italian), whereas microevolution refers to completed semantic changes that occur within the same language (Zalizniak, 2025). We initially focused on these categories because the analysis of Ürker and Boleda (2026) was restricted to completed semantic changes. We then extended the dataset by incorporating polysemy cases, by adding polysemy realizations to the filtered data to enable comparisons between completed semantic changes and cases of synchronic polysemy.
After filtering the relevant data, we addressed directionality issues in the extracted dataset. We normalized the directional information: in the database, some shifts are expressed as source → target and others as target ← source, and we normalized it as source → target. In some instances, the direction of change—as well as the source and target meanings—were not explicitly specified. However, it was often possible to infer a direction (e.g., a shift involving Ancient and Modern Greek, where the only plausible direction is from Ancient to Modern Greek). In these cases, we made the direction information explicit.
After these steps, we obtained 2,782 individual realizations from 792 languages, representing 234 distinct semantic change patterns.
Annotation Process
The annotation was conducted independently by three annotators (the two authors of this paper and a colleague with training in linguistics) using an annotation guideline developed on the basis of the theoretical literature on semantic change types. Rather than annotating each realization separately, the annotators annotated the underlying semantic change pattern. During the process, annotators tagged the source and target meanings with their corresponding PoS and annotated each semantic change pattern according to an annotation guideline2 based on widely accepted types of semantic change in the literature. They also provided brief rationales for their annotations and assigned confidence scores to increase transparency and enhance the reusability of the data. The two authors of this paper met to build a consensus version of the annotation.
We adopted the following typology of semantic change in our dataset (Geeraerts, 2020; Traugott, 2000, 2017).
Changes in Extension. This category includes types of semantic change in which the size of a word’s extension is altered, without changing the broad ontological category that is denoted by it. The extension of a word refers to the set of entities to which the word can be correctly applied, e.g., the extension of flower includes all the flowers in the world (Crystal & Alan, 2023). Subcategories:
Broadening: An increase in extension. moon: satellite of Earth → satellite of any planet (Geeraerts, 2020)
Narrowing: A decrease in extension. deer: any wild animal → deer (Traugott, 2017)
Changes in Emotive Value. This category includes changes in the affective associations of a word (i.e., shifts toward more positive or negative connotations). Subcategories:
Amelioration: A shift toward more positive connotations. nice: ignorant → pleasing (Traugott, 2000)
Pejoration: A shift toward more negative connotations. stink: to smell → to smell obnoxious (Traugott, 2000)
Changes in the Semantic Frame. This category includes changes in which the target meaning is related to the source meaning through the following semantic relationships (Barcelona, 2019; Traugott, 2017):
Metaphor: Changes in which the source and target meanings are linked through analogy and/or similarity. before: in front → earlier (Traugott, 2000)
Metonymy: Changes in which the source and target meanings are linked through contiguity, i.e., spatial, temporal, or causal association (Geeraerts, 2020; Traugott, 2017). board: table → people sitting around a table; governing body (Traugott, 2000)
During annotation, we did not treat types of semantic change as mutually exclusive unless they belonged to the same category. For example, broadening can co-occur with metaphor, metonymy, pejoration, or amelioration, but not with narrowing. For each semantic shift pattern from DatSemShift, annotators selected at most one type of change per category, using three separate annotation columns. When reporting confidence scores, annotators provided a single overall confidence score for each semantic change pattern rather than evaluating their confidence for each individual type of change. The confidence scores reported in the dataset represent a combined measure of the confidence scores from all annotators.
When this manual annotation was completed, we automatically extended the annotations to the individual realizations belonging to each pattern. However, during this process, we observed that some change patterns included bidirectional realizations (i.e., in some cases, the source and target meanings were reversed). We also found cases in which either the source or the target meaning corresponded to a near-synonymous meaning within the same pattern but differed in connotation. Because these differences can affect the annotations, particularly those that include change in extension and emotive value, the first author of this article manually reviewed the entire dataset to identify and correct such cases.
3 Dataset Description
Repository name
The dataset is hosted on Zenodo (https://doi.org/10.5281/zenodo.20785477, DOI: 10.5281/zenodo.20785477).
Object name
TypesofSemanticChangeData.csv: The annotated dataset
AnnotationGuide.pdf: The annotation guide
Languages.csv: Languages appear in the dataset and their frequencies
Format names and versions
The annotated dataset, “TypesofSemanticChangeData.csv” and list of languages, “Languages.csv” are provided in Comma-Separated Values (CSV) format. The annotation guide, “AnnotationGuide.pdf”, is provided in Portable Document Format (PDF).
Creation dates
The data was extracted and annotated in October 2025.
Dataset creators
Ecesu Ürker (Universitat Pompeu Fabra): Data extraction from DatSemShift, determining definitions of each type of semantic change and the criteria of annotation, creation of the Annotation Guide, data annotation, writing
Gemma Boleda (Universitat Pompeu Fabra, Catalan Institution for Research and Advanced Studies): Determining definitions of each type of semantic change and the criteria of annotation, creation of the Annotation Guide, data annotation, reviewing, editing
Paola Gonzalez Triana (Universitat Pompeu Fabra): Data annotation
Language
The metalanguage of the dataset is English. The data in the dataset come from 792 languages and a list of these languages and their frequencies can be accessed within the published dataset.
License
Creative Commons Attribution 4.0 International
Publication date
28-04-2026
4 Reuse Potential
Although considerable attention has been devoted to detecting semantic change and developing datasets for this purpose, the different types of semantic change have often been overlooked; as Cassotti et al. (2024) noted, more studies and resources focusing on this phenomenon are needed. To our knowledge, this dataset is the first annotated resource of semantic change types designed to capture cross-linguistic patterns. Unlike previous resources, it does not rely on hand-picked examples (Blank, 2012), and is annotated according to widely accepted categories of semantic change from the theoretical literature.
As such, this dataset can be reused in future quantitative studies of semantic change. Since it includes both completed semantic changes and cases of synchronic polysemy, it may also be used in future research investigating the differences between these phenomena. In addition, the dataset could serve as a benchmark for computational approaches aimed at detecting different types of semantic change in multi-lingual settings. Finally, the semantic change patterns included in the dataset may provide illustrative examples for theoretical research and teaching purposes.
Furthermore, the rationale for the annotations provided in the dataset as well as the confidence scores facilitate its use and adaptation for different scientific purposes. Last but not least, the annotation guidelines developed for this work, and released with the dataset, can be reused in future annotation studies.
The main limitation of the dataset is its size: even though it contains 2,782 data points (realizations, or words undergoing the relevant shifts) in 792 languages, they all come from 234 semantic change patterns. Therefore, findings derived from it should be interpreted with caution. Furthermore, because the initial data were extracted from DatSemShift, the dataset may reflect biases present in that source. Nevertheless, the dataset can serve as initial source of quantitative evidence and as a foundation for future work, as it may be expanded to create a larger and more comprehensive resource covering different types of semantic change and languages.
Notes
[1] The code for data extraction can be found at https://github.com/ecesuurker/DatSemShift_Semantic_Change_Types.
Acknowledgements
The authors thank the COLT group, the Evolang 2026 reviewers and audience, and the SLE 2026 reviewers for useful feedback on this project.
Author Contributions
Ecesu Ürker: conceptualization, methodology, software, writing-original draft.
Gemma Boleda: conceptualization, supervision, writing-review & editing.
