
ChroniclItaly 3.0: An Enriched Italian American Newspaper Corpus for Migration History and Quantitative Diachronic Linguistics
Abstract
ChroniclItaly 3.0 is an open dataset of digitized front pages from ten Italian-language newspapers published in the United States between 1898 and 1936. Hosted on Zenodo, the collection comprises 8,653 issues and more than 21.4 million words. It includes both original and processed text together with enrichment outputs such as named entities, geocoding, sentiment, and network-analysis files. Originally assembled to study migrant narratives and the Italian American diaspora through the ethnic press, the dataset also has strong reuse potential for quantitative diachronic linguistics, including lexical change, language contact, discourse variation, and historical NLP. Its dual value as both heritage collection and historical corpus makes it suitable for interdisciplinary reuse.
© 2026 Lorella Viola, published by Ubiquity Press
This work is licensed under the Creative Commons Attribution 4.0 License.