Skip to main content
Have a personal or library account? Click to login
ChroniclItaly 3.0: An Enriched Italian American Newspaper Corpus for Migration History and Quantitative Diachronic Linguistics Cover

ChroniclItaly 3.0: An Enriched Italian American Newspaper Corpus for Migration History and Quantitative Diachronic Linguistics

By:   
Open Access
|Jul 2026

Abstract

ChroniclItaly 3.0 is an open dataset of digitized front pages from ten Italian-language newspapers published in the United States between 1898 and 1936. Hosted on Zenodo, the collection comprises 8,653 issues and more than 21.4 million words. It includes both original and processed text together with enrichment outputs such as named entities, geocoding, sentiment, and network-analysis files. Originally assembled to study migrant narratives and the Italian American diaspora through the ethnic press, the dataset also has strong reuse potential for quantitative diachronic linguistics, including lexical change, language contact, discourse variation, and historical NLP. Its dual value as both heritage collection and historical corpus makes it suitable for interdisciplinary reuse.

DOI: https://doi.org/10.5334/johd.566 | Journal eISSN: 2059-481X
Language: English
Page range: 94 - 94
Submitted on: Apr 15, 2026
Accepted on: Jun 18, 2026
Published on: Jul 14, 2026
Published by: Ubiquity Press
In partnership with: Paradigm Publishing Services

© 2026 Lorella Viola, published by Ubiquity Press
This work is licensed under the Creative Commons Attribution 4.0 License.