(1) Overview
Repository location
Context
ChroniclItaly 3.0 was developed in the context of the project Deep teXT minER (DeXTER) (Viola and Fiscarelli 2021a), a deep learning, interoperable, critical workflow to contextually enrich digital collections and interactively visualize them. The corpus extends earlier versions of the resource, namely ChroniclItaly (Viola 2018) and ChroniclItaly 2.0 (Viola 2019). These first two versions, described in Viola (2021) and covering 4,810 issues from seven titles from 1898 to 1920 and 16,624,571 words, are effective resources to access migrant narratives through the Italian immigrant press. While versions 1.0 and 2.0 share the same seven-title core corpus, ChroniclItaly 3.0 substantially extends the collection in terms of titles, chronological range, corpus size, and enrichment layers, moving from a heritage corpus and entity-annotated dataset to a more fully contextually enriched resource. Table 1 shows the progression of the three versions.
Table 1
Comparison of ChroniclItaly, ChroniclItaly 2.0, and ChroniclItaly 3.0.
| FEATURE | CHRONICLITALY | CHRONICLITALY 2.0 | CHRONICLITALY 3.0 |
|---|---|---|---|
| Repository/DOI | Utrecht University 10.24416/UU01-T4YMOW | Utrecht University 10.24416/UU01-4MECRO | Zenodo 10.5281/zenodo.4596345 |
| Public release year | 2018 | 2019 | 2021 |
| Chronological coverage | 1898–1920 | 1898–1920 | 1898–1936 |
| Titles | 7 | 7 | 10 |
| Issues | 4,810 | 4,810 | 8,653 |
| Word count | 16,624,571 | 16,624,571 | 21,454,455 after preprocessing; 30,752,942 before intervention |
| Geographical coverage | Italian-language newspapers published in U.S. immigrant communities; California, Pennsylvania, Vermont, and West Virginia | Same core corpus as v1 | California, Connecticut, Pennsylvania, Vermont, and West Virginia |
| Material included | Digitized front pages of seven Italian immigrant newspapers | Same as v1 | Digitized front pages of ten Italian immigrant newspapers |
| Main enrichment/annotation layer | Corpus compilation | Entity annotation for people, places, and organizations | Processed and unprocessed text, NER, geocoding, entity sentiment analysis, and network-analysis outputs |
| Preprocessing/curation profile | No enrichment-oriented preprocessing pipeline | Entity annotation added to the core corpus | Explicit preprocessing pipeline: tokenization; removal of numbers, dates, very short tokens, and isolated special characters; correction of split words; selective retention of punctuation and stopwords |
| Main contribution relative to previous version | Establishes the reusable core corpus | Adds structured entity annotation to the original corpus | Expands corpus size, titles, and time span, and adds contextual enrichment layers |
(2) Method
Steps
ChroniclItaly 3.0 contains digitized front pages of ten Italian-language newspapers published in California, Connecticut, Pennsylvania, Vermont, and West Virginia between 1898 and 1936. The corpus contains 8,653 issues and 21,454,455 words. The ten titles are L’Italia, Cronaca sovversiva, La libera parola, The patriot/Il Patriota, La ragione, La rassegna, La sentinella del West Virginia, L’Indipendente, La Sentinella, and La Tribuna del Connecticut. The material was machine-harvested from Chronicling America,1 the Library of Congress’s open repository of historical U.S. newspapers. The newspaper OCR underlying ChroniclItaly 3.0 were not generated directly by DeXTER but inherited from Chronicling America, which in turn reflects the selection and digitization logic of the National Digital Newspaper Program. In that framework, newspapers are prioritized on the basis of their historical significance, their value as papers of record, their geographical representativeness, the length of their chronological runs, and, crucially, the existence of complete or near-complete, good-quality microfilm (Beals and Bell 2020). The ChroniclItaly collections therefore inherit both the advantages and the limitations of this infrastructure: they benefit from access to historically significant and preserved newspapers, but they also reflect earlier archival decisions about what was microfilmed and digitized, as well as uneven OCR quality linked to the material and technical conditions of digitization.
The ChroniclItaly 3.0 workflow combined corpus acquisition with computational enrichment. The enrichment steps were designed to support the exploration of migrant narratives, diasporic public discourse, and the role of the Italian immigrant press as a site of transatlantic information exchange (Viola, 2023). The collection includes two main textual states: an unprocessed version, corresponding to the harvested OCR text, and a processed version produced through a series of pre-processing interventions, including tokenization (word-level), removal of numbers and dates, removal of words shorter than two characters, removal of isolated special characters, and correction of words incorrectly split by line breaks, whitespace, or punctuation. We deliberately retained punctuation and stopwords when they were deemed useful for subsequent enrichment tasks, especially named entity recognition (NER), geocoding, and entity-level sentiment analysis. These pre-processing steps were intended to reduce the high number of OCR errors. However, users of the corpus can still access the unprocessed version included in the archive for their own purposes.
After this stage, the processed collection totaled 21,454,455 words. Table 2 reports for each of the ten newspaper titles included in ChroniclItaly 3.0, the temporal coverage, number of issues, token counts before and after pre-processing, and retention rate after intervention.
Table 2
Title-level composition of ChroniclItaly 3.0.
| TITLE | COVERAGE PERIOD | SPAN (YEARS) | ISSUES | TOKENS BEFORE INTERVENTION | TOKENS AFTER INTERVENTION | RETENTION AFTER INTERVENTION (%) |
|---|---|---|---|---|---|---|
| La sentinella del West Virginia | 1911-02-18 to 1912-05-11 | 1.2 | 53 | 124,988 | 86,839 | 69.5 |
| L’Italia | 1897-01-25 to 1919-12-31 | 22.9 | 6,489 | 24,584,681 | 17,123,478 | 69.7 |
| Cronaca Sovversiva (Barre, Vt.) | 1903-06-06 to 1919-05-01 | 15.9 | 771 | 1,825,758 | 1,289,239 | 70.6 |
| La Libera Parola | 1918-04-20 to 1922-12-23 | 4.7 | 241 | 1,079,122 | 752,309 | 69.7 |
| Il Patriota | 1914-08-08 to 1921-10-22 | 7.2 | 227 | 786,762 | 547,361 | 69.6 |
| La Ragione | 1917-04-25 to 1917-08-23 | 0.3 | 40 | 126,937 | 86,588 | 68.2 |
| La Rassegna | 1917-04-07 to 1917-08-25 | 0.4 | 25 | 86,401 | 60,116 | 69.6 |
| L’Indipendente | 1907-01-01 to 1936-05-23 | 29.4 | 48 | 133,872 | 90,758 | 67.8 |
| La Sentinella | 1920-04-17 to 1930-12-27 | 10.7 | 518 | 1,674,336 | 1,193,015 | 71.3 |
| La Tribuna del Connecticut | 1906-03-03 to 1908-12-09 | 2.8 | 130 | 330,085 | 224,752 | 68.1 |
The corpus also includes derivative outputs for NER, geocoding, sentiment analysis, network analysis, and metadata/documentation files, alongside code made available through the DeXTER workflow repository (Viola and Fiscarelli 2021a).
NER was performed with a deep-learning sequence-tagging approach based on BiLSTM-CRF (Riedl and Padó 2018), implemented in TensorFlow and using FastText-based embeddings. The Italian model was trained on the Italian Content Annotation Bank (I-CAB)2 corpus and tagged four major entity classes: person, organization, location (LOC), and geopolitical (GPE) entity. The initial run extracted 547,667 entities occurring 1,296,318 times; after manual intervention and exception handling, the adjusted output counted 521,954 unique entities occurring 1,205,880 times.
Geocoding was applied to geopolitical entities using the Google Cloud Natural Language API,3 with ‘Italian’ set as the processing language. In this workflow, we decided to geo-code only GPE entities. Though not optimal, the decision was made considering that GPE entities are generally more informative as they typically refer to countries and cities (though it was found to retrieve also counties and States) while LOC entities are typically rivers, lakes, and geographical areas (e.g., the Pacific Ocean). Depending on individual researcher’s needs, LOC-entities can also be geocoded by using the code and the scripts in the dedicated folder of the DeXTER GitHub repository (Viola and Fiscarelli 2021a). In total, 2,160 GPE-type entities were geocoded, corresponding to 283,879 mentions in the corpus. Since the Google Places database is based on the contemporary geopolitical landscape, the geocoding stage required a number of critical manual interventions. These were necessary to address historical place names that had changed over time, places whose political status had shifted, or locations that no longer exist. A detailed list of these interventions can be found in the DeXTER GitHub repository (Viola & Fiscarelli, 2021a) alongside the output dataset preceding the intervention.
Entity sentiment analysis was then carried out on sentence-level extracts containing high-frequency entities selected using a logarithmic weighting function to balance differences in newspeapers’ size. This involved segmentation into 677,030 sentences using punctuation-based sentence delimiters (i.e., full stop, semicolon, colon, exclamation mark, question mark), selection of the most frequent entities via logarithmic weighting function, yielding 228 entities, extraction of 133,281 sentences containing those entities, and sentiment analysis through Google Cloud tools, as no suitable Italian-language model was available at the time.
Network analysis was used to examine relationships between entities across the corpus rather than treating them in isolation. Entities were modelled as nodes in a co-occurrence network, with links created when entities appeared in the same sentence; these links were enriched with attributes such as newspaper’s issue, date, co-occurrence frequency, and sentiment, while node-level attributes were derived by aggregating the properties of connected edges. Figure 1 shows an example of the entity-co-occurrence network.

Figure 1
Example of one ChroniclItaly 3.0 entity co-occurrence network.
Sampling strategy
ChroniclItaly 3.0 is a purpose-built corpus selected to represent different strands of the Italian immigrant press, including mainstream (prominenti), anarchic (sovversivi), and independent newspapers. This design supports a differentiated view of migrant public discourse. At the same time, coverage varies across titles because of the historical survival of issues, the uneven publishing histories of newspapers, and the economic instability of immigrant press enterprises (Viola 2021). Users should therefore treat the corpus as historically rich but unevenly distributed across time and titles.
Quality control
While formal error rates were not systematically measured across all enrichment stages, data integrity was addressed through a combination of pre-processing, benchmarked NER models, manual post-processing of the enrichment outputs, and quality control operated at multiple stages. First, pre-processing was used to reduce OCR noise while preserving linguistically useful material where possible. This critical intervention affected about 30% of the material. This intervention improved interpretability but also entailed possible information loss. Second, for the NER workflow we reported benchmark scores of 98.15% accuracy and 82.88 F1; its output on ChroniclItaly 3.0 was then manually revised to correct misclassified, duplicated, or spurious entities before downstream geocoding and sentiment analysis. A list of exceptions was compiled and the entity output adjusted accordingly. This historically informed post-processing step is important for reuse because it improves downstream accuracy for geocoding and sentiment analysis. Third, the dataset is distributed with both original and processed versions, which allows users to evaluate the effects of intervention and choose the level of processing best suited to their own research questions (Viola and Fiscarelli 2021b).
(3) Dataset Description
Repository name
Zenodo
Object name
ChroniclItaly 3.0. A deep-learning, contextually enriched digital heritage collection of Italian immigrant newspapers published in the USA, 1898–1936.
Format names and versions
ZIP
Creation dates
2018-06-21 to 2019-09-23 (iterative corpus development across ChroniclItaly versions)
Dataset creators
Lorella Viola, Vrije Universiteit Amsterdam. Corpus creation, conceptual development, quality control.
Antonio Maria Fiscarelli, Experis Italia. Contextual enrichment and computational processing.
Language
Italian, English
License
CC BY 4.0 license
Publication date
2021-03-21
(4) Reuse Potential
ChroniclItaly 3.0 has already demonstrated value for migration, heritage, and digital-history research (Viola, 2021, 2022), but its reuse potential extends further into quantitative diachronic linguistics.
First, the corpus can support lexical-semantic change research, as already demonstrated in Melis et al. (2023) where it was incorporated into a larger historical Italian corpus used to model diachronic semantic shifts with distributional methods. Spanning nearly four decades, ChroniclItaly 3.0 enables the study of changing vocabularies of migration, labor, nationhood, class, citizenship, ethnicity, religion, and political identity. Because the corpus contains newspapers of different political orientations, it also allows comparison between ideological registers and the tracing of how key terms acquire, lose, or shift meanings over time. This makes the resource useful not only for Italian diaspora history but also for historical semantics and discourse analysis as well as for comparative work on broader patterns of semantic evolution in Italian.
Second, the dataset is valuable for studying language contact and multilingual influence. Italian immigrant newspapers published in the United States preserve traces of contact between Italian and English, including loanwords, calques, code-switching practices, orthographic variation, and changing references to American institutions and places. This opens the dataset to sociolinguists and historical linguists interested in diaspora Italian, contact-induced change, and immigrant-language adaptation in print culture.
Third, the corpus can be reused for historical NLP and corpus-linguistic method development. Since the deposit includes both unprocessed and processed text as well as NER, geocoding, sentiment, and network outputs, it is suitable for benchmarking OCR-sensitive pipelines, evaluating named-entity recognition on historical Italian, comparing annotation strategies, and testing reproducible workflows for non-English historical corpora. The availability of both raw and intervened data is especially valuable for methodological work because it makes the effects of pre-processing visible rather than hidden.
Fourth, the dataset has strong reuse value for spatial and networked diachronic analysis. The geocoding and network-analysis files make it possible to study changing geographies of belonging, transatlantic place-reference patterns, and relational structures among actors, institutions, and locations. These same layers can also support research on historical event diffusion, mobility narratives, and transnational public spheres.
Fifth, the corpus is useful in teaching and training. Because it is open, historically bounded, and accompanied by derived files, it can be used in courses on corpus linguistics, digital humanities, migration history, historical sociolinguistics, and computational text analysis. Students can work with different processing levels, compare raw OCR and curated outputs, and reflect critically on how pre-processing and annotation shape research results.
There are, however, limitations and barriers to reuse. The corpus consists of front pages only, so it does not capture the entirety of each issue. It is also uneven across titles and years, reflecting survival biases, publication histories, and source availability. OCR quality varies, and the enrichment layers inevitably inherit errors from both OCR and model-based annotation. Some of the computational tools were selected because suitable Italian alternatives were limited at the time, which means future users may wish to re-run parts of the workflow with newer historical-Italian models. Finally, because pre-processing removed a substantial proportion of material, researchers should choose carefully between original and processed files depending on whether their task prioritizes textual fidelity or computational tractability.
Notes
AI Declaration
Generative AI was used in the final stage of this manuscript to support wording revision. All factual content, dataset descriptions, and references are by the author who takes full responsibility for the submission.
Author Contributions
Lorella Viola: Conceptualization; Data curation; Methodology; Investigation; Validation; Writing – original draft; Writing – review & editing.
