Abstract
This paper presents a curated, ready-to-use corpus of 10,898 French novels published between 1800 and 1950, derived from the Fictions littéraires de Gallica collection. The corpus is designed as an infrastructure for the empirical study of nineteenth-century popular fiction — the literature hors d’usage that once constituted the bulk of novelistic production but has since fallen out of circulation and out of literary memory. A reproducible pipeline filters low-quality OCR scans, removes editorial duplicates through a hybrid procedure combining MinHash near-duplicate detection and rule-based metadata matching, and enriches each record with normalized author identities and BnF/Wikidata-cross-referenced biographical metadata. Two genre labels — adventure (1,428 novels) and detective fiction (830 novels) — are provided as flags in the metadata table, illustrating one form of downstream reuse. The dataset, the pipeline, the per-novel decisions, and the audit logs are openly released.
© 2026 Jean Barré, published by Ubiquity Press
This work is licensed under the Creative Commons Attribution 4.0 License.
