
A Diachronic Corpus of Modern Indonesian Literature Between 1920 and 2000
Abstract
This dataset presents a diachronic corpus of modern Indonesian literature comprising 58 literary works published between 1920 and 2000. The corpus contains approximately 1.97 million tokens and captures three orthographic periods in Indonesian literary writing. Texts were digitized from printed sources using Optical Character Recognition (OCR), then semi-manually cleaned, tokenized, orthographically standardized, lemmatized, and annotated for historical orthography. Each work is documented using a metadata schema adapted from Dublin Core, supplemented with fields for literary genre, orthography system, and annotation level. The repository provides open corpus extracts and accompanying metadata for works eligible for public distribution. The dataset supports research in corpus linguistics, language change, Indonesian literary studies, and digital humanities.
© 2026 Mohammad Rokib, A. Noorhidawati, Moh. Mudzakkir, Lutfiyah Alindah, published by Ubiquity Press
This work is licensed under the Creative Commons Attribution 4.0 License.