Skip to main content
Have a personal or library account? Click to login
A Diachronic Corpus of Modern Indonesian Literature Between 1920 and 2000 Cover

A Diachronic Corpus of Modern Indonesian Literature Between 1920 and 2000

Open Access
|Jul 2026

Abstract

This dataset presents a diachronic corpus of modern Indonesian literature comprising 58 literary works published between 1920 and 2000. The corpus contains approximately 1.97 million tokens and captures three orthographic periods in Indonesian literary writing. Texts were digitized from printed sources using Optical Character Recognition (OCR), then semi-manually cleaned, tokenized, orthographically standardized, lemmatized, and annotated for historical orthography. Each work is documented using a metadata schema adapted from Dublin Core, supplemented with fields for literary genre, orthography system, and annotation level. The repository provides open corpus extracts and accompanying metadata for works eligible for public distribution. The dataset supports research in corpus linguistics, language change, Indonesian literary studies, and digital humanities.

DOI: https://doi.org/10.5334/johd.564 | Journal eISSN: 2059-481X
Language: English
Page range: 99 - 99
Submitted on: Apr 10, 2026
Accepted on: Jun 26, 2026
Published on: Jul 22, 2026
Published by: Ubiquity Press
In partnership with: Paradigm Publishing Services

© 2026 Mohammad Rokib, A. Noorhidawati, Moh. Mudzakkir, Lutfiyah Alindah, published by Ubiquity Press
This work is licensed under the Creative Commons Attribution 4.0 License.