
Classical Greek in the Open Greek and Latin Project
Abstract
This discussion paper outlines the forty-year development and current state of the Classical Greek component from the Open Greek and Latin project. Reflecting a shift from proprietary search systems and manual data entry to an open-access ecosystem, the collection aggregates approximately 41 million words synthesized from multiple repositories, including the Perseus Digital Library and the First One Thousand Years of Greek project. The paper details the methodological evolution of the collection, describing the transition from initial SGML encoding and double-keying in the 1980s to modern workflows that leverage high-quality OCR and Large Language Models (LLMs) for complex document recognition and data curation. A central contribution is the provision of chronological metadata for 1,946 texts, enabling diachronic linguistic analysis of Classical Greek from the 8th century BCE through the 13th century CE.
© 2026 Gregory Crane, Alison Babeu, Lisa Cerrato, Rhea Lesage, Leonard Muellner, Bruce Robertson, Lucie Stylianopoulos, published by Ubiquity Press
This work is licensed under the Creative Commons Attribution 4.0 License.