1. Context and Motivation
The collection of Classical Greek1 described here reflects almost forty years of continuous development and began long before the term Digital Humanities had become current. A history of this collection follows more than forty years of evolution in the application of digital methods to the humanities. For reasons that will be discussed below, the collection is split into two complementary repositories.2 Each GitHub repository also stores its data on Zenodo, with updates synced.3 The two repositories have been developed in parallel so that they can be combined without duplication. We will, for convenience, refer to the combined collection as Open Greek. The combined collection contains a total of 41 million words, which have been postprocessed and are available via the Perseus Scaife Viewer.4 Of these, we have assigned 34 million words to particular centuries for diachronic analysis.
For those who wish to analyze classical Greek language over time in the short term, linguistically annotated corpora, largely derived from the corpus described here, are available and provide an immediate starting point: Alex Keersmaaker’s GLAUx5 – 20 million words: (Keersmaekers et al., 2024) and Giuseppe Celano’s Opera Graeca Adnotata6 – 45 million words: (Celano, 2024). That said, no linguistic annotation dataset has fully exploited the metadata encoded in, or aligned with, the Open Greek corpus. Those looking to do more extended work on the evolution of classical Greek from Homer through the classicizing Greek of the 19th century should be able to link downstream annotations with the original corpus.
The work described here began in 1982, when the lead author, then a graduate student, began developing an information retrieval system for Greek texts from the Thesaurus Linguae Graecae (TLG). This initial corpus covered a number of the most important sources and established the beginning for the most comprehensive collection of Classical Greek. Even though the TLG distributed its data on magnetic tape and later on CD ROMs, the collection was not open.
Licensing severely restricted reuse and those restrictions were vigorously enforced. Ultimately, the TLG7 became a closed database available via its own proprietary search function.
For many students of Classical Greek, a proprietary database with unverifiable results was not only satisfactory but preferable. The fixed search interface, in essence, extended the model of print indices and concordances with which classicists had worked for generations. Traditional philologists were comfortable working with this model and did not need to develop skills in computation.
Those who wished not only to analyze textual data themselves but to create and enhance versions of texts needed open data. Already in our first work making Greek available, we added indices to speed up searching and to allow readers to browse the collection by author and work – until then, researchers could only search for words. For those who already had access to texts from the TLG we created a search system that was written in C, ran under the Unix operating system and built on the supplementary indices that we had created. We scrupulously left the TLG files untouched so as not to violate the licensing agreements (Crane, 2004; Crane, 2022).
It soon became clear that we would need to be able to modify Greek texts directly. The TLG used a system called Beta Code that defined not only a (still very useful) transliteration scheme but also a complex set of code by which to represent the appearance of the original printed pages. Line breaks and even hyphens were preserved. Features such as italics, bold and expanded text (common because bold and italic Greek fonts were not readily available) were encoded. Block quotes were represented as indents. The old Beta Code scheme provided mechanisms to represent the surface structures that appeared in a wide range of Greek editions.
The first proposal for Perseus was submitted in September 1985. A planning phase began in 1986, followed by prototyping and then full-scale development in 1987. As part of the larger Perseus budget, we were able to include funding to enter an initial set of core Greek texts (c. 1 million words). Three aspects of this initial, and now almost antediluvian, phase deserve note.
First, almost all the data entry relied upon a professional data entry firm. Two individuals would separately key in each text and a third person would correct those places where the two initial copies deviated. This method provided textual data in which at least 99.95% of the keystrokes were correct. The cost at the time was $2,300 per million keystrokes (e.g., $2,300 for the Greek New Testament or c. $19,000 for one million words), with each diacritic counting as a separate keystroke. For an additional fee, we could have had a third person key in the text and thus then received a guaranteed accuracy of 99.995%. In practice, the data entry firm exceeded the 99.95% contractual minimum.
Optical Character Recognition (OCR) systems did already exist in the 1980s. They were dedicated machines that cost c. $5,000 and had a limited capacity for training. The University of Chicago used their Kurzweil reader to perform the initial data entry for Sophocles and Herodotus (the XML logs still record the names of those who cleaned up this work in 1988). In effect, we needed to retype the entire text and we would not use OCR again to enter Greek for another twenty years (Stewart, Crane & Babeu, 2007; Boschetti et al., 2009).
Second, early in our planning process, we began working with Ellie Mylonas and her colleagues at Brown University. Together they would publish a now classic paper, “What is text, really?” (DeRose et al., 1990). Ellie and her colleagues convinced us of two things. First, we needed to go beyond encoding surface features such as italics and capture instead the reason why italics had been used (e.g., to mark a title or a foreign language quotation or just to add emphasis).
Second, we needed to follow some shared mechanism by which to represent this markup and not invent yet another ad hoc system. We adopted the Standard Generalized Markup Language (SGML),8 an ISO standard (ISO 8879) that was established in 1986 and that would be replaced by the Extensible Markup Language (XML),9 that was, among other things, more efficient for processing.
Third, our work encoding Greek texts, English translations, and the Middle Liddell Intermediate Greek Lexicon began before establishment of the Text Encoding Initiative (TEI).10 Initially developing an ad-hoc SGML tagset while awaiting the TEI recommendations, we subsequently revised our markup to bring it into compliance with the final standards.
Ironically, after investing an enormous amount of labor (at least one person year of work) in adding SGML tags to our texts, we immediately stripped the markup out of the text to publish them in Apple’s HyperCard system, which was the delivery vehicle for Perseus 1 and Perseus 2. A few years later, however, the markup proved its worth: David A. Smith was able to use the markup to create the first Web version of Perseus in 1995, rapidly converting SGML into an HTML format that could represent all the formatting that we wished to communicate. We had no separate funding for Web Perseus. David Smith was only able to create this as a side project because of the investment we had put into the data.
When the TEI was first being established, some of us at least imagined TEI encoded texts merging smoothly into interoperable corpora that would, in turn, enable new forms of research. In practice, the TEI provided such a rich markup language that annotators can use different terms to represent the same object (e.g., <rs type=”person”> vs. <persName>). An even larger challenge emerges when we decide how to encode the complex citation schemes for canonical texts. We tried multiple approaches over the years until deciding upon a schema that we felt was effective. Even so, it took years to convert all of our encoded texts into the new format.
Shifting from TEI milestones to containers (<div>) requires restructuring features such as speeches (e.g., breaking down a single speech into multiple chunks, each within a separate container).
These shifts in encoding standards occurred alongside the steady expansion of our holdings over several decades. Below, we use the following milestones to summarize the history of the collection:
1987–1992: Perseus 1.0: the development of first 1 million word SGML corpus of Greek
1992–1996: Perseus 2.0 and Perseus on the Web (overlapping development): the collection expands to cover most widely read sources (c. 4 million words).
1996–2007: incremental additions added, with small amounts of data entry funded as parts of multiple funded projects. Data continue to be added by double keying.
2001: planning begins for the Perseus Treebanks of Greek and Latin.
2007: David Bamman, then a researcher at Perseus, coordinates the development of a shared tagset, based on the Prague Dependency Grammar, for treebanking Latin with Marco Passarotti, while maximizing compatibility with the treebanking tagset that the Norwegian Proiel project adopted.
2006: Following the example of the Papyri.info, Perseus adopts Creative Commons and releases its data under a CC-BY-SA license.
2008–2013: OCR technology reaches a point where it can provide a usable starting point and we begin adding texts by doing in-house correction. The first OCR engines at this point could transcribe Greek characters with reasonable accuracy (Stewart, Crane & Babeu, 2007) but not Greek diacritics. We developed mechanisms to add accents semi-automatically. Our copy of the Greek Anthology was added at this point and still needs final correction.
2009–2010: As part of the joint Canadian, UK and American Dynamic Variorum project,11 Mount Allison’s Bruce Robertson began optimizing open-source OCR platforms for Classical Greek. For the first time, we could use OCR to convert page images into accented Greek at an accuracy that was good enough for efficient post-correction. We continued to use data entry firms for most of the corrections until 2019, but we could pay for more and we could much more readily do our own correction. Robertson continued enhancing open-source OCR for more than ten years (Robertson, 2019) and laid the foundation for our ability to go from 10 million to more than 40 million words.
2013: As Alexander von Humboldt Professor at Leipzig, Crane creates the Open Philology Project, with subprojects such as Open Greek and Latin (OGL). We create OGL12 so that we have a framework that is distinct from, but closely aligned with, Perseus. Maxim Romanov, as part of the Humboldt Chair, plays a key role in establishing the Open Islamicate Texts Initiative.13
2013: Perseus and Open Philology adopt the Canonical Text Services (CTS) data model (Smith, 2009) so that they can precisely cite any token in any version of their text collections. An XML implementation of the work begins on retrofitting Perseus texts
2016: The CapiTainS Software Suite14 and Guidelines for Citable Texts (Clérice et al., 2017, May 2) provides tools by which to validate CTS-compliant TEI XML files and directory structures.
2016: Representatives from the Humboldt Chair, the Harvard College Library, Harvard’s Center for Hellenic Studies (CHS), Mount Allison University, the University of Virginia Library, and Tufts University develop a consortium to expand the amount of openly licensed Greek data. This will become the First One Thousand Years of Greek Project.15 The primary goal of this project is to provide strong coverage for Greek from the Homeric Epics through the third century CE, although First1KGreek does include some later texts of great importance (e.g., the Byzantine Encyclopedia called the Suda and Scholia to authors such as Homer and Pindar).
2017–2026: While some texts are added, attention shifts to curating and augmenting the existing Greek corpus (e.g., problematic texts processed with earlier OCR-workflows such as Philo of Alexandria and Lucian are upgraded).
2018: Funded by the CHS and the Humboldt Chair, the Scaife Viewer goes live. Built on CapiTainS and on the CTS data mode, the Scaife viewer provides access to (and demonstrates compatibility with) the CTS data model.
2024: OCR for Classical Greek, produced by Google and published through the HathiTrust, begins to exceed the quality of what can be produced with public systems such as Tesseract and Kraken. In particular, the new OCR can switch back and forth between Greek and non-Greek words in sections on an edition such as the textual notes. One major problem remains: the Google OCR often gets confused when dealing with multiple columns and even where there are spaces between chunks of words (as in textual notes).
June 2025: The Harvard Institutional Data Initiative published ID1: the new Google OCR-generated text for just under 1 million books, published through the beginning of the 20th century, from the Harvard library collections (Cargnelutti et al., 2025). While we can download individual books for curation, this new aggregate collection makes it possible for us to search for Greek editions at a far wider scale.
August 2025: David Smith, Alison Babeu and Cliff Wulfman create a new corpus of Classical Greek of 600 million words. Roughly half of this text – 300 million words – consists of new editions for Greek texts already available in the curated XML collection. The other 300 million words consist of new Greek texts, including more than 100 million words of classicizing Greek from the 19th century (katharevousa).
November 2025: Google introduces Gemini 3, a model that has the ability not only to provide very accurate OCR generated text for mixed Greek and Latin fonts but has resolved problems with within lines being scrambled. Furthermore, the LLM is able to recognize complex document types such as critical editions (and the accompanying textual notes), dictionaries (with multiple word senses) and grammars (with coverage of reasonably complex paradigms).
2. Dataset description
Github repositories for ongoing contributions: https://github.com/PerseusDL/canonical-greekLit/ and https://github.com/OpenGreekAndLatin/First1KGreek Repository Locations: Perseus Canonical Greek: https://doi.org/10.5281/zenodo.19581321; the First Thousand Years of Greek: doi.org/10.5281/zenodo.596723.
Repository name: Zenodo
Format names: epiDoc TEI XML; URNs based on the Canonical Text Services protocol.
Most recent Zenodo creation dates: Perseus: May 29, 2026; First1KGreek: May 29, 2026
License: Creative Commons Attribution Share Alike 4.0 International
Publication Date: April 19, 2026
The Scaife Viewer currently aggregates 41 million words of Greek from Perseus and First 1KGreek repositories. From this corpus, we have been able to provide chronological data that assigns, where practical, centuries of composition to works in the repository for 1,946 texts with 34.7 million words of text, ranging from the 8th century BCE through the 13th century CE.16 Our goal is not to provide a canonical chronology – there is too much debate as to when many Greek works were produced. And, of course, some authors (such as Philo of Alexandria and Plutarch) had careers that spanned centuries. Assigning centuries to Greek works seems, however, to us to be a pragmatic starting point. We expect others to revise our proposed dates according to their own views and surely provide (and hopefully redistribute) alternate chronological categories.
3. Method
This larger collection includes works to which we have not assigned a century of composition. The commentary preserved in medieval manuscripts (scholia) include textual data from multiple periods and do not reflect any given century. In some cases, we simply have not been in a position to add detailed markup for the components of a work. The Greek Anthology, for example, is a famous collection of shorter poems by different authors and reflecting different periods. In other cases, the date of composition simply cannot be determined: the Battle of Frogs and Mice (Batrachomyomachia) is a mock-epic poem for which proposed composition dates have varied over 800 years, from the Homeric epics through the Roman Empire.
The Open Greek collection varies widely across centuries, with best coverage for the 4th century BCE and the 2nd and 3rd centuries CE. When we are able to include the often voluminous Church Fathers, Byzantine and classicizing Greek produced after the fall of Constantinople in 1453 CE, the amount of later Greek will increase substantially.
As illustrated in Figure 1,

Figure 1
Chronological distribution of the Open Greek collection. The left panel displays the total word count by century (in millions), while the right panel shows the number of individual text files. Data is categorized by source: Perseus (blue) and First1KGreek (green).
The scale of this generic imbalance is visualized in Figure 2. While many – and perhaps most – students of Classical Greek focus on canonical poetic works such as the Homeric Epics and Athenian Drama, the vast majority of surviving Classical Greek is prose.

Figure 2
Distribution of genres in the Open Greek collection (Prose vs. Poetry). This visualization illustrates the disparity between the “canonical” focus of traditional Greek studies and the physical reality of the surviving record, where prose — encompassing historiography, philosophy, medicine, and technical treatises — constitutes the vast majority of the total word count.
While Classical Greek poetry does not include prose, many prose works do quote poetry. Those wishing to study the language of prose in the fifth and fourth centuries, for example, need to be able to distinguish the language of Thucydides and Plato from the poetic language of quotations from poetic sources such as Homer or from oracles delivered in poetic form. While some texts in the collection originally used the <lb/> milestone to identify line breaks in poetry and some texts still need updating, the vast majority of the texts in this collection now use the <l></l> tags to identify lines of poetry, making it easy for programs to recognize and separate out poetic lines quoted in a primarily prose text.
We have made initial, though by no means exhaustive, efforts to identify two other categories of text:
First, we use the <quote> tag to identify passages that one text has imported from another. The <quote> tag is particularly important when one prose work includes extracts from another prose work. The <quote> tag then becomes the mechanism by which to keep the imported text from affecting analysis of the main text.
Second, we use the <q> and <said> tags to represent speeches and dialogue to mark speech that is a part of a larger text (e.g., reported speech in Herodotus). The goal is to enable comparison of narrative from quoted speech.
The functional difference between these tags is visualized in Figure 3. First, we use the <quote> tag to identify passages that one text has imported from another. Second, we use the <q> and <said> tags to represent speeches and dialogue within a single text. As Figure 3 demonstrates, this distinction enables a more granular comparison of narrative versus quoted speech.

Figure 3
Distinguishing intertextual citations from internal speech. This figure illustrates the divergent use of TEI elements: the <quote> tag is employed to sequester imported text for independent analysis, while the <q> and <said> tags delineate speech acts from the surrounding narrative voice.
The distinction between narrative and speech is crucial for many texts, prose and poetic alike, from the Homeric Epics onwards. Markup in the Open Greek corpus represents only a first step in a much longer process of adding more and more refined metadata about speech. While our structural markup provides the basic boundaries of speech, the Canadian-German DICES Project (Forstall & Verhelst, 2026) demonstrates the potential for more refined, standoff metadata. DICES published standoff markup identifying speeches, speakers and other features in Greek and Latin epics from Homer to Nonnus. We will return to this below.
We emphasize one final feature of the Open Greek collection design. Many of the text search engines that researchers use include a single version of each text that has been judged the best available. Search and linguistic analysis of such collections cannot distinguish between texts where printed editions vary widely and any single edition has a high degree of scholarly conjecture (e.g., Aeschylus) from those texts where the textual tradition seems solid and editions vary little (e.g., Pindar). For those studying the reception of Greek literature, the most important versions of a text are not 21st century editions but editions that circulated in the past.
Most of our effort has gone into expanding coverage and collecting at least one edition for as many Greek works as possible. Nevertheless, the architecture of this collection has been designed to support multiple versions of the same work. Readers can see the beginnings of this attempt to provide multiple versions by viewing texts of Aeschylus’ Agamemnon,17 Lucian’s Zeuxis,18 and Aristotle’s Poetics.19 These examples demonstrate an ability to manage and align multiple editions. A great deal of work remains to provide automated services for analyzing and visualizing differences across versions and editions of the same work. CTS allows us to distinguish between editions. By contrast, the TLG assigns a new work ID to different editions of the same work; e.g., Aelius Aristides’s Against Plato is 0284.045 in the edition by Dindorf and now also 0284.061 in the edition by Lenz and Behr. The CTS architecture, using a third level “version” after “textgroup” and “work”, is far better in this respect, also considering conceptual models such as FRBR or, most recently, IFLA LRM
To standardize our markup practices across the corpus, we utilized the EpiDoc subset of the TEI XML tags20 (Elliott et al., 2006). Although designed originally for inscriptions, the EpiDoc TEI XML subset proved to be a good foundation for our needs. For now, the Perseus Greek corpus – for the most part – follows a single edition of each work without textual notes. Adding textual notes – and especially, information about how different editions differ – is a major focus for further work.
The most far-reaching – and, arguably, the most useful – design decision was to adopt the Canonical Citation Services (CTS) data model (which is part of the CITE architecture: Blackwell& Smith, 2020). For convenience, we use identifiers from the TLG, but this does not mean we used the same edition.
CTS uses five different fields to precisely define sections of a canonical text.
Textgroup: This is typically the author but can also describe a conventional group of texts by multiple hands (e.g., the Homeric Hymns, the Greek New Testament and the Greek Anthology are textgroups).
The work: e.g., the Iliad, the Gospel of John.
The version: This can be either primary witness (such as a papyrus or a manuscript) or a critical edition.
The canonical citation (e.g., book+line in the Iliad or book+chapter+section in Thucydides).
Token+index: This field allows us to define precise tokens. There are five separate instances of the Greek word kai (“and”) in the Henry Stuart Jones edition of book 1, chapter 1, section 1 of Thucydides’ History of Peloponnesian War. We can precisely cite any of the five by using kai[1], kai[2]. kai[3] etc.
The complete CTS URN for the example cited appears as follows: urn:cts:greekLit:tlg0003.tlg001.perseus-grc2:1.1.1:καì[2].
Perseus and the OGL Project adopted the CapiTainS directory structure, which itself builds upon the CTS data model. The main directories21 contain a directory “data/” within which there are directories for CTS textgroups and then subdirectories for works. Under this structure, the following GitHub directory collects all editions and translations of the Homeric Iliad: https://github.com/PerseusDL/canonical-greekLit/tree/master/data/tlg0012/tlg001.
The directory tlg0012 designates the Homeric epics while the subdirectory tlg001 aggregates versions of the Iliad.
Perseus and OGL conventionally include “-grc” in the edition field for Greek texts: e.g., data/tlg0012/tlg001/tlg0012.tlg001.perseus-grc2.xml designates a Greek version of the Iliad. There are some cases where works combine Greek and Latin (e.g., the Hermeneumata Pseudodositheana22) but for now a program can use this convention to extract Greek editions from the Open Greek repositories.
The CTS URNs work very well for prose texts structured as book/chapter/section, chapter/section and section where the sections begin and end at sentence breaks.
The Perseus Greek corpus implements the CTS data model in XML by providing information in the header. The TEI <refsDecl/> element (“Reference Declaration”) communicates how a given text can be segmented:
Figure 4 shows a <refsDecl> that allows the work to be segmented by chapter or by chapter+verse. Developers can create programs that scan each header for the appropriate ways to segment a file and then break the file up into its constituent components.

Figure 4
TEI <refsDecl> element specifying hierarchical citation schemes. This declaration defines the mapping between the XML structure and the Canonical Text Services (CTS) URNs, allowing the text to be addressed and retrieved at the level of chapter or chapter-and-verse.
Some texts – including two of the most important surviving authors – do not have citation schemes based on logical chunks of a text. Most scholars cite the works of Plato and Aristotle by referencing particular sub-portions of the page in earlier editions: editions of Plato roughly follow the 1578 edition published by Henricus Stephanus (a citation scheme that has its own Wikipedia page23) while editions of Aristotle typically use the page breaks and sections from the edition of Immanuel Bekker (Berlin, 1831). Researchers analyzing texts with citation schemes like these typically use a program to break the sentence into sentences (and hopefully keep track of which sentences go with which Stephanus/Bekker pages).
Poetry presents a similar – and much less arbitrarily inflicted – problem. We cite poetry by line number or by book+line number. Line breaks, however, do not, of course, always coincide with sentence breaks — Homeric epic exploits the fact that sentences can begin and end in the middle of poetic lines for poetic effect, beginning with the second line of the Iliad: line 1, “Goddess, sing the wrath of Achilles, son of Peleus,” is grammatically complete but line 2 begins with a modifier oulomenên, “murderous,” which attaches to “wrath.” In many cases, those analyzing the language of poetry need to break the text into sentences that can begin and end anywhere in the line.
Of course, Greek drama confronts us with the overlapping hierarchy problem as soon as we choose line numbers as our basic chunk. We use the TEI XML part attribute and then add an alphabetic suffix to the line number.
Figure 5 shows how we deal with split lines. The “part” attribute indicates that we do not have a whole line and a program can then prepare to identify the next part. We identify the first and second half of the line as lines 123 and 123b.

Figure 5
A line from Euripides’ Phoenician Women split across two speakers.
4. Results and Discussion
The use of CTS compliant epiDoc TEI XML has made it possible for different projects to reuse and augment these texts. A detailed discussion of how these texts have been reused must be deferred to another space. For now, we will mention a few examples. The major treebanks developed by hand and then by using manually curated treebanks to create models to annotate Greek collections with tens of millions of words have largely been built on Perseus texts (Celano, 2024; Keersmaekers et al., 2024). The University of Chicago republishes (and often enhances the markup of) Perseus texts with PhiloLogic 4.24 The DICES database uses CTS URNs to provide standoff markup identifying speeches and speakers and other features in Greek and Latin epic from Homer through Nonnus (Forstall & Verhelst, 2026). A number of individual scholars, who began by reusing proprietary Greek textual data, decided to shift to the Greek texts published by Perseus and OGL so that they could republish their data as well as their conclusions (Sansom, 2021; Sansom & Fifield, 2023).
5. Implications/Applications
The Open Greek collection provides a useful start for open and transparent – and thus genuinely philological – analysis of Classical Greek but an enormous amount remains to be done. Forty years of part time work by small groups has only begun to capture, in an open machine actionable form, thousands of years of manuscript and hundreds of years of print culture about Classical Greek. Machine learning has already played a foundational role through its applications in services such as OCR and automatic morphosyntactic analysis. Large Language Models (LLMs) now allow us to bootstrap new categories of linguistic annotation (e.g., co-reference resolution, classification of agents as human/divine or animate vs. inanimate) with which to augment our collections. More importantly, LLMs provide the first serviceable machine translation services for Classical Greek, making it possible for a much wider range of researchers to study Classical Greek language over time (Zainaldin et al., 2026; Ströbel & Maier, 2025; Wannaz & Miyagawa, 2024).
Improvements to machine learning and to LLMs will play a key role in three major areas of development.
First, a great deal of work can be done to enhance the markup within the existing Greek collection to make sure that we have annotated features such as external quotations and quoted speech. Much of this work can be done by analyzing the collection itself but an increasing amount can also be done by reapplying back to the Open Greek corpus new open data from projects such as DICES (with its data about speeches and speakers in Greek and Latin epic) or the DraCor Project (which has enhanced information about speakers in Greek in Perseus texts of Greek drama).
Second, we are in a position to expand open Greek data massively. Three foundational changes make this possible.
Google improved the performance of its OCR engines for languages such as Classical Greek and for texts that switch back and forth between Greek and Roman fonts. The default OCR-generated text available from the HathiTrust is now better than that which we have been able to produce by optimizing open-source OCR systems such as Tesseract and Kraken.
LLMs such as Gemini have now not only internalized this improved OCR but can also identify the structure of complex documents such as critical editions and create useful, if imperfect, TEI XML automatically. We have been experimenting with this new capability since Gemini 3 appeared in fall 2025. There are challenges. LLMs are so proficient that they will often rewrite passages that they find difficult and we need to be able to minimize and then later detect such changes. And, of course, LLMs are slow. We can process books one at a time easily enough but we will need to confront engineering challenges if we wish to apply LLMs to much larger corpora.
In June 2025, the Harvard Institutional Data Initiative (Cargnelutti et al., 2025) published Google’s improved OCR output for almost 1 million books from the Harvard libraries that had been published through the early 20th century. In August 2025, researchers working with Perseus were able to extract editions from this collection that included 600 million words of Greek. Roughly half of this – 300 million words – consists of new editions for works already in the open corpus. We thus have the raw materials by which to trace the evolution of our texts produced in multiple editions over generations. The other half of the extracted Greek – another 300 million words – adds Greek texts that are not yet in the Open Collection in any form. These include texts produced not only through the fall of Constantinople (1453 CE), the conventional end of classical Greek but also 100 million words of classicizing Greek produced through the 19th and early 20th centuries. We have the raw materials to study the evolution of Classical Greek for almost 3 millennia.
Notes
[1] The Library of Congress applies the term Ancient Greek through the fall of Constantinople in 1453 CE. While that designation is appropriate for all of the Greek that we have so far published, the conclusion of this paper points to an additional 100 million words of Greek, produced after 1453 CE, that is now available from Harvard’s million books of OCR (Cargnelutti et al., 2025). At least some – and hopefully all – of that Greek will be available in CTS-compliant TEI XML within the short term. We use Classical Greek as a more flexible term to include the classicizing katharevousa Greek produced through the 20th century.
[2] https://github.com/PerseusDL/canonical-greekLit/ and https://github.com/OpenGreekAndLatin/First1KGreek.
[3] CanonicalGreekLit-https://zenodo.org/records/20451365; First1KGreek https://zenodo.org/records/3827343.
[10] https://tei-c.org/.
[13] https://openiti.org/.
[14] https://capitains.org/.
[15] With data available at: https://github.com/OpenGreekAndLatin/First1KGreek.
[20] https://epidoc.stoa.org/.
Acknowledgements
After forty years of planning and development, we cannot hope to do justice to the people and institutions who have contributed along the way. Lisa Cerrato and Alison Babeu have for decades labored to develop and improve the Greek corpus. In the past ten years of work, Lenny Muellner and Gregory Nagy from the CHS, Rhea Lesage of the Harvard College Library, and Lucie Stylianopoulos of the University of Virginia Library, Matt Munson and Thomas Koentges at Leipzig have worked for years in planning and execution. Bruce Robertson of Mount Allison labored for more than a decade, tirelessly optimizing OCR systems for Classical Greek and developing a post-correction environment that allowed us to create far more open Greek than would otherwise be feasible. For more than 35 years Tufts University has provided a stable and supportive home that made our other collaborations possible.
Author Contributions
Gregory Crane: Conceptualization, Data Curation, Funding Acquisition, Investigation, Methodology, Project Administration, Resources, Software, Supervision, Writing – original draft.
Alison Babeu: Conceptualization, Data Curation, Funding Acquisition, Investigation, Methodology, Project Administration, Resources, Software, Supervision, Writing – original draft.
Lisa Cerrato: Conceptualization, Data Curation, Funding Acquisition, Investigation, Methodology, Project Administration, Resources, Software, Supervision.
Rhea Lesage: Conceptualization, Data Curation, Funding Acquisition, Investigation, Methodology, Project Administration, Resources, Software, Supervision.
Leonard Muellner: Conceptualization, Data Curation, Funding Acquisition, Investigation, Methodology, Project Administration, Resources, Software, Supervision.
Lucie Stylianopoulos: Conceptualization, Data Curation, Funding Acquisition, Investigation, Methodology, Project Administration, Resources, Software, Supervision.
