
Figure 1
Chronological distribution of the Open Greek collection. The left panel displays the total word count by century (in millions), while the right panel shows the number of individual text files. Data is categorized by source: Perseus (blue) and First1KGreek (green).

Figure 2
Distribution of genres in the Open Greek collection (Prose vs. Poetry). This visualization illustrates the disparity between the “canonical” focus of traditional Greek studies and the physical reality of the surviving record, where prose — encompassing historiography, philosophy, medicine, and technical treatises — constitutes the vast majority of the total word count.

Figure 3
Distinguishing intertextual citations from internal speech. This figure illustrates the divergent use of TEI elements: the <quote> tag is employed to sequester imported text for independent analysis, while the <q> and <said> tags delineate speech acts from the surrounding narrative voice.

Figure 4
TEI <refsDecl> element specifying hierarchical citation schemes. This declaration defines the mapping between the XML structure and the Canonical Text Services (CTS) URNs, allowing the text to be addressed and retrieved at the level of chapter or chapter-and-verse.

Figure 5
A line from Euripides’ Phoenician Women split across two speakers.
