(1) Context and motivation
(1.1) Introduction to TheSu XML and Objectives
This paper presents the first online dataset using TheSu XML, an annotation schema designed for analysing ideas and their contexts in historical sources.1 The dataset contains stand-off TheSu XML annotations of selected passages on lead white – a pigment historically produced from lead – from four Greco-Roman authors (Plato, Theophrastus, Dioscorides, and Plutarch) and two modern studies, together with the source files, authority records, and graph visualisations generated from the annotations. This dataset is intended primarily as a first reference implementation of the TheSu XML schema: a curated example showing how the schema can be applied to historical sources, how its components can be combined in practice, and how the resulting annotations can support both close reading and visual analysis. Its primary purpose is therefore methodological and demonstrative rather than documentary. At the same time, because the annotations focus on ancient and modern discussions of lead white, the dataset also offers reusable material for a diverse scholarly audience, with particular relevance for researchers working on ancient science and technology or historical pigment production.
The TheSu XML schema takes its name from its two core components, ‘thesis’ and ‘support’. In brief, a ‘thesis’ is a statement or idea conveyed by a source, while a ‘support’ is a discursive context that helps justify, explain, expand, or frame another statement or discourse component. Table 1 introduces these two elements through examples, abridged definitions, and simplified XML structures. The schema is intended to provide a method for the digital study, indexing, and mapping of ideas conveyed explicitly or implicitly by a source and of the relationships that connect them to their contexts. Tailored for research in the history of ideas, philosophy, science, and technology, it allows users to annotate individual theses with details such as thematic classifiers, formal structure, and associated speakers – features especially useful for systematic cataloguing and indexing. Building on this, the schema also enables users to model the supports that connect theses to arguments, explanations, examples, objections, or framing passages. By representing these links explicitly, TheSu XML makes it possible to map the functional interconnections within a discourse. This approach is intended to benefit scholars studying sources (enhancing their analysis), the wider academic community (producing controllable and reusable datasets), and the public (providing digital aids for understanding the annotators’ interpretations of sources).
Table 1
Introductory presentation of the ‘thesis’ and ‘support’ elements in TheSu XML.
| THESIS | SUPPORT | |
|---|---|---|
| Examples |
|
|
| Abridged definitionc | A thesis is an explicit or implicit instantiation of a proposition in a text, i.e. a minimal declarative sentence representing the stance of its speakers. | A support is a segment of text that is presented by its speakers in function of a part of the discourse that is conveyed by the same text. A support can be: argumentative; expository; expansive; contextualizing. These four functions may be cumulative. |
| Simplified XML structured | ![]() | ![]() |
[i] a The first ‘thesis’ example is visualised in Figure 1; the second in Figure 2, both are discussed in sec. 4.
b These ‘support’ examples are drawn respectively from the Plutarch and Plato visualisations discussed with Figures 1 and 2 in sec. 4.
c These definitions abridge the technical documentation for the elements THESIS and SUPPORT. For full technical definitions, readers should consult https://alchemeast.eu/thesu/ns/1.0/documentation/TheSu.html, using the table of contents on the left side of the page to focus on the elements thesu:THESIS and thesu:SUPPORT; see also Morrone (2022, pp. 6–13) for a scholarly explanation of the schema.
d For a reader-oriented guide to the main features illustrated by these simplified XML structures, see https://thesu.io/documentation. For the formal schema documentation, readers may consult the same technical documentation referenced in note c and focus on thesu:THESIS and thesu:SUPPORT through the table of contents on the left side of the page. Note that the sequence shown inside sequencesGroup is included only for illustration: ‘sequences’ are not required in the annotation of all ‘theses’, and the two ‘thesis’ examples listed in this table are annotated without ‘sequences’.
TheSu XML can be used to annotate argumentative relationships between statements, similarly to functionalities found in other argumentation annotation frameworks (major models are reviewed in Lopes Cardoso et al., 2023), but is based on different methodological assumptions and models a wider range of connections between discourse components.2 This broader scope is intended to serve the requirements of historical, philological, and epistemological research.
A key distinctive trait, relevant to the current paper, is its inclusion of features intended to support comparative analysis of ideas and discourses across different parts of the same source or multiple sources. For this purpose, the schema distinguishes between specific statements occurring within sources (‘theses’) and the abstract ideas (‘propositions’) they broadly represent.3 This feature is absent from other published argumentation annotation frameworks. For example, the thesis “Lead white is the most cooling of deadly drugs”, attested in Plutarch’s Quaestiones convivales (691b), can be interpreted to represent and instantiate abstract propositions such as “Lead white is a harmful substance” and “Lead white is a cooling substance”.4 By linking multiple statements interpreted as expressing, essentially, a common core idea under a single ‘proposition’ – even across different contexts or sources – users can digitally track how that idea varies in its details and is presented, used, or supported throughout a text or corpus.
Beyond the topic of lead white, which is of specialised interest, it is worth noting that this mechanism can support other, far-reaching research questions in the history of ideas. For example, in a variety of sources concerning the historical debate between geocentrism and heliocentrism, statements such as “the sun revolves around the earth” and “the earth revolves around the sun” could be linked to competing ‘propositions’ that represent them, thus allowing the corresponding source-level ‘theses’ to be comparatively studied in their argumentative contexts. This would make it possible to compare how different authors connect these stances to issues such as the earth’s motion, stellar parallax, and scriptural interpretation; to distinguish astronomical, theological, and philosophical uses of the same claims; and to trace whether later writers repeat inherited positions, answer specific objections, or introduce new arguments to the debate.
Both the thesis-support relationships within an individual source and the comparative connections enabled by linking theses to propositions across contexts directly result in network structures. A key advantage of modelling the annotated relationships this way is that the networks can be automatically visualised: as small maps to aid close reading, or as larger graphs for overviews, distant reading methodologies, and comparative analysis. Later in this paper (sec. 3.2 and sec. 4), I present these visualisation strategies and demonstrate their value through concrete examples, from close reading of individual arguments to comparative analysis across sources.
I based these sample visualisations on a TheSu XML annotation document created specifically for this paper. This document is intended to demonstrate the utility of applying TheSu XML to the annotation of historically attested ideas, arguments, and (chemical) processes, in order to show its potential as a standard framework for digital work on such materials.5 It showcases a wide range of TheSu XML’s capabilities, both in its structural features (diverse thesis-support relationships and proposition-based connections for comparative analysis) and in its variety of annotated content (from philosophical argumentation to the more descriptive discourse of technical treatises). Furthermore, the dataset demonstrates TheSu XML’s specific features for detailed, step-by-step analysis of processes and recipes. These also enable mapping of alternative accounts of processes or sets of instructions to each other, to identify divergences, commonalities, and changes. This is a unique functionality, particularly suited to research in the history of science and technology.6
Designed for sharing and reuse, the dataset is offered as a foundation for future research using TheSu XML. The hope is that this initial contribution will encourage further collaborative efforts (especially once the planned release of an annotator GUI will facilitate its use), leading to a growing corpus of TheSu XML-based resources. As this corpus expands, the increasing availability of such interoperable data would permit comparative and quantitative analyses of progressively wider scope, benefiting the history of ideas and related fields.
(1.2) Specific Application: the Case of psimúthion/cerussa
The dataset presented here is part of a broader project studying references to lead in Greco-Roman sources, spanning the history of philosophy and chemistry. One key area of examination concerns psimúthion (ψιμύθιον, Greek) or cerussa (Latin), today known as lead white. This substance, used primarily as a white pigment but also medically, was one of the few chemical compounds artificially produced in antiquity, through the controlled corrosion of metallic lead, and it features prominently in ancient writings (Caley & Richards, 1956, p. 190).7 Today, the expression ‘lead white’ properly refers to basic lead carbonate (also known as ‘white lead’), but the precise chemical identity of the substance produced by ancient methods and called psimúthion or cerussa remains debated.8
The sources mentioning lead white provide surprisingly fertile ground for demonstrating TheSu XML’s capabilities. Not only do they contain technical descriptions of recipes and production processes,9 which can be effectively analysed using the schema’s dedicated functionalities for modelling ‘sequences’, but also philosophical discussions and arguments in which lead white features as an example.10 Including both these types of content, the dataset is intended to showcase TheSu XML’s versatility in handling discourse structures typical of technical texts as well as complex argumentative constructions. Lead white also had significant medical, artisanal, and even socio-cultural dimensions (e.g., gender connotations), all extensively documented in the ancient sources.11 These aspects, though part of the project’s scope, are mostly not covered by the annotations presented here, which focus on lead white as a chemical product and as an argumentative example embedded in philosophical discourse.
The focus on lead white also places the dataset in dialogue with existing digital resources on artisanal and technical knowledge. For example, the ARTECHNE database, a large collection of digitised sources on artisanal techniques from the period 1500–1900, documents extensive historical material concerning pigment production and use.12 Its indexed sources include 151 records associated with the glossary term “White lead”, including many recipes and descriptions concerning the substance’s production, processing, and use. These resources are not annotated at the same fine-grained interpretative level as TheSu XML: for instance, they do not systematically distinguish recipe steps from other kinds of descriptions. They nevertheless provide the kind of digitised source base to which TheSu XML could be added as a higher-level annotation layer for comparative study. The present dataset is intentionally small, but it is readily extensible or connectable to annotations of additional corpora like ARTECHNE, thus allowing recipes, technical descriptions, and their surrounding discursive contexts to be compared across collections.
(1.3) Source corpus composition
The dataset’s small collection of TheSu XML stand-off annotations, contained entirely within the file ‘ancient-lead-white_JOHD.xml’, covers exemplary passages concerning lead white. See Table 2 for details on each source. Primarily, the annotated passages derive from a subset of Ancient Greek texts chosen to effectively illustrate TheSu XML’s core functionalities. However, to further demonstrate the schema’s versatility beyond ancient materials, the collection also includes two modern sources: the commentary section of a book edition of an ancient work (Caley & Richards, 1956) and an academic article (Katsaros et al., 2010). These were selected specifically because they discuss replication attempts of ancient recipes for lead white (particularly Theophrastus’s), which makes them suitable for a demonstration of TheSu XML’s comparative mapping capabilities. Following publisher guidelines regarding open access materials,13 the Caley & Richards source was treated as copyrighted. Hence, a method was developed for publishing TheSu XML annotations that respect copyright restrictions through fair use, avoiding substantial reproduction of the original text.14
Table 2
Sources included in the dataset.
| AUTHOR | WORK TITLE | PASSAGE THEME/TYPE | LICENCE | REFERENCE | DOCUMENT IN THE DATASET | DIGITAL SOURCE |
|---|---|---|---|---|---|---|
| Plato | Lysis | Philosophical | CC BY-SA 4.0 | (Burnet, 1903) | tlg0059.tlg020.perseus-grc2.xml | PerseusDL/canonical-greekLit |
| Theophrastus | De lapidibus | Chemical, Procedural | CC BY-SA 4.0 | (Wimmer, 1862) | tlg0093.tlg004.1st1K-grc1.xml | OpenGreekAndLatin/First1KGreek |
| Dioscorides | De materia medica | Medical, Chemical, Procedural | CC BY-SA 4.0 | (Wellmann, 1906) | tlg0656.tlg001.1st1K-grc1.xml | OpenGreekAndLatin/First1KGreek |
| Plutarch | Quaestiones convivales | Philosophical, Chemical, Procedural, Medical | CC BY-SA 4.0 | (Bernardakis, 1892) | tlg0007.tlg112.perseus-grc2.xml | PerseusDL/canonical-greekLit |
| Caley, E. R., & Richards, J. F. C. | Theophrastus. On stones (Commentary section) | Chemical, Procedural | Copyrighted | (Caley & Richards, 1956) | today_caley,richards_1956_copyrighted.xhtml | OCR PDF scan → XHTML |
| Katsaros, T., Liritzis, I., & Laskaris, N. | Identification (Article) | Chemical, Procedural | CC BY-NC-ND 4.0 | (Katsaros et al., 2010) | today_katsaros,liritzis,laskaris_2010.xhtml | Publisher PDF → XHTML |
(2) Dataset Description
Repository location https://doi.org/10.5281/zenodo.20290847
Repository name Zenodo
Object name lead-white_JOHD
Format names and versions XML, XSD, CSS, XHTML, DOT, GEPHI, SVG, PDF, PNG
Creation dates 2025-02-21 – 2025-06-25; minor corrections 2026-05-19
Dataset creators Daniele Morrone
Language English, Ancient Greek
License CC-BY-SA-4.0
Publication date 2026-04-16
(3) Method
(3.1) Annotation methodology and preparation of digital editions
I produced the annotations presented in this dataset through close reading of the selected source passages. I employed TheSu XML to encode ideas concerning lead white along with their immediate discursive contexts (micro-contexts). In practice, this means that if a statement about lead white serves a specific point, that point and its connection with the statement are annotated. I generally excluded further points derived from such connections, or broader framing contexts (macro-contexts) like section introductions, unless essential for understanding the statements’ immediate purposes and reasons.
TheSu XML offers a versatile toolbox of elements and attributes adaptable to varying research needs: annotators can be selective or exhaustive with these components depending on their projects’ goals. For this annotation, I chose to utilise a broad selection to showcase the schema’s potential. Primarily, this involved identifying and encoding all relevant declarative sentences (‘theses’) explicitly or implicitly conveyed by the source passages, starting from those specifically concerning lead white. These ‘thesis’ annotations were then enriched with further details directly supported by the text, including indexing information like associated speakers, themes, and keywords,15 as well as notes on formal structure like the presence of embedded aetiological explanations or analogies.16
After identifying the ‘theses’, the analysis focused on their relationships with the ‘supports’ that employ or target them. Annotating these connections to reconstruct the argumentative and expository discourse is an inherently interpretative task. This is consistent with TheSu XML’s design principles: rather than simply mapping surface-level syntax, the goal is to discern and encode the most reasonable argumentative and rhetorical links implied by the textual evidence – including implicitly conveyed assumptions, intermediate premises, conclusions, and other alluded meanings.17 This arguably mirrors the interpretative effort involved in crafting a scholarly translation or explanatory commentary in the fields of history and philology: the aim is generally to remain fair to the source while acknowledging the mediation (and subjectivity) of the interpreter, welcoming divergences in interpretation. Accordingly, it should be kept in mind that this dataset only presents one possible interpretation of the sources, formalised in a machine-readable format, rather than an absolute, objective, representation of their content.
Still, the principle of charity that I deployed in my analysis prioritises fidelity to authorial language, context, and discernible intent, also at the cost of reconstructing potentially fallacious reasoning.18 Different annotators using TheSu XML are free to make different choices based on varying principles or judgments. Given this variability, it should be considered crucial to track provenance in the annotations, e.g., through author and timestamp metadata associated with their core components. This feature is not yet implemented in the current version of the schema – and is accordingly not represented in the dataset – but will be incorporated in the future.19 For the moment, readers may simply bear in mind that all analyses and visualisations presented in this dataset are a product of my own interpretative reading of the sources, consistent with the criteria outlined above.
Finally, each recipe for lead white was modelled as a single complex ‘thesis’ containing one or more ‘sequences’. Each ‘sequence’ breaks down the process described in the thesis into constituent ‘phases’, annotated with individual paraphrases and details such as ingredients, duration, and required repetitions. The schema provides features for mapping ‘variant’ recipes to their corresponding ‘primary’ process (whether explicitly designated as primary or merely presented earlier in the source), through detailed step-by-step encoding of commonalities and divergences. These features were fully utilised in the XML annotation file (‘ancient-lead-white_JOHD.xml’), but the detailed phase-to-phase mappings they encode are not yet supported by the current visualisation script. Therefore, though variant recipes do appear as distinct ‘thesis’ nodes in the graph files and diagrams (particularly Dioscorides’s in Figures 3 and 7 below, sec. 4), their internal steps and specific links to the primary recipe ‘sequence’ are visible only by inspecting the base XML file.20
The annotation relies on separate authority record files, referenced within the base XML document, for indexing tags, people, types of ‘supports’, bibliography, and other details.21 This practice is meant to promote controllability, standardisation, and reuse. A goal for future schema versions, in order to better align with FAIR data principles (Wilkinson et al., 2016), is to incorporate even further metadata – including, for instance, detailed provenance for annotations (as mentioned above) and additional linked data identifiers for sources (using standards like URNs, especially valuable for canonical Greek and Latin texts, following the precedent of the Perseus Digital Library).22 Another planned development is the establishment of an automated conversion pipeline of compatible TheSu XML components into the Argument Interchange Format (AIF, see Argumentation Research Group, 2011), which is presently the leading encoding format for the digital annotation of arguments.23
TheSu XML employs stand-off annotation. Therefore, for the purposes of this dataset, source texts were segmented into individually identifiable XML elements (words, punctuation, numbers, whitespaces), each provided with its own ‘xml:ID’ attribute. In the annotation document, components like ‘theses’ and ‘supports’ reference text spans in the associated sources by pointing to these ID-bearing elements through dedicated descendant ‘segment’ elements. These ‘segment’ elements, specifically, use ‘xlink:href’ attributes to point to the IDs of the elements beginning and ending the text spans. The main advantages of this method are that it supports references to discontinuous text segments (through inclusion of multiple sibling ‘segment’ elements within the same ‘textRef’ parent)24 and that it circumvents the potential issue of overlapping hierarchies.25 It also preserves the source file’s structure and semantic markup, which is especially helpful in history and philology for precise referencing (e.g., allowing ‘div’ sections and TEI ‘milestones’ to be automatically retrievable for any annotated passage, which would be impossible if the source were annotated as plain text).26
Though advantageous, stand-off annotation is currently laborious, especially given the absence of a dedicated graphical user interface (GUI). The present dataset was created within Oxygen XML’s ‘Author mode’ view, with the TheSu XML annotation and source files open side-by-side (using tailored rudimental CSS for the visualisation of the former); each annotation component (like ‘theses’ and ‘supports’) was created individually and then linked to the relevant text segments as specified above. Relying largely on manual work – assisted by the editor’s limited graphical features and some custom automation scripts – this text referencing process is inevitably time-consuming and lacking immediate visual confirmation of correctness. This adds to the well-known challenges of detailed annotation labour, which is demanding even without employing manual stand-off methods.27 Despite this, the utility of the stand-off approach remains clear. An additional advantage is that it allows TheSu XML to apply potentially beyond text, to any digital medium where components can be assigned IDs (images, audio segments, audiovisual clips, etc.) – anything conveying information that is modellable as declarative sentences.28
To maintain the flexibility of stand-off annotation while addressing its workflow challenges, developing a dedicated GUI is crucial. Planned for release by 2027 is a first version of a stand-alone, offline, portable application, packaged for easy execution on the main operating systems, that will simplify and accelerate the process by enabling the user to generate TheSu XML annotation components via direct text highlighting, contextual menus, and hotkeys. The envisioned design will include on-the-go visualisation generation features and automation aids (such as keyword suggestions based on string similarity, or online GenAI-based recommendations for paraphrases and argumentative links –29 to be used with caveats, to ensure these suggestions do not excessively sway the annotator’s expert judgment). Since a specific GUI for displaying and editing TheSu XML annotations is not yet available, the CSS stylesheet used while creating the base document ‘ancient-lead-white_JOHD.xml’ is also included in this dataset. It requires Oxygen XML’s ‘Author mode’ for correct rendering, as it utilises editor-specific functions. The XML structure itself is viewable in standard text editors or web browsers.
As a consequence of the stand-off methodology employed, the published dataset includes not only the primary annotation file (‘ancient-lead-white_JOHD.xml’) but also the segmented source documents to which it links. For sources available under open licences, the original, pre-segmentation digital editions are also provided for completeness. However, to respect copyright restrictions, the original text file for the Caley & Richards (1956) commentary is not included. Instead, only the segmented version required for the stand-off annotation links is provided. This version was processed using a script that replaced the majority of words with placeholder character sequences of equivalent length; only the initial and final words of each text span referenced by an annotation component were left unobscured. This method allows users to verify the annotation references against an independently accessed copy of the work while avoiding substantial reproduction of the copyrighted text, thus complying with fair use principles. To enable precise referencing and facilitate such independent verification, structural metadata within this processed digital edition were not obscured, including page and paragraph markers.30 A full TEI-encoded citation of the work is also available, in the bibliography authority record file associated with the TheSu XML document.
(3.2) Conversion of TheSu XML annotations into discourse graphs
A Python script was created to visualise the discourse structures encoded in TheSu XML.31 As previously noted (sec. 1.1), both standard thesis-support relationships and proposition-based comparative links naturally form networks that can be visualised. This script uses these connections to automatically convert select linked annotation components (‘theses’, ‘supports’, etc.) into graph formats. The script’s design is inspired by argumentation mapping principles, inasmuch as it relies on established techniques to visually represent claims and their inferential connections.32 An original feature is the visualisation of individual ‘phases’ within procedural ‘theses’ (e.g., recipes), integrated with the surrounding discourse map. Besides helping users understand procedural ‘theses’ within their discursive contexts, this design allows to display very specific relationships within the maps, such as an argument supporting a single step instead of the overall recipe in which this is included.33
Argumentation mapping literature describes various tools and designs.34 Scalability, which is a known problem, likely affects all: visualising the discourse of extensive texts in an effective and easily readable way is much harder than mapping single arguments or short passages. Published examples – including those in the present paper – typically focus on such simpler cases, while attempts to map longer or more complex discourses easily result in visualisations that are difficult to read as the diagrams become overcrowded with elements and crossing lines, and the automated arrangement algorithms struggling with their complexity.35 Even detailed maps of short passages can look remarkably cluttered – for instance, if they are exhaustive in their inclusion of implicit premises and conclusions.36
These challenges suggest that effective visualisation tools should follow several key requirements. First, these tools should offer strong user control through interactivity, providing options to filter nodes and edges, modify layouts, and apply decluttering methods.37 While static, non-interactive visualisations may be suitable for illustration or teaching if carefully designed, the primary analytical uses of TheSu XML annotations (supporting research, enhancing text comprehension), also considering their variety,38 will certainly require interactive views. Second, the interfaces presenting such visualisations should be optimised for displaying them alongside the source texts and annotations they represent; for example, providing interactive linking, so that selecting an element in the visualisation immediately highlights the corresponding source text segment or annotation component in the TheSu XML full document.39 Such linking, ideally working in both directions (from text to graph and vice versa), would significantly speed up the process of relating the visualisation to the source passages and enhance the user’s control over the analysis. Finally, a dual visualisation strategy, allowing users to switch between two different map designs, one more detailed, the other for larger scale overviews, seems an ideal solution to the scalability problem. The detailed design, more closely aligned with argumentation mapping trends, would be used with short passages to facilitate close reading; the other design, presenting network graphs as commonly seen in network analysis research, would work better for both passages presenting complex discourse structures and larger annotations, potentially extending to the discourse of an entire work or corpus. Enabling both kinds of visualisation within the same interface would allow for smooth movement between close and distant reading methods, making it easy for users to verify their distant reading intuitions through careful examination of the individual network nodes and their micro-contexts.40
Full implementation of the interactivity and integrated linking features described above is a central objective, also tied to the planned graphical user interface (see above, sec. 3.1). The current Python visualisation script, though producing static and unlinked graphs, already implements the dual visualisation strategy. It involves two main conversion processes. The first converts data from the TheSu XML annotation and its associated sources into the DOT language for Graphviz processing: its outcomes are a DOT document containing a list of nodes, clusters, and edges representing a structured ‘argumentation map’, and SVG, PDF, or PNG files produced by Graphviz to display that map (following custom parameters). Such maps usually work better when arranged as hierarchical tree structures (through the ‘dot’ engine), to effectively show argumentative and expository dependencies (e.g., premises appear above conclusions), and they display significant details for each node, including full paraphrases for the ‘theses’. The second process thoroughly revises the DOT document produced by the first to ensure it correctly represents a ‘network graph’ and optimises it for import into the network analysis software Gephi.41 This results in a more abstract network representation that shows less information directly on the nodes, requiring user interaction within the Gephi environment for full understanding. Gephi provides customisable layout algorithms (like ‘ForceAtlas2’ and ‘Fruchterman Reingold’, used for this dataset’s graphs) that can produce much clearer and more controllable arrangements for complex networks than Graphviz’s built-in engines. As shown in the next section, this also makes it suitable for visually comparing different discourses linked via ‘proposition’ nodes, even across different sources.42 All the files resulting from the conversion processes mentioned above are included in the dataset, along with the Gephi project files created after DOT import and application of layout algorithms.
It should be noted that the visualisation workflows presented here use only a selection of TheSu XML features. Annotation documents encoded using TheSu XML are – or can be – much richer in details, based on the wide variety of elements and attributes defined in the XML Schema Definition. Developers are therefore invited to create custom visualisation scripts for graphical representations alternative to the one presented here, or visualise other aspects of the encoded information, according to specific research questions and interests.
(4) Results and Discussion
This section showcases some of the possibilities afforded by TheSu XML annotations in terms of discourse structure visualisation. As previously mentioned, the machine-readable nature of TheSu XML annotated discourse not only allows for such automatic visualisation of functional relationships between discourse components, but also enables direct quantitative analyses. In principle, such analyses could, for instance, include comparisons of tendencies in argumentation patterns across different sources or within the same text, in the same vein of Wang et al. (2022). I have already illustrated in previous work some of these possibilities, through sample analyses of an integral annotation of Plutarch’s Aquane an ignis utilior sit.43 The size of the annotated corpus in the current dataset is arguably insufficient for revealing interesting patterns through quantitative analysis. Nevertheless, it is sufficient for generating visualisations that demonstrate the schema’s capabilities, limitations, and potential applications to larger corpora in future research.
Let us start with case studies of argumentation and exposition visualisation. Figure 1, a detail from the Graphviz visualisation in directory ‘1.Plutarch’, presents a minimal example from Plutarch’s Quaestiones convivales (6.5, 690f-691c). It shows a single argumentative link while demonstrating TheSu XML’s capacity to encode additional kinds of relationships between ‘thesis’ statements, including one defined as ‘entailment’.44 The visualisation displays how Plutarch’s statement that “lead white is the most cooling of deadly drugs” is ‘included’ within (hence ‘entailed’ by) a broader procedural observation about lead producing lead white when rubbed with vinegar. This serves as evidence for the claim that “lead is among the naturally cold substances”. To make the logical progression more obvious, I annotated an ‘implicit’ mediating premise – that “lead can produce the most cooling of deadly drugs” – marking it as ‘entailed’ by the procedural thesis.45

Figure 1
Detail of the image ‘1.graphviz_Plutarch.png’, generated through Graphviz (‘dot’ engine) from the DOT file ‘1.graphviz_Plutarch.dot’. Visualisation of the TheSu XML annotation for the passage in Plutarch, Quaestiones convivales 6.5, 690f–691c.
The hierarchical tree arrangement, generated through Graphviz’s ‘dot’ engine, displays extensive information: full paraphrases, speaker attribution, passage references, and original text snippets, with ‘implicit’ components visually distinguished from the ‘explicit’.46 The amount of information displayed produces graphs ill-suited to small views like print, certainly requiring digital navigation with scrolling and zooming. Nonetheless, after learning the graphical conventions, readers can readily identify through these maps how each thesis functions within its immediate discoursive context. In the dataset, the same relationships in Plutarch’s passage are also represented in a Gephi network graph (using the ‘ForceAtlas2’ layout). This version has the advantage of showing the entire discourse structure in a single view, making relationships immediately evident. However, nodes display only their element type, requiring mouse interaction to reveal full information. This is already possible in Gephi, but its interface has limitations. Better solutions will be implemented in the planned TheSu XML annotation interface.
Figure 2 presents the Gephi visualisation in directory ‘2.Plato’, displaying a more complex discourse structure from Plato’s Lysis (217b-e). In this passage, Socrates, in a dialogue with Menexenus, reflects on the nature of qualities and their presence in objects.47 Socrates advances examples involving the application of lead white to Menexenus’s blond hair, repeating some of the points for emphasis and alternating between explicit statements and rhetorical questions (conveying ‘implicit’ theses). The lead white examples serve a dual purpose: they first function as ‘expository supports’ for Menexenus, clarifying Socrates’s unintuitive claim that a dyed object’s colour differs from that of the dye present to it; then, they become argumentative premises supporting Socrates’s ultimate moral point, concerning the relationship between what-is-neither-good-nor-bad and badness, when this is present in the former.48 The resulting complexity of the discourse makes it difficult to read it in the form of a Graphviz ‘dot’ hierarchical tree (included in the dataset), also considering that many of the statements are presented as contrasting with each other – hence annotated with ‘contextualising supports’ – and that Menexenus occasionally responds with confirmations – annotated, similarly to argumentative premises, as ‘justifying supports’. As with Plutarch’s example, the Gephi version, without manually added captions like those in the figure, requires digital interaction to reveal node and relationship information.

Figure 2
Image ‘2.gephi_Plato.png’, generated through Gephi (‘ForceAtlas2’ layout) from the DOT file ‘2.gephi_Plato.dot’, with manually added captions, box, and arrows. Visualisation of the TheSu XML annotation for the passage in Plato, Lysis, 217b–e.
Moving from small-scale discourse analysis to larger-scale network representations, Figure 3, from the Gephi graph in directory ‘3.Plato-Plutarch-Dioscorides’, demonstrates how TheSu XML’s ‘proposition’ element can be used to compare discourses across sources. The graph connects the passages from Plutarch and Plato discussed above with selected sections from Dioscorides’s De materia medica (5.81, 5.82, 5.88) on lead white’s properties and production.49 Perhaps surprisingly, despite Plutarch’s Platonism, his discussion shares no common propositions with Plato’s passage. Instead, both independently connect with Dioscorides’s treatment: Plutarch’s chemical-pharmacological mentions of lead white as a “deadly”, “cooling” drug and its production process find parallels in Dioscorides’s detailed account,50 while Plato’s focus on whiteness connects with Dioscorides’s remark that properly produced lead white “comes about white and effective”.51 This visualisation approach can be expected to scale effectively to larger corpora, reliably revealing thematic clusters: chemical-pharmacological discussions would group near Plutarch’s, chromatic (and metaphysical) ones near Plato’s, and encyclopaedic accounts like Dioscorides’s would mediate between them. In the dataset, I also included a Graphviz ‘dot’ hierarchical version for comparison, only to show that it is unsuited to such cross-source analyses.

Figure 3
Image ‘3.gephi_Plato-Plutarch-Dioscorides.png’, generated through Gephi (‘ForceAtlas2’ layout) from the DOT file ‘3.gephi_Plato-Plutarch-Dioscorides.dot’, with manually added captions and boxes. Comparative visualisation of the TheSu XML annotations for the passages in Plutarch, Quaestiones convivales 6.5, 690f–691c, Dioscorides, De materia medica 5.81, 5.82, 5.88, and Plato, Lysis, 217b–e.
Now, let us examine how TheSu XML handles the annotation and visualisation of recipes and their comparative analysis, focusing on the production processes for lead white described in our sources. Figure 4, from the Graphviz visualisation in directory ‘4.Theophrastus’, demonstrates this through Theophrastus’s recipe in De lapidibus 55 – the earliest attested production process for lead white. The ‘sequence’ within the thesis appears broken down into its constituent ‘phases’, each with a full paraphrase, numbered and grouped in a cluster connected to the parent ‘thesis’ node. The integration of this view within the argumentation map design ensures that the recipe’s structure is revealed together with its discursive context. For instance, the thesis at hand also includes the statement that lead “acquires thickness in a maximum of 10 days”. Since this is not a procedural step, but a descriptive physical statement, it cannot be annotated as a ‘phase’; it can, however, be annotated as an ‘included thesis’ connected through an ‘expository support’ to the step “they wait for lead to acquire thickness”, as in the figure. I have also added bibliographic references to step 6 to document previous interpretations assuming that it involves the use of water;52 these references are currently not displayable in the graph. The advantages and disadvantages of visualising recipes via Graphviz or Gephi are the same as those discussed for Plutarch’s example of analysed argumentation (Figure 1). As regards the Gephi version of the graph, a peculiarity of Gephi-visualised recipes is that phase-to-phase connections are removed. In network graphs, edges should always represent meaningful relationships, rather than serving ordering purposes. Therefore, in this design, each phase links instead directly to its pertaining ‘thesis’ parent node.53

Figure 4
Detail of the image ‘4.graphviz_Theophrastus.png’, generated through Graphviz (‘dot’ engine) from the DOT file ‘4.graphviz_Theophrastus.dot’. Visualisation of the TheSu XML annotation for the passage in Theophrastus, De lapidibus 55–56.
We can now move to the comparative visualisation of recipes. In Graphviz, the spring-model engine ‘neato’ proved more effective for this purpose than the hierarchical ‘dot’ engine. In Figure 5, from directory ‘5.Diosc.recipe-Thphr.recipe’, the first recipe for lead white presented by Dioscorides (De materia medica 5.88.1) is visually compared with a recipe encoded as a ‘proposition’. This proposition models Theophrastus’s variant procedure for lead white. Arrows connect corresponding steps, with labels characterising their relationships, important for cases in which the correspondences are only partial (see in the figure, e.g., ‘extends’, ‘alters’). Non-corresponding steps appear spatially distant and dimmer-coloured, making divergences immediately apparent. Here too, the Gephi version included in the dataset sacrifices content readability, removing paraphrases, but displays as labels each step’s number in the sequence, for reference. It uses the ‘Fruchterman Reingold’ layout, which for these comparative purposes appears to be more effective than ‘ForceAtlas2’.

Figure 5
Detail of the image ‘5.graphviz_Diosc.recipe-Thphr.recipe.png’, generated through Graphviz (‘neato’ engine) from the DOT file ‘5.graphviz_Diosc.recipe-Thphr.recipe.dot’. Visualised comparison between the TheSu XML annotations of the first recipe for lead white in Dioscorides, De materia medica 5.88.1 (in green) and a procedure modelled from Theophrastus’s recipe in De lapidibus 55 (in purple).
The same Graphviz ‘neato’ design also enables effective three-way comparisons. In directory ‘6a.KLL.recipe-Thphr.recipe-Diosc.recipe’, the modern replication of Theophrastus’s process by Katsaros et al. (2010, pp. 2–3) is visually mapped to both ancient recipes at the same time. As shown in a separate argumentation-graph visualisation (directory ‘6b.KLL.recipe’), the replication’s authors explicitly state in a previous page of their paper that they combined Dioscorides’s recipe with Theophrastus’s “incomplete” description to achieve “a complete method”.54 However, they do not precisely specify which steps derive from which source. The comparative diagram, through spatial distancing and colour coding, makes these attributions clear. For instance, their step 5, “We placed the vessel under the sun”, appears in neither ancient source directly, exemplifying an original contribution on the researchers’ part. However, their step 8, “We took the […] lead out”, though not matched to either ancient text, is clearly implied by both procedures, reminding us that these comparisons concern similarities and differences in textual presentation rather than in the procedures underlying them. A similar comparative visualisation of Caley and Richards’s (1956, p. 188) replication (in directory ‘6c.Caley,Richards.recipe-Thphr.recipe-Diosc.recipe’) reveals significant departures from Theophrastus’s recipe in the final part, primarily related to their interpretation that Theophrastus’s step 6 involves water.55 It should be kept in mind that these visualisations simply highlight differences for analytical, not directly evaluative, purposes: each divergence should be assessed individually for its methodological and procedural value, and even a procedure diverging entirely from its source may constitute a valid test of its chemical rationale. Principe’s method of modelling recipes to isolate specific variables, which he applied to Theophrastus’s recipe for lead white, provides a good example (Principe, 2018, p. 160, n. 4).
Figure 6 (from directory ‘7.all.recipes’) demonstrates how Gephi visualisations prove more effective than Graphviz visualisations for comparing more than three recipes simultaneously. The network graph compares all recipes discussed above – including Plutarch’s compressed two-steps description – against the model recipes from Theophrastus and Dioscorides, with spatial closeness providing an approximated indication of shared steps.56 According to my tests, the ‘Fruchterman Reingold’ layout better represents expected relationships between recipes by positioning similar steps closer together than ‘ForceAtlas2’, but it is important to note that, even in this layout, some shared steps do not appear as centrally positioned between related recipes as one may expect.57 Despite this limitation, the interactive nature of Gephi enables effective verification of visual insights, whereas Graphviz is proven to be inadequate for such large-scale comparisons (a ‘neato’ visualisation, included in the dataset for comparison, evidently produces an overly cluttered diagram, with the ‘dot’ engine performing even worse).

Figure 6
Image ‘7.gephi_all.recipes.png’, generated through Gephi (‘Fruchterman Reingold’ layout) from the DOT file ‘7.gephi_all.recipes.dot’, with manually added captions and boxes. Visualised comparison between the TheSu XML annotations of (in green) lead white recipes from Dioscorides, De materia medica 5.88.1, Theophrastus, De lapidibus 55, Plutarch, Quaestiones convivales 6.5, 691b, the replications by Caley & Richards (1956) and Katsaros et al. (2010), and (in purple) procedures modelled from Dioscorides’s first recipe and Theophrastus’s.
Finally, Figure 7 (from directory ‘8.everything’) presents the most comprehensive Gephi visualisation in the dataset, encompassing all annotated discourse structures, including sequential theses, their phases, and connections to propositions. Similar to Figure 3 but at a larger scale, the graph distributes sources in clusters around proposition nodes they connect to, while also showing which ‘thesis’ recipe steps correspond to ‘propositional’ steps in the model recipes. The predominance of technical-chemical sources over philosophical ones (represented only by Plato) reflects a deliberate choice: while the dataset is also meant to demonstrate TheSu XML’s effectiveness for annotating philosophical discourse, one of its core purposes was to showcase the schema’s capabilities in analysing recipes across ancient and modern sources. A broader corpus of ancient texts, including those discussing medical applications or critiquing lead white’s use as a cosmetic, would reveal more interesting patterns in how ideas cluster and spread across sources, making it easier to trace intellectual traditions and tendencies.

Figure 7
Image ‘8.gephi_everything.png’, generated through Gephi (‘Fruchterman Reingold’ layout) from the DOT file ‘8.gephi_everything.dot’, with manually added captions and boxes. Comparative visualisation of the TheSu XML annotations for the passages in Plato, Lysis, 217b–e, Dioscorides, De materia medica 5.81, 5.82, 5.88, Plutarch, Quaestiones convivales 6.5, 690f–691c, Theophrastus, De lapidibus, 55–56, Caley & Richards (1956), and Katsaros et al. (2010). All primary recipes and replications annotated for the sources are visually compared, with step-by-step mapping, to the two procedures modelled from Dioscorides’s first recipe and Theophrastus’s.
Any such visual insights, of course, will still require verification through close reading. The visualisations presented in this section serve primarily as heuristic tools for interactive exploration; not as definitive representations. Force-directed algorithms like ‘ForceAtlas2’ and ‘Fruchterman Reingold’ generate procedural approximations that prioritise readability over objective spatial relationships.58 These visualisations require user interaction: readers should engage with them critically, using them as interactive maps for discovering patterns to be individually verified.59 Gephi already offers basic interaction with node attributes for examining, for instance, ‘thesis’ content and metadata.60 However, as mentioned above, a dedicated TheSu XML interface would greatly facilitate quick access to this data, as well as to annotated contexts and to linked source segments.
For this dataset specifically, a primary goal was to verify that the visualisation method faithfully represents the relationships encoded in TheSu XML annotations, establishing standards for reliability assessment before scaling to larger corpora. Spatial relationships in these visualisations warrant particular caution: visible proximity between nodes does not necessarily indicate thematic similarity.61 Future applications of these network representations may also enable quantitative analyses based on absolute metrics instead of approximations, proceeding from the graphs’ underlying data structure rather than their visual rendering for human readers.
(5) Implications/Applications
The application of the TheSu XML schema, in this paper, to select sources mentioning lead white demonstrates its value for research in the history of ideas, philosophy, science, and technology. Users of TheSu XML, thanks to the schema’s systematic encoding of ideas and their discursive contexts, can transform their interpretations of sources into structured datasets suitable for computational analysis, automated visualisation, and reuse by other researchers. The schema supports rigorous textual analysis by allowing users to annotate implicit argumentative elements, various functional relationships between statements, and phase-by-phase details of descriptions of processes. The annotation framework even accommodates scholarly uncertainty, having dedicated elements for speculative interpretations and conflicting views. Furthermore, its visualisation tools can already aid individual and comparative understanding of complex textual structures, at both micro and macro levels.
The analytical encoding of procedures, broken down into their individual steps, each with its own details, also serves experimental archaeology and the history of science. For instance, comparative visualisations of historical procedures, later technical descriptions, and published replication attempts can aid the assessment of their similarities and differences, potentially also inspiring new laboratory replications. Although demonstrated here on a small corpus centred on lead white, this approach is extensible to larger collections of technical recipes and artisanal knowledge.
The machine-readable nature of TheSu XML annotations enables quantitative studies of how ideas and arguments develop within and across textual sources, provided a large and representative enough corpus is annotated.62 Several planned improvements are intended to encourage wider adoption of the schema and the production of better and larger annotated corpora. For clarity, these planned improvements may be divided into high-priority developments, expected by the end of 2027 or sooner, and lower-priority desiderata for later work. The first group includes: generation of GEXF files instead of DOT for the ‘network graph’ visualisations, improving compatibility with Gephi and allowing the inclusion of richer graph data such as edge weights;63 support in the visualisation script for displaying more detailed phase-to-phase mappings of variant recipes;64 a dedicated user-friendly, open-source GUI for producing, examining, and visualising TheSu XML annotations, with intuitive markup generation through direct text highlighting, contextual menus, and hotkeys, interactive visualisation, graph filtering, source-text/annotation linking, and optimised access to node and edge data;65 completion and refinement of the element tree for ‘sequences’;66 richer metadata for FAIR-oriented reuse, including annotation provenance and linked-data identifiers for sources;67 and further improvements to the lead-white dataset itself, including fuller definitions in the authority records and possible extensions of the annotated corpus.68 Lower-priority desiderata, to be addressed after 2027, include: revising the schema’s definitions so that they explicitly encompass the annotation of non-textual media and extending the schema with dedicated elements optimised for such media;69 designing a web platform for presenting and comparing annotations of various provenance while displaying richer contextual and bibliographic information;70 and establishing an automated conversion pipeline from compatible TheSu XML components into other encodings such as the Argument Interchange Format (AIF).71
The present dataset has several limitations affecting its immediate reuse and scalability. The corpus is intentionally small and selective, and is meant primarily as a reference implementation rather than as a representative basis for quantitative conclusions. Moreover, the annotations reflect a single researcher’s close reading of the sources. This reflects a methodological choice in the present TheSu XML pipeline, whose purpose is to support close reading and make specific interpretative reconstructions explicit, inspectable, and reusable. For this reason, inter-annotator agreement is not currently treated as an implementation objective in the same way as it would be in annotation projects designed to produce consensus labels. In this respect, TheSu XML annotations are comparable to expert translations or interpretative commentaries of historical texts, where scholarly disagreement is expected and where validation depends partly on the interpreters’ explicit positioning. Accordingly, users are encouraged to document their non-trivial interpretative choices through the schema’s features for bibliographic referencing and evidence-based justification.72 Future projects wishing to fork and extend the system, especially after the release of the planned Annotator GUI, may of course complement this approach with inter-annotator agreement features; such developments would be welcomed and supported by the author.
Readers wishing to inspect, reproduce, or extend the present dataset can begin with the following workflows. To inspect the existing TheSu XML annotation, they should download the dataset without altering its directory structure and open ‘ancient-lead-white_JOHD.xml’. If Oxygen XML Editor is installed, the associated CSS stylesheet will be loaded correctly in ‘Author’ mode, simplifying inspection and modification of the most important elements and attributes. To reproduce the visualisations presented in this paper, or to change their parameters, readers may instead download the standalone TheSu- repository released as Morrone (2026), which includes a copy of the same annotation file and provides detailed reproduction instructions in ‘JOHD_dataset_reproduction.md’. As explained in that tool’s ‘README’, this workflow requires Python, Graphviz, and several Python dependencies. Finally, readers wishing to create a new TheSu XML annotation may reuse the CSS included in the dataset, since it provides some automation in Oxygen XML Editor, including contextual menus and automatic insertion of required elements and attributes. This route still presupposes some familiarity with the XML encoding. Schema-import instructions are provided at the beginning of the formal technical documentation, which also supplies human-readable definitions and usage guidance for individual elements and attributes.73 Readers without experience in XML are advised to wait for the planned user-friendly GUI, which is intended to require no prior experience with digital encodings or programming.
The initial dataset presented here serves primarily as a reference implementation of the TheSu XML schema, demonstrating through concrete examples how its annotations can be created, reused, and converted into Graphviz and Gephi visualisations of the relationships between ‘theses’, ‘supports’, and ‘propositions’. At the same time, its focus on lead white provides a small but reusable basis for future annotation projects on ancient science, technology, and historical pigment production. Supported by its methods for annotating copyrighted materials and its commitment to FAIR data principles, this work sets the stage for the next steps in TheSu XML development and paves the way for broader implementation across academic disciplines.
Notes
[2] For an accessible introduction to TheSu XML, see https://thesu.io, especially the ‘About’ and ‘Documentation’ sections, which provide an extensive presentation of the schema including examples and annotation snippets. For a more detailed, scholarly presentation of TheSu XML’s conceptual framework and purposes see Morrone (2022, pp. 6–13), though describing a previous, more limited, version of the schema. The technical documentation of the latest version of the schema is accessible online at https://alchemeast.eu/thesu/ns/1.0/documentation/TheSu.html. The latest version of the XML Schema Definition file is included in the current dataset and mirrored on the KU Leuven RDR repository (Morrone 2023).
[3] See Morrone (2022, pp. 6–8) for further bibliographic references and elaboration on TheSu XML’s positioning with respect to comparable annotation initiatives. TheSu XML shares some analytical concerns with such frameworks but differs in its methodological assumptions, in the range of discourse connections it models (for instance, ‘supports’ model not only argumentative premises, but also explanatory examples, introductory statements, and other types of discourse components instrumental to others), and in its orientation toward interpreter-led historical inquiry rather than pragmatic-linguistic analysis. For an example of a crucial distinctive feature, see the following paragraph. See also below, n. 17.
[4] Abridging the technical documentation (see Table 1, note c): a ‘proposition’ is the abstract, semantic content of a declarative sentence, i.e. the uninstantiated counterpart of a ‘thesis’ that is recognised by an annotator as being synonymous with it or very similar in meaning. In many cases, the relationship between a ‘proposition’ and a ‘thesis’ may parallel the type/token distinction familiar to philosophers of language and logicians; the schema does not, however, require strict type/token correspondence, so as to allow ‘propositional’ links based on humans’ subjective recognition of any partial similarity. This allows the schema to support more efficiently the interests of comparative research. For instance, the thesis “Lead white is the most cooling of deadly drugs” is not strictly a token of the type “Lead white is a harmful substance”, yet it is pertinent to the study of references to lead white as a harmful substance. Annotators can also qualify such links to signal divergences in meaning between a thesis and the proposition it is connected to. In the example at hand, the ‘thesis”s link to “Lead white is a harmful substance” is marked in the annotation file as both ‘generalised’ and ‘partial’ – ‘generalised’ because the ‘proposition”s predicate is more general than “is the most cooling of deadly drugs” while still encompassing its meaning; ‘partial’ because it omits the information on coolness. Each available qualifier for a ‘matching proposition’ references is defined in the technical documentation. For further examples of such qualifiers for comparative links, see the edge labels in Figure 5 (discussed in sec. 4).
[5] The ‘thesis’ is firstly visualised in Figure 1, discussed in sec. 4; the ‘propositions’ appear in the comparative graph in Figure 3.
[6] Though the TheSu XML Schema Definition (XSD) is already available and at an advanced stage of development, with no significant changes expected in the next years, some of its sections are still work in progress. Especially the element tree for ‘sequences’ (on which see below, sec. 3.1) is being completed and refined as part of the current project.
[7] Note that the analysis of technical processes represents just one application of the features for annotating ‘sequences’. These can also model any ordered series of events described in sources, making them potentially also suitable, for example, for analysing and comparing historical narratives.
[8] On lead white being the only artificially produced white pigment in antiquity see the references in Liou et al. (1995, p. 175); on it being virtually never extracted in antiquity from naturally occurring cerussite see Kunzelman (2008, p. 101, 116–17), with references.
[9] While Foster (1934, p. 225) still assumed it to be basic lead carbonate, Bailey (1932, p. 204) influentially argued that it could only be lead acetate. See Principe (2018, pp. 163–165) for further references and criticism of Bailey’s unwarranted conclusions, building on Stevenson (1955). Caley & Richards (1956, p. 187) suggested that the terms could refer to both basic lead carbonate and lead acetate. Archaeological findings consistently point to lead carbonate (see, e.g., Foster, 1934, p. 225; Katsaros et al., 2010, KER-1 sample; Photos-Jones et al., 2020, pp. 2–3, 5–10), but post-depositional alteration prevents using these findings as reliable evidence for the substance’s original chemical state.
[10] Most of these are reported in translation in Gliozzo & Ionescu (2022, pp. 4–5).
[12] For medical uses, see Photos-Jones et al. (2020, p. 2); its toxicity was known (Caley & Richards, 1956, p. 190). For artisanal uses (coating, wall painting, artworks) see the references in Gliozzo & Ionescu (2022, p. 12); for its processing into red lead (minium) and other chemical uses see Liou et al. (1995, pp. 175–176) and Gliozzo & Ionescu (2022, pp. 14, 17). Its strong association with feminine cosmetics is discussed by Shear (1936, pp. 315–316), Caley & Richards (1956, p. 190), and Photos-Jones et al. (2020, pp. 1–2), and corroborated by archaeological evidence like the pyxides commonly found in female graves (Shear, 1936, p. 314).
[13] See Artechne Team & Digital Humanities Lab (2019).
[14] See The Ohio State University Press Open Access Initiative guidelines: https://ohiostatepress.org/books/Openaccess.htm.
[15] For canonical-greekLit and First1KGreek sources, I used the latest versions available on the projects’ GitHub collections on February 14, 2025. For the Zenodo records of these corpora, see Cerrato et al. (2019) and Crane et al. (2022), respectively.
[16] Specifying ‘speakers’ was crucial for Plato’s Lysis, a dialogue where the main character advancing claims and arguments is Socrates (a voice different from the author’s). See also below, sec. 4 about the annotated contributions from Menexenus, another character in the dialogue.
[17] In the annotation document, see, for example, the analogies analysed within Dioscorides’s thesis ‘tlg0656_tlg001.T211310’ (“The best lead dross is that which looks like lead white […]”) and the aetiology annotated within Plutarch’s ‘tlg0007_tlg112.T151258’ (“It is likely that thinner waters are dominated by the cold because of their weakness”).
[18] This interpretative stance acknowledges that arguments presented in natural language – including those attested in most historical sources – are usually enthymematic, requiring pragmatic reconstruction (see Lopes Cardoso et al., 2023, p. 4). Recent computational models of argumentation sharing a similar spirit include IAMT, which, on the spectrum between the approaches defined as “reconstruction” and “annotation”, privileges the former approach, yet aiming to integrate it with surface-level “annotation” of syntactic relationships signalled by discourse markers (Rocci & Lucchini, 2025, paras 2, 6–7, 11–13, cf. Doury & Pilon, 2025, paras 5–10, 72–85). The design of TheSu XML is even further oriented towards the “reconstruction” pole, treating surface-level information only as textual evidence for the annotated interpretation of the conveyed ‘discourse’, rather than as an object of analysis in itself.
[19] It may be classified as a “sub-maximal principle of charity”, according to the taxonomy proposed by Shields (2025), pp. 9, 11.
[20] Future web platforms for the presentation of TheSu XML annotations could ideally utilise such provenance metadata to allow users to compare different annotations when present. This would be analogous to examining multiple translations, commentaries, or editions of the same text.
[21] See especially Dioscorides’s recipes, starting from theses ‘tlg0656_tlg001.T213927’ (the first described procedure, including minor variants specified for individual steps) and ‘tlg0656_tlg001.T165902’ (a variant recipe with more significant divergences).
[22] These external XML files (included in the dataset) list controlled vocabularies for all tags used in the annotation and provide full TEI-encoded bibliographic entries for cited works (either as published editions of the annotated textual sources or as containing passages relevant to the discussion of annotation components, e.g. within the elements ‘externalRef’ or ‘scholarContra’). In the current dataset, the authority records do not report definitions for each tag; these should ideally be included in future projects utilising TheSu XML.
[23] See Perseus Stable URIs (n.d.).
[24] Of course, not all relations annotated in TheSu XML are convertible to AIF; as an encoding specific to argumentation, it cannot, for example, accommodate the additional discourse functions modelled by ‘support’ elements.
[25] This allows for precision in the selection of textual evidence associated with any element. For example, Plutarch’s thesis ‘tlg0007_tlg112.T145739’ (“Lead white is the most cooling of deadly drugs”) can in this way link to the discontinuous text segment τὸ ψυκτικώτατον τῶν θανασίμων φαρμάκων […] ψιμύθιον, skipping the irrelevant word ἐξανίησι.
[26] This expression refers to formally invalid cases in XML where multiple distinct elements refer to partially shared but non-identical text spans. For example, Plutarch’s support ‘tlg0007_tlg112.S143800’ is associated with the phrase ὅς γɛ τριβόμɛνος ὄξɛι τὸ ψυκτικώτατον τῶν θανασίμων φαρμάκων ἐξανίησι, which partially overlaps with the text span associated with the thesis described in the previous footnote. The word occurring after ἐξανίησι, i.e., ψιμύθιον, is essential to the thesis but irrelevant to the support, which only employs thesis ‘tlg0007_tlg112.T143932’ (“Lead can produce the most cooling of deadly drugs”) as an argumentative premise for ‘tlg0007_tlg112.T133442’ (“Lead is among the naturally cold substances”) – on this interpretation see below, sec. 4, with n. 44. Without stand-off annotation, encoding these two elements would require either omitting part of the relevant text or falsely including unrelated material, since XML does not allow elements to cross over one another.
[27] The stand-off approach pervades the schema. For example, keywords associated with a ‘thesis’ are not annotated as strings within the ‘thesis’, but as references to separate ‘keyword’ elements. This design allows for multiple ‘thesis’ elements linking to the same ‘keyword’, for cases in which these ‘theses’ are associated with common text segments. In this way, for example, the same keyword ‘tlg0007_tlg112.K141236’ for the adjective ‘cooling’ can be shared by both theses ‘tlg0007_tlg112.T145739’ and ‘tlg0007_tlg112.T143932’ described in the previous two footnotes, preventing redundancies and enabling better indexing.
[28] Within Argumentation Mining, see, e.g., Lopes Cardoso et al. (2023, pp. 12–13); Wang et al. (2022, p. 867).
[29] The current schema is exclusively text-oriented, but the future development roadmap includes expanding its features and definitions specifically for the annotation of non-textual media.
[30] Cf. Lenz & Bergmann (2025).
[31] The digital edition was produced from a PDF scan processed with Tesseract OCR, converted first to HTML and then to XHTML via custom Python scripts. PDF page breaks were retained, paragraphs were identified based on Tesseract’s OCR box detection, and page numbers were adjusted using another script to match the original printed pagination.
[32] See the tool provided as open source in Morrone (2026), a refactored and polished version of the script used while creating the dataset associated with this paper. Using that tool, it is possible to replicate exactly the same graphs and visualisations included in this dataset.
[33] For a recommended introduction to argumentation mapping, though presented within the context of Argumentation Mining, see Stede & Schneider (2019), pp. 36–43.
[34] For an example see below, sec. 4, with Figure 4.
[35] For an excellent recent overview see Mardah (2024), pp. 58, 60–67. This is not exhaustive: also consider, for example, Putra et al. (2023) for their tool TIARA 2.0.
[36] Such outcomes conflict with the objective referred to by Rathkopf as “informational encapsulation” (Rathkopf, 2024, pp. 7–8), according to which the map’s design should isolate argument components to limit the user’s cognitive load in trying to understand them – precisely the opposite of the effect produced by visual clutter.
[37] See the IAMT diagram generated using OVA+ for a brief passage in Rocci & Lucchini (2025, Figure 12). The authors do acknowledge that their model’s applicability is currently “limited to collections of short argumentative passages extracted from a larger corpus for in-depth analysis” (para. 58).
[38] Cf. the decluttering techniques implemented in TIARA 2.0 (Putra et al., 2023).
[39] Cf. Doury & Pilon (2025, para. 100), who ultimately “dream” of a tool capable of generating representations depending on the preferred analytical lens, through a flexible filtering system.
[40] A similar side-by-side display with interactive linking was implemented in DIGAT (Kirschner et al. 2015; Ubiquitous Knowledge Processing Lab 2020). Other frameworks contest this design; for example, Putra et al. (2023, p. 21) argue for a “dual-view” interface requiring users to switch between text and graph.
[41] Cf. Drakman & Gelfgren (2022), who argue for a mixed-method approach to the history of science and ideas, combining close and distant reading.
[42] Future versions of the script will generate the Gephi-optimised file directly and in GEXF format, which offers better compatibility with the software and full customisation options (e.g., encoding edge weights).
[43] This method and its purposes have similarities with Herfeld’s (2025) network-based approach to the study of philosophical concept formation. The use of TheSu XML differs inasmuch as the schema supplies a new kind of semantic data to analyse, focusing on epistemic, rhetorical, and argumentative content, instead of objective patterns of co-citation from which to infer conceptual diffusion and model transfer across scientific domains.
[44] See Morrone (2022, pp. 8–13).
[45] The fuller graph represents my interpretation, previously outlined in Morrone (2022, pp. 86–87, 93–94, 137–139), that both the cooling and “thinning” effects Plutarch attributes to lead are explained through its essential coldness. Given this interpretation’s speculative and potentially controversial nature, the relevant annotation components were provided ‘scholarPro’ child elements with bibliographic references – one of several ways TheSu XML can link components to supporting, opposing, or neutrally relevant scholarship.
[46] By marking this premise as ‘implicit’, TheSu XML signals its status as an interpretative reconstruction of an element not explicitly attested in the source, but supposed to match the author’s or speaker’s communicative intention in the pertinent passage.
[47] Cf. Rocci & Lucchini (2025, paras 27–29), whose IAMT model uses grey for “reconstructed implicit standpoints”.
[48] This dialogue segment contributes to a broader exploratory argument about whether the neither-good-nor-bad loves the good because of the bad, which Socrates ultimately dismisses as “nonsense” at 220b-221c (see Penner & Rowe, 2005, pp. 99–124). I have not annotated this macro-context, but TheSu XML allows encoding such broader framing information: this would be valuable for users reading the graph to understand the passage’s purpose and validity within the fuller dialogue.
[49] While the broader argument is rejected later in the dialogue (see the preceding footnote), the individual theses and arguments composing it, including the lead white examples, probably retain their value (see Penner & Rowe, 2005, pp. 110, 153–156).
[50] Dioscorides presents variant recipes and alternative steps. In the annotation document, I created a model procedural ‘proposition’ from the first recipe only, and connected all variant recipes to the ‘thesis’ that represents the first. I used advanced TheSu XML features to annotate the variants, specifying which steps of the first recipe are repeated, which are altered, and how. Since the current visualisation script does not yet support display of this information, only the first recipe is represented with connected ‘phase’ nodes, and it is the only one included in the view.
[51] My annotation connects Plutarch’s compressed report on the formation of lead white to both Dioscorides’s and Theophrastus’s recipes as it shares similarities with both. On these recipes see below.
[52] Since neither author states, or means to specifically communicate, that lead white is white in colour, I marked these theses as ‘extrinsic’. Unlike ‘implicit’ theses (see above, n. 45), ‘extrinsic’ theses represent points that may be interpreted as logically entailed by the text, but which the author does not intend to communicate. This greater interpretative distance is reflected in the visualisation, where ‘extrinsic’ thesis nodes have no edges connecting them to ‘source’ nodes, but only to the ‘thesis’ nodes from which they are extrapolated.
[53] See Caley & Richards (1956, pp. 57, 187–188); Eichholz (1965), pp. 79, 125; Principe (2018), pp. 167–169; Photos-Jones et al. (2020, pp. 3, 10).
[54] Similar simplifications apply to argumentative and expository discourse connections: mediating label-nodes like ‘EMPLOYED IN’, ‘EXPLAINS’, and ‘ENTAILS’ (all visible in Figure 4) are removed in the Gephi version, replaced by direct labelled edges.
[55] The Graphviz ‘dot’ visualisation in directory ‘6b.KLL.recipe’ is meant to demonstrate how integrating the visual analysis of the recipe within its mapped discourse structure can reveal important contextual information that might not be immediately apparent from reading the procedure description alone.
[57] The graph only includes the first recipe presented by Dioscorides. See above, n. 49, on the current limitations in displaying variant recipes.
[58] For example, Theophrastus’s ‘propositional’, light-violet, step 6 in the lower side of the graph (“Filter off what you are grinding”), also attested in Dioscorides’s first recipe and Caley & Richards, appears between these two procedures but relatively far from the Theophrastus recipe on which it is modelled.
[59] For ‘ForceAtlas2’, see Jacomy et al. (2014).
[60] Cf. van den Berg et al. (2018), who emphasise that visualisations without provable guarantees of faithful data representation may mislead users, as in their case study of a Gephi force-directed layout suspected of distorting relationship patterns in textual analysis.
[61] In Gephi’s ‘Overview’ tab, clicking on a node using the ‘Edit’ function reveals the node’s attributes: e.g., for a ‘thesis’, text, paraphrase, passage reference, speaker, etc.
[62] See the above considerations on Figure 6, but also, for example, the proximity in Figure 7 between nodes from Plato’s cluster and Caley & Richards’s, certainly not reflecting similarity in content (which would be signalled by mediating ‘proposition’ nodes).
[63] As mentioned above, sec. 3.1, the schema’s potential extends even beyond textual sources: TheSu XML may in principle be used to annotate any medium where components can be represented as declarative sentences and receive unique digital identifiers.
[66] See above, sec. 3.1, on the planned annotator GUI; sec. 3.2, with nn. 37 and 39, on interactive filtering and linked graph/text views; and sec. 4, with n. 60, on Gephi’s limited features for accessing node and edge data.
[73] In the formal schema documentation referenced above in Table 1, note c, see in particular thesu:scholarshipBundle under Element Groups and thesu:proof under Elements; see also above n. 21.
[74] For the formal schema documentation and the reader-oriented guide to the main features of TheSu XML, see Table 1, notes c, d.
Data Accessibility Statement
The dataset described in this article is openly available on Zenodo at https://doi.org/10.5281/zenodo.20290847, under a CC-BY-SA-4.0 licence. The content of a segmented source file derived from the copyrighted commentary by Caley & Richards (1956) is largely redacted to avoid substantial reproduction of the original text, as explained in sec. 3.1.
Acknowledgements
The author would like to thank the editors of the Data-Driven History of Ideas special collection of the Journal of Open Humanities Data, Arianna Betti and Hein van den Berg, for their feedback following the acceptance of the initial abstract. Their comments helped substantially improve the paper’s conception and its alignment with best practices in digital humanities. The author is also grateful to Margherita Fantoli for feedback on the text, which especially helped improve clarity, and to the three anonymous reviewers for their helpful requests and suggestions.
The features discussed in this paper for encoding, visualising, and comparing ‘sequences’ are presented here in a scholarly publication for the first time. They were first presented on 12 July 2022 at the AlchemEast ERC project’s monthly seminar, “What’s in a Recipe? Epistemology and practical understanding of recipe literature”, organised by Gabriele Ferrario and Matteo Martelli. The author would like to thank all members of the AlchemEast project for inspiring this approach and indirectly contributing to the design of the ‘sequence’ encoding components and their visualisations. During the initial design and testing phase of these features, Lucia Raggetti and Gabriele Ferrario generously provided broken-down translations of medieval Arabic and Hebrew recipes with which to test the annotation schema on those languages. The author is also especially grateful to Caterina Manco for an hours-long brainstorming session, during that same phase, which helped optimise the element tree and attributes used for encoding recipes as ‘sequences’.
Author contributions
Daniele Morrone: Conceptualisation; Data curation; Formal analysis; Funding acquisition; Investigation; Methodology; Project administration; Resources; Software; Validation; Visualisation; Writing – original draft; Writing – review & editing.


