1 Context and motivation
The field of medieval Greek literature offers a rich variety of textual genres that reflect complex literary, cultural, and historical dynamics across centuries. Among these, book epigrams, i.e. brief versified writings added in manuscript margins, represent a distinctive mode of literary engagement, shedding light on how medieval scribes, readers, and poets interacted with canonical texts. These marginal epigrams, typically concise and self-referential, encapsulate a wealth of interpretive and intertextual nuances essential for understanding the transmission and reception of medieval Greek poetic traditions (Bernard & Demoen, 2019). At the same time, they provide a valuable window into the contemporary concerns and intellectual environment of their scribes.
In this context, book epigrams stand out as a particularly important corpus. Appended to or interspersed within manuscripts, they are not merely decorative or peripheral texts. Functioning as micro-commentaries, they often thematically echo or instead contrast with the primary text, thereby enriching the reader’s experience and offering a glimpse into the medieval intellectual milieu. Due to their interconnected nature (Ricceri et al., 2023), book epigrams lend themselves especially well to detailed semantic analysis. The study of semantic similarity among verses in the book epigrams is not simply an academic exercise. It addresses a pressing practical challenge for scholars in both philology and digital humanities. Manuscript-based research has traditionally relied on laborious manual collation and comparison of text passages to identify thematic parallels or variant readings. This approach, while meticulous, becomes impractical when dealing with large corpora or extensive manuscript traditions. The ability to automatically detect semantically similar verses would therefore be transformative, enabling the rapid identification of relevant material for textual criticism, commentary, and comparative literary analysis.
Recent advances in natural language processing (NLP) and computational semantics present promising avenues to support such academic endeavours. However, computational models and benchmarks for measuring semantic similarity are largely developed for modern languages and prose texts (Agirre et al., 2012). Although benchmark resources for semantic similarity in Ancient Greek have recently emerged, notably AGREE (Stopponi et al., 2024), no dataset currently targets expert-annotated, graded semantic similarity at the verse level in highly formulaic poetic corpora such as medieval Greek epigrams. This absence underscores the need for specialised annotated datasets and evaluation frameworks that capture the nuanced semantic relationships inherent to these texts. At the same time, within ancient language processing (ALP) a strong line of research on semantics is thriving, with work on semantic similarity benchmarking (Stopponi et al., 2024), polysemy (McGillivray et al., 2019), and sentence level meaning (Krahn et al., 2023). Related work also includes computational approaches to poetic formulaicity and variation, such as HoLM, which models linguistic unexpectedness in Homeric poetry rather than explicitly annotated semantic similarity (Pavlopoulos et al., 2024). Within philological applications, however, similarity detection has so far been predominantly limited to orthographic similarity detection, with—if so desired—relaxation (Deforche et al., 2023; Swaelens et al., 2024).1
A further challenge concerns the annotation of similarity itself. In contrast to tasks where absolute labels can be assigned, similarity is inherently relational and often difficult to capture on a fixed scale, especially in highly allusive and context-dependent material such as poetry. Comparative annotation strategies, such as pairwise comparison and best–worst scaling, have therefore been proposed as more robust alternatives, allowing annotators to express relative judgements rather than absolute scores (Fechner, 1860; Louviere et al., 2015; Thurstone, 1927). These approaches are particularly suited to domains in which similarity cannot be reduced to a single dimension, but emerges from the interaction of multiple, partially competing criteria.
Focusing on book epigrams from the Database of Byzantine Book Epigrams (DBBE) (Demoen et al., 2023), the present study seeks to bridge this gap by initiating the annotation of a similarity benchmark and by exploring computational methods to similarity detection, tailored to the linguistic and stylistic particularities of medieval Greek poetry. The outcomes are intended to expand the toolkit available to philologists and digital humanists, supporting more scalable and systematic investigation into similarity detection. More broadly, the work contributes to the ongoing integration of computational techniques into the study of classical and medieval literature, generating new insights into textual transmission, literary intertextuality, and cultural continuity.
2 Dataset description
Repository location
Repository name
https://github.com/coswaele/ByzantineGreekDatasets/blob/main/STS_benchmark_medieval Greek.tsv
Object name
STS - Byzantine Greek
Format names and versions
TSV
Creation dates
2025-03-01 — 2025-05-31.
Dataset creators
Colin Swaelens (curator & annotator), Flo Christiaens, Kyra De Backer, Gaelle Lamens, Elise Pype, Lode Sans, Lisa van Boven, Jutta Van Daele, Maxime Van Houdt, Nieke Van Impe, Camiel Vanden Eynde Wessel, Marie Vandensande, Nina Vannieuwenhuyse (annotators).
Language
Ancient Greek.
License
CC0.
Publication date
2026-03-30.
3 Method
In the context of our research, pairwise annotation means presenting a pair of verses together with a reference verse to one or more annotators. The judge is then required to select one of the two verses in the pair, indicating a preference in the specific context of the comparison, even if this does not reflect their overall preference. The situation becomes more complex when three or more objects must be compared.2 All objects (A, B, C) are evaluated in pairs (AB, AC, BC), and each comparison—typically—results in a preference for one of the two objects. Each of the three pairwise comparisons has two possible outcomes, leading to a total of eight possible result patterns: six of the type [2 1 0] and two of the type [1 1 1]. In the former, one object (in this example, A) receives two wins, another one win, and the third none; in the latter, each object has exactly one win. This last phenomenon is called a circular triad (Kendall and Smith, 1940) and illustrates that not all preferences can be linearly ordered without contradiction.
Despite such inconsistencies, the resulting pairwise judgements can be modelled using established statistical techniques (Bradley and Terry, 1952; Thurstone, 1927), which can be used to estimate latent similarity scores across all annotated items. These models allow us to construct an internally coherent similarity scale without requiring annotators to impose absolute values. For all these reasons, we adopted pairwise comparison as the core annotation strategy for building our gold standard.
3.1 Pilot study
Before annotating the set of verses that will eventually constitute our gold standard, we conducted a pilot study to assess both the feasibility of pairwise annotation and the level of detail required in the annotation guidelines. Ten philologists from two different institutions participated in the study, which consisted of two parts. Firstly, verse 1 served as the reference verse; secondly, verse 1 did. The remaining verses were presented in all possible pairwise combinations, as visualised in Figure 1. Annotators were given a single instruction: to select the verse from each pair they considered most similar to the reference verse.

Figure 1
Screenshot of the pilot study set-up.
The decision not to provide more detailed annotation guidelines was the subject of considerable discussion. Ultimately, we opted for a minimal-instruction setup. In pairwise annotation, annotators are required to make a choice even when both candidates exhibit little or no meaningful similarity, which limits the effectiveness of predefined hierarchies of criteria (e.g. semantic over syntactic or metrical similarity). Rather than enforcing potentially arbitrary distinctions, we left the notion of similarity intentionally underspecified, allowing annotators to rely on their own interpretative strategies. Consequently, the dataset should not be understood as capturing “semantic similarity” in a narrow sense, but rather a broader notion of perceived similarity. To gain further insight into the decision-making process, annotators were subsequently asked to reflect on the criteria they employed; these reflections are analysed in Section 4.
To examine the feasibility of the task, we analysed the presence of circular triads and computed inter-annotator agreement. No circular triads were observed, indicating consistent and non-contradictory judgements. Inter-annotator agreement was calculated using Cohen’s kappa coefficient (Cohen, 1960), as all annotators evaluated all pairs. Agreement was computed pairwise between annotators (Table 1) and averaged, resulting in an overall score of 0.5769. This corresponds to moderate agreement, which is encouraging given the subjective nature of the task (Artstein and Poesio, 2008).
Table 1
Cohen’s kappa between all ten annotators.
| a | b | c | d | e | f | g | h | i | j | |
|---|---|---|---|---|---|---|---|---|---|---|
| a | - | 0.3571 | 0.3571 | 0.3571 | 0.5199 | 0.0323 | 0.3077 | 0.3571 | 1 | 0,3571 |
| b | 0.3571 | - | 1 | 1 | 0.7391 | 0.5199 | –0.125 | 1 | 0.3571 | 1 |
| c | 0.3571 | 1 | - | 1 | 0.7391 | 0.5199 | –0.125 | 1 | 0.3571 | 1 |
| d | 0.3571 | 1 | 1 | - | 0.7391 | 0.5199 | –0.125 | 1 | 0.3571 | 1 |
| e | 0.5199 | 0.7391 | 0.7391 | 0.7391 | - | 0.3571 | 0.0323 | 0.7391 | 0.5199 | 0.7391 |
| f | 0.0323 | 0.5199 | 0.5199 | 0.5199 | 0.3571 | - | 0.1724 | 0.5199 | 0.0323 | 0.5199 |
| g | 0.3077 | –0.125 | –0.125 | –0.125 | 0.0323 | 0.1724 | - | –0.125 | 0.3077 | –0.125 |
| h | 0.3571 | 1 | 1 | 1 | 0.7391 | 0.5199 | –0.125 | - | 0.3571 | 1 |
| i | 1 | 0.3571 | 0.3571 | 0.3571 | 0.5199 | 0.0323 | 0.3077 | 0.3571 | - | 0.3571 |
| j | 0.3571 | 1 | 1 | 1 | 0.7391 | 0.5199 | –0.125 | 1 | 0.3571 | - |
(1)
a.
Ὥσπερ ξένοι χαίρουσιν ἰδεῖν πατρίδα,
Just like travellers rejoice upon seeing their homeland,
b.
καὶ οἱ θαλαττεύοντες εὑρεῖν λιμένα,
and sailors upon finding a harbour,
c.
καὶ οἱ στρατευόμενοι ἰδεῖν τὸ νῖκος,
and soldiers upon seeing victory,
d.
καὶ οἱ πραγματευόμενοι ἰδεῖν τὸ κέρδος,
and craftsmen upon seeing some profit,
e.
οὗτω καὶ οἱ γράφοντες ἰδεῖν βιβλίου τέλος.
so do scribes upon seeing the end of the book.
3.2 Annotation Process
After the successful pilot study, carried out to validate our setup, we started the annotation of our set of 100 verses. The annotations were carried out by 11 annotators: one expert in Byzantine studies and ten students from the classics department. As the task involves pairwise comparison, directionality becomes a potential issue: annotators may assign different similarity judgements depending on which verse is presented as the reference. This asymmetry must be taken into account when interpreting the resulting similarity scores.
The dataset was divided into batches, each centred around a single reference verse and the other 99 candidate verses. Rather than directly assigning a similarity score to each reference–candidate pair, annotators were presented with all pairwise combinations of the 99 candidate verses, resulting in 4,851 comparison pairs per batch ( with t = 99). For each comparison pair, annotators selected the verse they considered more similar to the reference verse. The batches were presented through a custom web application, in which annotators evaluated sets of 10 pairs at a time (Figure 2). In each set, three pairs were annotated twice, allowing for the continuous estimation of inter-annotator agreement.

Figure 2
A screenshot of the web application used for the pairwise annotation. This shows the first three out of ten pairs presented on the page.
Since not all annotators completed the full dataset, agreement was computed using Fleiss’ kappa, resulting in a score of 0.6584. This moderate agreement reflects both the partial overlap in annotated items and the inherent difficulty of the task. Each annotator worked for approximately ten hours in total, in sessions of no more than 30 minutes, following feedback from the pilot study that sustained attention is difficult to maintain for this task. In total, 18,897 verse pairs were annotated (three full batches and 1,448 duplicated pairs for agreement estimation). These pairwise preference judgements were subsequently aggregated into counts for each candidate verse, based on how often it was preferred over another candidate. The resulting counts were normalised to similarity percentages, yielding the final benchmark of 296 reference–candidate similarity scores released in the dataset (cf. Table 2).
4 Results and discussion
The inter-annotator agreement study displays a moderate agreement, which means that the annotation of the gold standard is quite reliable. Nevertheless, a qualitative analysis of this study will provide insight into the reasons for disagreement, which we expected to stem either from verses that are extremely similar or, on the contrary, from verses that do not resemble to the base verse at all. As pointed out in Section 3.2, we annotated 3 sets with the paired comparison method. Examples 2, 2 and 3 were the respective reference verses to which each of the pairs were judged.
(2)
a.
Αἶνος θ(ε)ῶ χάρις τε καὶ δόξα πρέπει·
b.
τῷ δόντι τέρμα τῆς γραφῆς φθάσαι σθένος
to Him who gives strength to reach the end of this writing.
(3)
Ἐνταῦθα νῦν σκόπησον ὀρθῶς ὁ βλέπων.
Thou, who sees, behold now here correctly (…)
As we were only able to annotate three out of the hundred sets, we highlight here the annotation of each of the three reference verses (16977 1, 16977 2, and 17148 1) against one another in both directions. Table 3 presents the outcomes of the inter-annotator agreement study for these verse pairs. Notably, both of the annotators independently converged on the same preferences in all three cases, indicating full agreement. For example, when comparing 16977 1 with 16977 2 and 17148 1, both annotators preferred 17148 1 over the others, and similarly consistent patterns emerge in the other two rows.
Table 3
The results of the inter-annotator agreement study of the three reference verses. Preference 1 is the choice of annotator 1, preference 2 is the choice of annotator 2.
| REFERENCE | VERSE 1 | VERSE 2 | PREFERENCE 1 | PREFERENCE 2 |
|---|---|---|---|---|
| 16977 1 | 16977 2 | 17148 1 | 17148 1 | 17148 1 |
| 16977 2 | 16977 1 | 17148 1 | 16977 1 | 16977 1 |
| 17148 1 | 16977 1 | 16977 2 | 16977 2 | 16977 2 |
Table 4 complements this analysis by displaying the similarity scores between each pair of reference verses. These scores should be read from left to right and represent how closely one verse is judged to resemble another. Interestingly, verse 16977 2 shows a very high similarity score toward 16977 1 (0.9545), whereas the reverse score from 16977 1 to 16977 2 is much lower (0.2336). This asymmetry suggests that, while 16977 2 appears strongly dependent on 16977 1, the reverse is not perceived as true to the same extent. The highest overall agreement in Table 3 coincides with some of the highest similarity scores, reinforcing the reliability of both the annotation and the underlying scoring model. Nonetheless, the substantial differences in perceived similarity between verses underline a methodological challenge we have not yet resolved. While our annotations were conducted as objectively and consistently as possible, these results also expose the complexity of directionality in similarity—a phenomenon whose interpretation requires taking the relative nature of the annotation task into account.
Table 4
The similarity scores between the verse pairs of the reference verses.
| 16977 1 | 16977 2 | 17148 1 | |
|---|---|---|---|
| 16977 1 | / | 0.233610 | 0.118850 |
| 16977 2 | 0.954545 | / | 0.123377 |
| 17148 1 | 0.646825 | 0.599206 | / |
To better understand the observed asymmetries and sources of disagreement, annotators were asked to reflect on the strategies they employed during the paired comparison task. All but one annotator reported following a broadly hierarchical decision process: lexical overlap between verses was considered first, followed by semantic similarity and, when necessary, broader semantic fields or thematic domains. Syntactic structure generally served only as a secondary criterion. One annotator diverged from this pattern by prioritising metrical similarity over semantic field similarity. These reflections suggest that human judgements were guided by a layered combination of lexical, semantic, and formal criteria, partially explaining the moderate disagreement between annotators.
Importantly, the observed directionality does not necessarily imply that annotators assessed a fundamentally different notion of similarity from the computational models. Rather, it partly reflects the design of the annotation task itself, in which verses were not assigned absolute similarity scores but were ranked relative to a fixed reference verse. The resulting benchmark therefore captures not only semantic similarity in a strict symmetric sense, but also relative semantic salience with respect to the reference verse. Model performance should therefore be interpreted with this asymmetry in mind.
Based on the data in Table 5, the five verses most similar to verse 16977 1 (Example 2) all revolve around expressions of divine praise, typically in reference to the beginning or end of a manuscript. Notably, all but one (verse 18886 1) contain the dative form θεῷ (“to God” or “with God”), which suggests a strong lexical and thematic signal driving the similarity. The two highest-ranked verses, each scoring a perfect 1.0, are essentially duplicates of the reference verse. By contrast, the five least similar verses show minimal thematic alignment, with most referencing the end of literary or dramatic works without invoking divine language. This divergence is reflected in their low similarity scores, which drop to as little as 0.02073.
Table 5
The 5 most similar verses to verse 16977 1 and the 5 least similar verses according to our annotations.
| OCCURRENCE & VERSE | TEXT | SIMILARITY SCORE |
|---|---|---|
| 20963 1 | τῷ θ(ε)ῷ ἡμῶν δόξα τῷ δόντι ἀρχὴν καὶ τέλος:+ | 1.0 |
| 22220 1 | δόξα τω θ[ε]ῶ δόντι καὶ ἀρχὴν (καὶ) τέλος˙ | 1.0 |
| 18150 1 | πῆναξ σὺν θεῶ ἁγ(ίας) τῆς βίβλου ταύτης | 0.86528 |
| 18034 1 | + τέλος σὺν θ(ε)ῶ τοῦ δευτέρου βιβλίου:+̃ | 0.82383 |
| 18886 1 | Ἀρχῆς καλλῆς κάλλιστον εἶναι καὶ τέλος | 0.79792 |
| … | ||
| 24479 1 | Εἴληφε τέρμα τῶν ἑπτὰ ἐπί θήβας | 0.04663 |
| 33429 1 | τέλος τῆς γάμα ὁμήρου ῥαψῳδίας | 0.04145 |
| 23122 6 | τοίγαρ ἀντολίηθεν ἐωσφόροι ἀμφανέντες | 0.04145 |
| 34901 3 | τῶν συλλαβῶν δὲ τὴν ἀρίθμησιν κύκλον | 0.02591 |
| 20839 1 | Τέλος τοῦ τρίτου δράματος σοφοκλέους | 0.02073 |
5 Implications/Applications
In the following section, we present a series of experiments based on this benchmark to assess the extent to which automated methods approximate human judgements of semantic similarity. Across all experiments, cosine similarity was employed as the distance metric, while both the type of embeddings and the level of analysis varied. Specifically, this study examined the performance of different embedding models in evaluating similarity at both the word and verse levels, with the aim of identifying which combinations best capture nuances of meaning in unedited ancient Greek. In order to evaluate the quality of the predicted similarity scores, the root mean squared error (RMSE) was computed against the gold standard annotations, quantifying the absolute deviation from human scores.
Given the inconsistent orthography of the book epigrams, we tried to reduce the noise as much as possible by lowercasing the texts and removing punctuation or special characters. The diacritics have not been stripped off.
5.1 Exploration of orthographic similarity
As a baseline, we first evaluated a recently developed hierarchical orthographic similarity measure (Deforche et al., 2024), which compares verses based on the similarity of their word forms. The method computes similarity between individual words and subsequently aggregates these scores into a verse-level distance measure using an adapted Levenshtein algorithm.3 When comparing the resulting orthographic similarity scores to our benchmark, the method achieved an RMSE of 0.2422. This exploratory experiment served two purposes: first, to assess whether orthographic similarity correlates with semantic similarity; and second, to establish a methodological baseline for the semantic similarity experiments presented in Section 5.2.
5.2 Word-Level
The experiments to compute similarity between the verses at the word level followed a consistent methodology. Firstly, four types of embeddings (Skip-Gram, CBoW, GloVe and BERT) were each trained in two variants: once on raw tokens and once on lemmatised forms. Embeddings are numerical vector representations of words, designed such that semantically related words occupy nearby positions in vector space. Each word in a verse was thus mapped to a vector, after which an adapted Levenshtein distance was computed to compare entire verses based on these word representations.
For each verse pair, tokenisation and lowercasing were applied. The tokens were subsequently aligned using an approach analogous to the classical Levenshtein distance. Substitution costs were derived from the semantic similarity between word vectors, such that semantically similar words incurred a lower penalty than unrelated words. Insertion and deletion operations were each assigned a fixed cost of 1. The resulting alignment cost was normalised by dividing it by the length of the longer verse. The semantic similarity score was then defined as the negative of this normalised alignment cost. This procedure was carried out twice: first using token-based embeddings, and subsequently using lemma-based embeddings.
Skip-Gram & CBoW
Initial experiments employed standard distributional embedding models, namely Skip-Gram and CBoW (Mikolov et al., 2013), to vectorise the tokens. Given the historical and orthographically variable nature of the corpus, particular attention was paid to out-of-vocabulary (OOV) handling. To mitigate this issue, we used the FastText implementation (Bojanowski et al., 2016), which can generate embeddings for previously unseen word forms based on subword information. The models were trained on both tokenised and lemmatised versions of the corpus. Lemmatisation reduces vocabulary size and consolidates contextual information across inflectional variants, but at the cost of removing orthographic variation, which is a defining feature of the book epigrams. In this case, lemmatisation was performed using the DBBErt lemmatiser (Swaelens et al., 2025a), trained on Classical Greek. The results are presented in Table 6. Embeddings derived from lemmatised input consistently outperform their token-based counterparts in terms of RMSE, suggesting that lexical normalisation facilitates more stable representations.
Global Vectors
GloVe embeddings (Pennington et al., 2014) were next employed to vectorise the tokens of all verses. As a static embedding model relying on global co-occurrence statistics, GloVe does not provide a mechanism for handling out-of-vocabulary (OOV) items, which is particularly problematic for historically variable corpora. In our benchmark dataset, 137 words were not present in the vocabulary. To preserve methodological consistency, all verse pairs containing at least one OOV token were excluded from the evaluation. This filtering step drastically reduced the usable dataset for token-based embeddings to 19 verse pairs, which limits the interpretability of the token-based results. In contrast, lemmatised input did not suffer from OOV issues, as lexical normalisation collapses inflectional and orthographic variation into a smaller, more stable vocabulary. Under these conditions, lemma-based GloVe embeddings achieve a relatively low RMSE (0.0815), though still slightly higher than that obtained with lemmatised CBoW embeddings. However, given the substantial reduction of the evaluation set for token-based input and the inherent simplifications introduced by lemmatisation, these results should be interpreted with caution. These results highlight the sensitivity of static embedding models to vocabulary coverage in low-resource historical corpora.
Transformers
The final word-level experiment employed the pretrained transformer model DBBErt, a domain-specific BERT model pre-trained on ancient Greek (Swaelens et al., 2025b), to generate contextualised embeddings for each token. Unlike static embeddings, contextualised embeddings represent a word differently depending on its surrounding context. Each verse was tokenised and passed through the model, with embeddings extracted from the final hidden layer. As in the previous experiments, cosine similarity between token embeddings was used to compute a modified Levenshtein distance between verses. This method was chosen for its ability to capture fine-grained semantic relationships at the token level, while taking into account the sequential structure of the input. DBBErt achieved an RMSE of 0.2904, indicating a moderate level of absolute deviation from the human similarity ratings. This might point to limitations in the embedding’s ability to capture the specific semantic nuances characteristic of the corpus.
5.3 Verse Level
In addition to the word-level experiments, we also conducted a series of experiments at the verse level, treating entire verses as the basic unit of comparison. In these experiments, each verse was represented by a single embedding, allowing semantic similarity to be computed directly between whole verses rather than between individual words. The only exception to this general methodology is the final transformer-based experiment, which follows a slightly different approach.
Static Embeddings
The first approach to computing similarity at the verse level builds on the static word embeddings introduced in the previous section. Instead of using word-level cosine similarity as input for a Levenshtein-style distance between verses, this method directly computes the cosine similarity between full verse embeddings. Each verse embedding is constructed by averaging the embeddings of its constituent words, yielding a simple but effective baseline representation that ignores word order and syntactic structure. Both Word2Vec variants (CBoW and Skip-Gram) were used to generate word vectors, each trained on both the tokenised and lemmatised versions of the corpus.
The same approach was applied to GloVe embeddings, which were also evaluated in the same way. As with the word-level experiments, only verses containing no out-of-vocabulary tokens were retained. On the verse level, GloVe performed slightly better than the lemmatised CBoW model, with its lowest RMSE at 0.2028. Here again, the scores for the token-level GloVe embeddings are not reported because there were too few examples. Full results are presented in Table 7.
Contextualised Embeddings
DBBErt
To complement the static embedding approaches, a second set of experiments was conducted using contextualised verse embeddings derived from DBBErt, as previously applied at the word level in Section 5.2. Rather than averaging pre-trained word vectors, this method leverages the transformer architecture’s ability to encode contextual dependencies and token interactions. Each verse was passed through DBBErt to obtain a sequence of contextual embeddings. Two strategies were evaluated to obtain a single verse representation from these embeddings: using the special classification token ([CLS]) or averaging all token embeddings (mean pooling), allowing us to compare two common approaches to deriving verse-level representations from transformer models. Cosine similarity was then computed between each pair of verse embeddings. This allows comparison between two common strategies for deriving verse-level embeddings from transformer models.
In practice, mean pooling substantially outperformed the CLS token representation, as shown in Table 7: the lowest RMSE for mean pooling was 0.3181 compared to an RMSE of 0.6819 for the CLS-based approach. Notably, the performance of mean-pooled DBBErt embeddings performed less well than the best-performing static embedding models. However, when compared only to token-based embeddings, mean-pooled DBBErt achieved the best performance. This suggests that, for the task of computing semantic textual similarity between verses, static embeddings may already capture sufficient relevant information if used on lemmatised corpora. These findings suggest that the added value of transformer-based architectures for low-resource, domain-specific semantic similarity tasks may depend strongly on the availability of sufficient training data.
Sentence Bert
To make optimal use of the limited annotated data available, a Sentence-BERT model, specifically designed to generate sentence-level embeddings for similarity tasks, was trained and evaluated using 10-fold cross-validation. This strategy maximises the use of the limited dataset while still allowing robust out-of-sample evaluation. For evaluation, cosine similarity was computed between the sentence embeddings of each pair in the validation set. Performance resulted in an RMSE of 0.3181.
5.4 Pilot Study
The results from our benchmark experiments suggest that computing semantic textual similarity for this specific corpus remains a challenging task. As we are still annotating the benchmark set itself, these results are tentative. Therefore, we carried out a pilot study, using the data available in DBBE, to determine whether we can detect verses and epigrams that have been manually labelled as similar. This pilot study illustrates one possible reuse scenario for the benchmark: the automatic detection of semantically related verses and epigrams in large textual databases.
This pilot is conducted on data derived from the DBBE, which encodes relationships between textual instances at multiple levels of granularity. The DBBE provides two mechanisms for grouping similar textual material: Verse Variants, which group verses with identical or near-identical content, and Types, which group semantically or formally related occurrences. Despite certain limitations, like the focus on orthography, the DBBE offers a valuable opportunity to evaluate our methods on a complex, real-world corpus. Moreover, both the verse variants and the Types–Occurrences groupings have previously been used in similarity-based research, primarily at the orthographic level (Deforche et al., 2024). In this study, we replicate those experiments using the same datasets, but instead of relying on orthographic similarity, we apply our best-performing semantic similarity measure as established in Section 5.
The data consists of a set of 750 verse variants, and 500 types-occurrences (Deforche et al., 2024). We conducted two experiments that followed the same methodological framework, differing only in the granularity of analysis: one at the verse level and the other at the type level. In the verse-level experiment, pairwise semantic similarity between verses was computed using the best-performing similarity measure identified in Section 5. A range of similarity thresholds was then applied to classify verse pairs as similar or dissimilar, and these predictions were compared against the manually curated verse variants using precision, recall, and F1-score.
The epigram-level experiment followed the same procedure, but similarity was computed between epigrams rather than individual verses. Performance was again evaluated using precision, recall, and F1-score.
The results of both experiments reveal a consistent pattern in the relationship between the similarity threshold and overall performance, as shown in Figures 3 and 4. At lower thresholds, recall remains high because more verse pairs are classified as similar, albeit at the cost of lower precision due to increased false positives. This results in modest F1-scores. As the threshold increases, precision improves steadily, while recall gradually decreases. The verse-level experiment reached its highest F1-score of 0.91 at a threshold of 0.6, with a precision of 0.95 and recall of 0.87 (Table 8). Similarly, the type-level experiment achieved its best F1-score of 0.80 at the same threshold (0.6), with a precision of 0.84 and recall of 0.77 (Table 9). These results suggest that the semantically informed edit distance is effective at identifying similar textual material at both levels of granularity, with optimal performance occurring at a mid-range threshold that balances precision and recall.

Figure 3
Plot of the precision/recall trade-off, displaying the ideal threshold for detecting semantic textual similarity on verse level.

Figure 4
Plot of the precision/recall trade-off, displaying the ideal threshold for detecting semantic textual similarity on epigram level.
Table 8
Verse variants experiments.
| THRESHOLD | PRECISION | RECALL | F1 |
|---|---|---|---|
| 0.05 | 0.144942 | 0.999065 | 0.253157 |
| 0.1 | 0.144942 | 0.999065 | 0.253157 |
| 0.15 | 0.145053 | 0.999016 | 0.253324 |
| 0.2 | 0.145664 | 0.99781 | 0.254216 |
| 0.25 | 0.14729 | 0.993381 | 0.256543 |
| 0.3 | 0.153433 | 0.990182 | 0.265695 |
| 0.35 | 0.16642 | 0.986245 | 0.284785 |
| 0.4 | 0.199558 | 0.98248 | 0.331735 |
| 0.45 | 0.263403 | 0.971629 | 0.414451 |
| 0.5 | 0.441683 | 0.952879 | 0.603588 |
| 0.55 | 0.782899 | 0.920792 | 0.846265 |
| 0.6 | 0.948479 | 0.873819 | 0.909619 |
| 0.65 | 0.98392 | 0.800984 | 0.883077 |
| 0.7 | 0.994334 | 0.690969 | 0.815348 |
| 0.75 | 0.99956 | 0.558366 | 0.716491 |
| 0.8 | 1 | 0.411811 | 0.58338 |
| 0.85 | 1 | 0.268996 | 0.423951 |
| 0.9 | 1 | 0.162992 | 0.280298 |
| 0.95 | 1 | 0.075787 | 0.140897 |
Table 9
Epigram level experiments.
| THRESHOLD | PRECISION | RECALL | F1-SCORE |
|---|---|---|---|
| 0.05 | 0.080121 | 0.996222 | 0.148313 |
| 0.1 | 0.083577 | 0.993363 | 0.154183 |
| 0.15 | 0.091631 | 0.981621 | 0.167615 |
| 0.2 | 0.107101 | 0.966612 | 0.192836 |
| 0.25 | 0.13508 | 0.956912 | 0.236741 |
| 0.3 | 0.178067 | 0.944354 | 0.299636 |
| 0.35 | 0.256167 | 0.930978 | 0.40178 |
| 0.4 | 0.322862 | 0.903512 | 0.475727 |
| 0.45 | 0.380734 | 0.874515 | 0.530505 |
| 0.5 | 0.568682 | 0.836533 | 0.677079 |
| 0.55 | 0.755732 | 0.797631 | 0.776116 |
| 0.6 | 0.840879 | 0.766183 | 0.801795 |
| 0.65 | 0.873855 | 0.730651 | 0.795863 |
| 0.7 | 0.888342 | 0.68154 | 0.77132 |
| 0.75 | 0.925937 | 0.612722 | 0.73745 |
| 0.8 | 0.967323 | 0.531958 | 0.68643 |
| 0.85 | 0.989766 | 0.454258 | 0.622717 |
| 0.9 | 0.99602 | 0.35777 | 0.526442 |
| 0.95 | 0.998206 | 0.227282 | 0.370259 |
6 Conclusion
This study presented the first benchmark for semantic textual similarity in ancient Greek, focusing on Byzantine book epigrams, an understudied yet richly rewarding corpus. Through a carefully designed annotation campaign based on pairwise comparison and involving domain experts, we constructed a dataset that captures graded similarity judgments at the verse level. This benchmark not only offers a resource for evaluating computational models but also serves as a reference point for broader questions in literary analysis, manuscript studies, and historical linguistics. In addition to the creation of the benchmark, we conducted a range of experiments at both the word and verse level, comparing static and contextual embeddings, and exploring the impact of lemmatisation on static embeddings. These experiments reveal that while contextual embeddings like Sentence-BERT yield encouraging results, traditional approaches such as Skip-Gram on lemmatised text remain competitive. The semantically informed edit distance method further demonstrated robust performance, especially at the epigram level in the case study, where precision and recall could be balanced effectively by tuning the similarity threshold.
However, not all results were straightforward. In particular, the interpretation of certain similarity scores is complicated by cases where the same verse pair was annotated in both directions, occasionally resulting in divergent scores. As the benchmark expands and stabilises, some of the present findings and interpretations will need to be revisited. This underlines the importance of continued annotation efforts to ensure a more robust basis for evaluating and refining computational models.
The work underscores the value of integrating computational techniques with philological expertise. By combining modern NLP methods with historically aware annotation, we move toward tools that are both technically robust and attuned to the complexities of ancient texts. Future research will focus on completing the benchmark, investigating the sources of annotation divergence, and exploring techniques such as sense-aware embeddings and cross-lingual transfer to address the unique challenges posed by historical and low-resourced languages.
Notes
[1] In this paper, the term Ancient Greek is used in a broad philological sense to refer to the historical stages of Greek preceding Modern Greek. We acknowledge that Byzantine Greek is often treated as a distinct period in more traditional periodisations; however, within computational language processing, these stages are frequently considered jointly due to substantial linguistic continuity and shared resource requirements.
[2] For clarity, we illustrate the principles using three objects; the same approach can be extended to larger sets.
[3] For a detailed description of the algorithm and its relaxation strategies, we refer to Deforche et al. (2024).
AI Declaration
Generative AI is used for manuscript refinement.
Author Contributions
Colin Swaelens: data curation, formal analysis, methodology, visualisation, writing — original draft; Maxime Deforche: software; Ilse De Vos: supervision, writing — review & editing; Els Lefever: supervision, writing — review & editing.
