
Figure 1
Screenshot of the pilot study set-up.
Table 1
Cohen’s kappa between all ten annotators.
| a | b | c | d | e | f | g | h | i | j | |
|---|---|---|---|---|---|---|---|---|---|---|
| a | - | 0.3571 | 0.3571 | 0.3571 | 0.5199 | 0.0323 | 0.3077 | 0.3571 | 1 | 0,3571 |
| b | 0.3571 | - | 1 | 1 | 0.7391 | 0.5199 | –0.125 | 1 | 0.3571 | 1 |
| c | 0.3571 | 1 | - | 1 | 0.7391 | 0.5199 | –0.125 | 1 | 0.3571 | 1 |
| d | 0.3571 | 1 | 1 | - | 0.7391 | 0.5199 | –0.125 | 1 | 0.3571 | 1 |
| e | 0.5199 | 0.7391 | 0.7391 | 0.7391 | - | 0.3571 | 0.0323 | 0.7391 | 0.5199 | 0.7391 |
| f | 0.0323 | 0.5199 | 0.5199 | 0.5199 | 0.3571 | - | 0.1724 | 0.5199 | 0.0323 | 0.5199 |
| g | 0.3077 | –0.125 | –0.125 | –0.125 | 0.0323 | 0.1724 | - | –0.125 | 0.3077 | –0.125 |
| h | 0.3571 | 1 | 1 | 1 | 0.7391 | 0.5199 | –0.125 | - | 0.3571 | 1 |
| i | 1 | 0.3571 | 0.3571 | 0.3571 | 0.5199 | 0.0323 | 0.3077 | 0.3571 | - | 0.3571 |
| j | 0.3571 | 1 | 1 | 1 | 0.7391 | 0.5199 | –0.125 | 1 | 0.3571 | - |

Figure 2
A screenshot of the web application used for the pairwise annotation. This shows the first three out of ten pairs presented on the page.
Table 2
Quantitative overview of the released benchmark and annotation procedure.
| COMPONENT | SIZE |
|---|---|
| Reference verses | 3 |
| Candidate verses per batch | 99 |
| Comparison pairs per batch | 4,851 |
| Duplicated pairs (agreement estimation) | 1,448 |
| Total annotated comparison pairs | 18,897 |
| Released benchmark entries | 296 |
Table 3
The results of the inter-annotator agreement study of the three reference verses. Preference 1 is the choice of annotator 1, preference 2 is the choice of annotator 2.
| REFERENCE | VERSE 1 | VERSE 2 | PREFERENCE 1 | PREFERENCE 2 |
|---|---|---|---|---|
| 16977 1 | 16977 2 | 17148 1 | 17148 1 | 17148 1 |
| 16977 2 | 16977 1 | 17148 1 | 16977 1 | 16977 1 |
| 17148 1 | 16977 1 | 16977 2 | 16977 2 | 16977 2 |
Table 4
The similarity scores between the verse pairs of the reference verses.
| 16977 1 | 16977 2 | 17148 1 | |
|---|---|---|---|
| 16977 1 | / | 0.233610 | 0.118850 |
| 16977 2 | 0.954545 | / | 0.123377 |
| 17148 1 | 0.646825 | 0.599206 | / |
Table 5
The 5 most similar verses to verse 16977 1 and the 5 least similar verses according to our annotations.
| OCCURRENCE & VERSE | TEXT | SIMILARITY SCORE |
|---|---|---|
| 20963 1 | τῷ θ(ε)ῷ ἡμῶν δόξα τῷ δόντι ἀρχὴν καὶ τέλος:+ | 1.0 |
| 22220 1 | δόξα τω θ[ε]ῶ δόντι καὶ ἀρχὴν (καὶ) τέλος˙ | 1.0 |
| 18150 1 | πῆναξ σὺν θεῶ ἁγ(ίας) τῆς βίβλου ταύτης | 0.86528 |
| 18034 1 | + τέλος σὺν θ(ε)ῶ τοῦ δευτέρου βιβλίου:+̃ | 0.82383 |
| 18886 1 | Ἀρχῆς καλλῆς κάλλιστον εἶναι καὶ τέλος | 0.79792 |
| … | ||
| 24479 1 | Εἴληφε τέρμα τῶν ἑπτὰ ἐπί θήβας | 0.04663 |
| 33429 1 | τέλος τῆς γάμα ὁμήρου ῥαψῳδίας | 0.04145 |
| 23122 6 | τοίγαρ ἀντολίηθεν ἐωσφόροι ἀμφανέντες | 0.04145 |
| 34901 3 | τῶν συλλαβῶν δὲ τὴν ἀρίθμησιν κύκλον | 0.02591 |
| 20839 1 | Τέλος τοῦ τρίτου δράματος σοφοκλέους | 0.02073 |
Table 6
Root Mean Squared Error of the static word embeddings tested on our gold standard.
| CBOW | SKIP-GRAM | GLOVE | BERT | ||||
|---|---|---|---|---|---|---|---|
| TOKEN | LEMMA | TOKEN | LEMMA | TOKEN | LEMMA | TOKEN | |
| RMSE | 0.2768 | 0.0782 | 0.3326 | 0.3106 | 0.0524 | 0.0815 | 0.2904 |
| Spearman | 0.2138 | 0.2701 | 0.1988 | 0.1063 | 0.1563 | 0.2686 | 0.0504 |
| p-value | 0.0002 | 0 | 0.0006 | 0.0679 | 0.5228 | 0.0003 | 0.3544 |
Table 7
Root Mean Squared Error of the static verse embeddings tested on our gold standard.
| CBOW | SKIP-GRAM | GLOVE | BERT | ||||||
|---|---|---|---|---|---|---|---|---|---|
| TOKEN | LEMMA | TOKEN | LEMMA | TOKEN | LEMMA | CLS | MEAN | SBERT | |
| RMSE | 0.3961 | 0.2146 | 0.5745 | 0.3343 | 0.1273 | 0.2028 | 0.6819 | 0.3181 | 0.3181 |

Figure 3
Plot of the precision/recall trade-off, displaying the ideal threshold for detecting semantic textual similarity on verse level.

Figure 4
Plot of the precision/recall trade-off, displaying the ideal threshold for detecting semantic textual similarity on epigram level.
Table 8
Verse variants experiments.
| THRESHOLD | PRECISION | RECALL | F1 |
|---|---|---|---|
| 0.05 | 0.144942 | 0.999065 | 0.253157 |
| 0.1 | 0.144942 | 0.999065 | 0.253157 |
| 0.15 | 0.145053 | 0.999016 | 0.253324 |
| 0.2 | 0.145664 | 0.99781 | 0.254216 |
| 0.25 | 0.14729 | 0.993381 | 0.256543 |
| 0.3 | 0.153433 | 0.990182 | 0.265695 |
| 0.35 | 0.16642 | 0.986245 | 0.284785 |
| 0.4 | 0.199558 | 0.98248 | 0.331735 |
| 0.45 | 0.263403 | 0.971629 | 0.414451 |
| 0.5 | 0.441683 | 0.952879 | 0.603588 |
| 0.55 | 0.782899 | 0.920792 | 0.846265 |
| 0.6 | 0.948479 | 0.873819 | 0.909619 |
| 0.65 | 0.98392 | 0.800984 | 0.883077 |
| 0.7 | 0.994334 | 0.690969 | 0.815348 |
| 0.75 | 0.99956 | 0.558366 | 0.716491 |
| 0.8 | 1 | 0.411811 | 0.58338 |
| 0.85 | 1 | 0.268996 | 0.423951 |
| 0.9 | 1 | 0.162992 | 0.280298 |
| 0.95 | 1 | 0.075787 | 0.140897 |
Table 9
Epigram level experiments.
| THRESHOLD | PRECISION | RECALL | F1-SCORE |
|---|---|---|---|
| 0.05 | 0.080121 | 0.996222 | 0.148313 |
| 0.1 | 0.083577 | 0.993363 | 0.154183 |
| 0.15 | 0.091631 | 0.981621 | 0.167615 |
| 0.2 | 0.107101 | 0.966612 | 0.192836 |
| 0.25 | 0.13508 | 0.956912 | 0.236741 |
| 0.3 | 0.178067 | 0.944354 | 0.299636 |
| 0.35 | 0.256167 | 0.930978 | 0.40178 |
| 0.4 | 0.322862 | 0.903512 | 0.475727 |
| 0.45 | 0.380734 | 0.874515 | 0.530505 |
| 0.5 | 0.568682 | 0.836533 | 0.677079 |
| 0.55 | 0.755732 | 0.797631 | 0.776116 |
| 0.6 | 0.840879 | 0.766183 | 0.801795 |
| 0.65 | 0.873855 | 0.730651 | 0.795863 |
| 0.7 | 0.888342 | 0.68154 | 0.77132 |
| 0.75 | 0.925937 | 0.612722 | 0.73745 |
| 0.8 | 0.967323 | 0.531958 | 0.68643 |
| 0.85 | 0.989766 | 0.454258 | 0.622717 |
| 0.9 | 0.99602 | 0.35777 | 0.526442 |
| 0.95 | 0.998206 | 0.227282 | 0.370259 |
