1. Introduction
A poem and the aesthetic experience it evokes in humans are inseparable companions. The neural responses accompanying aesthetic evaluation and aesthetic emotions are related to evolutionary basic emotions (Panksepp 1998) and presumably shared among our species. Against this universally common cognitive equipment, the linguistic differences language usage in poetry draws on and the cultural aesthetic preferences it is embedded in introduce undeniable differences. This creates a diversity of poetic forms that impose severe limits on cross-linguistic and cross-cultural comparison of the many facets of verbal art and the aesthetic evaluation and emotion it evokes. Given this diversity, it seems almost impossible for us to ever assess and compare the aesthetic value and function of verbal art produced in non-native languages and embedded within different cultural backgrounds, let alone the emotions experienced by contemporary listeners of poetry performed millennia ago in ancient languages, e.g., Homeric hexameters, or Zarathustra’s gāthās. However, an intuitive and straightforward approach can be to use translations. Translations are often the only access to poetry in particular of underresourced and/or extinct languages and distant cultures. Implementing and evaluating appropriate strategies to transfer the structural and semantic features and aesthetic and cultural values of poetry across languages is a central topic in translatology (Lefevere 1975; Boase-Beier 2004, 2015; Jones 2011). But the question remains of how we can assess whether, for example, a native speaker of Chinese is exposed to a similar aesthetic experience when reading translated Russian poetry as a Russian native speaker would. While there are studies that explore cross-linguistic and cross-cultural differences in responses to poetry (e.g. Chesnokova et al. 2017; Chesnokova et al. 2009), there are—as far as we can tell—practically no empirical studies exploring the full potential of poetry translations from a comparative perspective on aesthetic experience. We approach this question by assessing the degree to which translations preserve the aesthetic potential of the original.
Comparing structural and semantic properties of poetry originals and translations has been a huge challenge for centuries. Recently developed methods allow for the computation of fidelity measures and demonstrate the success of AI translations (Gao et al. 2024; Karaban and Karaban 2024). However, while such holistic measures provide an estimate of the structural and semantic closeness between original and translation, they cannot capture the degree to which readers’ experiences of aesthetic pleasure when reading a translation comply with that of linguistic and cultural “natives” when reading the original. And it remains unclear to what extent AI translations approximate, i.e., faithfully preserve, the aesthetic and emotional effects of their originals and how they perform compared to human translations.
Taking up Roman Jakobson’s analytical approach to poetry (Jakobson and Lévi-Strauss 1962; Fechino, Jacobs, and Lüdtke 2020; but see Riffaterre 1966), modern computational linguists investigate more or less ‘hidden’ text features, such as the type-token ratio or the adjective-verb quotient, that potentially drive affective and aesthetic responses to poetry and literature in general. There exist several extensive feature lists indicating promising candidates for the prediction of aesthetic reader responses to poetry (e.g., Leech and Short 2013; Simonton 1990; Dalvean 2015; Jacobs 2019, 2023). One of the first to empirically test the validity of these features, Jacobs (2023) transformed each line of each of the 154 Shakespearean sonnets into a 20-dimensional vector of experimentally validated features computed by the text analysis tool SentiArt2.0 (Jacobs, 2019, https://github.com/matinho13/SentiArt). On that basis he was able to predict individual readers’ ratings of the most striking line in each sonnet with very high accuracies between .94 and .97.
So far, these analytical approaches are confined to poems in the language of the originals. In this paper, we extend the potential of the predictive approach in order to probe into the value of translations for accessing the aesthetic and emotional effects of their originals. A first step into this direction was already taken in Rybová et al.’s (2025) proof-of-concept study on eusemy (the beauty of words) and euphony (the beauty of sounds) in translations of poetry from and into Russian, Polish, Czech, and Slovak. They show that across closely related languages sharing a highly similar cultural embedding, this approach returns sensible results. Extending this line of research, we consider more determinants of aesthetic potential and focus on two highly appreciated poems with a linguistically and culturally more distant background: Shakespeares’s Sonnet 27 and Pushkin’s Ja vas ljubil, and their translations into German and Russian, and German and English respectively. Sonnet 27 has been the subject of extensive close and distant reading (i.e. computational) as well as empirical studies (Jacobs et al. 2017; Papp-Zipernovszky et al. 2021; Xue, Jacobs, and Lüdtke 2020). Pushkin’s poem has been much discussed as a showcase of poetry devoid of imagery, i.e. figurative language, and relying on grammatical means instead (Jakobson 1979; Chesnokova and Peer 2020). By the differences in the importance of meaning vs. structure related features, these two poems pose different challenges for keeping their aesthetic and emotional impact across languages. This makes them an ideal test case for our approach to assess the faithfulness of both human and AI (machine) translations.
Computational stylometry analysis has shown that AI translations can closely mirror human translations in lexical, syntactic and content features (Yao, Kang, and McCosker 2025; Hiebl and Gromann 2023), but little is known about the faithfulness of stylistic devices in AI poetry translations. We thus combine our method to quantify the aesthetic faithfulness between original and translation in terms of empirically verified determinants of the perceived beauty of poetry with qualitative inspection of stylistic devices. This combination of computational quantification and philological interpretation yields a middle-reading approach that allows us to explore the usefulness of human and AI translations for accessing the aesthetic value of the originals.
We introduce our data basis, i.e. the poems and translations we consider, the features and methods we apply to quantify the aesthetic and emotional faithfulness of human and AI translations in Section 2. In addition to the description of data used (Sect. 2.1) and quantification of features (Sect. 2.2), we describe the methodology (title prediction) applied for validating the features and their contributions (Sect. 2.3); Sect. 2.4 presents the close-reading approach used in this investigation. The results are reported in Section 3 and discussed in Section 4, which also highlights the prospects of our middle-reading approach. Section 5 offers a brief conclusion and situates our work in a larger context.
2. Data and methods
We quantify faithfulness in terms of correlations between original and translation (Section 2.1) along six features representing six dimensions of aesthetic space (Section 2.2) and validate their reliability by means of title prediction (Section 2.3). Expanding on what Gambino et al. (2020) investigated for Sonnet 27, we apply a qualitative close reading to explore the distribution of stylistic devices in Pushkin’s poem Ja vas ljubil and its translations (Section 2.4).
2.1 Poems
We analyzed two poems, Shakespeare’s Sonnet 27 and Pushkin’s Ja vas ljubil, and a convenience sample of translations thereof into German and Russian, and German and English, respectively, including one machine translation per language, obtained from Google Translate on 18 October 2025, see (1) and (2). We chose freely available, custom Google translate even though it is known to be outperformed by other systems (e.g. Gao et al. 2024), since this paper is not about comparing the performance of different machine translation models.
Shakespeare: Sonnet 27 (English)
Pushkin: Ja vas ljubil (Russian)
While both poems are about love, describing it as tension and longing, they differ in the linguistic means used to evoke this emotion. Shakespeare’s Sonnet 27 is the first of the rather quiet and meditative ‘travel sonnets’ where the poet writes about the agony of being separated from his friend on a journey. Employing rich imagery, i.e. figurative language, it develops the metaphor of love, addressing various emotional states such as feelings of longing (passion, love), tiredness (exhaustion) and anxiety (discomfort, frustration, distress). Its scenic rather than narrative drama about night, restlessness, and jealousy offers few events but numerous fresh scenes. The metaphor of a journey is the dominant poetic device; other stylistic figures, such as repetition of various kinds, are employed in a comparatively sparse number and arranged with little overlap. This yields a fairly regular poetic texture (Gambino et al. 2020).
At first sight, Pushkin’s Ja vas ljubil likewise seems to be characterised by the absence of major deviations from the structure of everyday language (Chesnokova and Peer 2020). It is nearly devoid of imagery, uses only five nouns and no attributive adjectives or modifiers. However, it abounds in grammatical figurativity, induced in particular by a specific arrangement of pronouns (‘I’, ‘you’, ‘the other’) as the sole means to refer to the dramatis personae, a characteristic usage of verbal aspect and a figura etymologica based on the root ljub- ‘love’ (Jakobson 1979, 244–48). This masterly interplay of grammatical forms and syntactic arrangements, in combination with a lack of imagery, leads us to expect a more complex layering of stylistic devices than identified for Sonnet 27. What is more, the prevalence of morphological and syntactic stylistic devices, i.e. devices closely tied to the grammatical system of Russian, leads us to expect considerable challenges for keeping the aesthetic and emotional impact in translation.
2.2 Quantification
The language-specific models used for quantification differ with respect to their empirical validation: The English and German models are empirically validated and proved predictive for human ratings of, e.g., Shakespearean poetry or stories by E.T.A. Hofmann (Jacobs 2023). The Russian model includes the label lists developed in Rybová et al. (2025) to quantify eusemy. These lists correspond to the lists used for English and German, but lack empirical validation. Based on these models, we contrast three different combinations of original and translation language depending on the availability of empirically validated models: a) models of both original and translation language are empirically validated (English and German); b) the model of the language of the original is empirically validated, the model of the translation language is not (English and Russian); c) the model of the language of the original is not empirically validated, the model of the translation language is empirically validated (Russian and German/English).
As for features, Rybová et al. (2025) demonstrate and integrate a subset of the text features identified by Jacobs (2023) as being important for shaping reader responses to poetry. In our study, we exclude euphony features because potential differences between the Russian, German, and English phonetic systems likely are difficult to interpret. In particular, many factors that contribute to euphony depend on the grammar of the languages and are therefore not entirely at the translator’s power and hence beyond artistic intentions. In short poems this likely biases results. To illustrate, the most common genitive and nominative plural marker in English is s, which usually scores much lower on sonority scales (e.g., Rybová et al. 2025) than the common Russian genitive and nominative plural markers a and i. Moreover, the exact structure of the sonority scale is contested, and it might not be an accurate predictor of sound related aesthetic responses in the first place. In addition, sound features are reported to play a minor role, at least in predicting reader responses to the 154 Shakespeare sonnets (Jacobs 2023).
Five of the features considered here have been shown in Jacobs’ (2023) study of Shakespeare’s sonnets to be predictive in readers’ appreciation of poems, whether consciously or subconsciously. For instance, they have been shown to accurately predict single word valence ratings in different languages (e.g., English, German, Dutch), valence ratings for sentences, entire paragraphs, song lyrics, poems (by different authors and of different types), or stories (in German and English). They are therefore good candidates for analysing the faithfulness of poetry translations. In addition, we consider a novel affective semantic feature, viz. emotion entropy. All features are computed using the text analysis tool SentiArt2.0.
2.2.1 Affective Aesthetic Potential (AAP)
AAP is a measure of the association between a word and a set of special labels for prototypical concepts that represent positive and negative affects and emotions experienced during daily life as well as the reading of literature. It is based on semantic models that convert words into vectors, specifically the word2vec method (Mikolov et al. 2013), and estimates the degree to which each word in a text is affectively and aesthetically positive or negative. Semantic values are based on 60 positively connotated anchor words or labels (e.g. art, beautiful, romantic, etc.) and 60 negatively connotated ones (e.g. ugly, apathy, horror). The anchor words have been empirically validated with human ratings in a number of studies for English, German and Dutch (Jacobs 2023), and adapted to Russian (Rybová et al. 2025). A word has a positive AAP value if its mean association strength with the positive labels is greater than its association with the negative labels. It is zero if the difference is zero. Depending on whether raw, or centered, or standardized/normalized values are used there can be different ranges of variation. What matters are not the absolute values, however, since they depend on e.g. the language model used, the extent of the text etc. Only standardized values (from –3 to +3) are interpretable across texts or languages.
2.2.2 PosNeg Quotient
The quotient of positive to negative words is computed by the AAP measure and is indicative of the overall mood a line, stanza or poem can induce. It thus complements the AAP and other emotional features used here at the meso- and macro levels of analysis. Since the PosNeg Quotient relies on the AAP, it might seem redundant at first sight. However, when it comes to predicting human ratings of poetry using partially redundant features is no wasted effort as long as it remains unclear which of a potentially unlimited set of text features are the most prominent ones in shaping explicit and implicit reader responses (Jacobs 2023). Although the PosNeg Quotient was a less important predictor of human ratings in Jacobs’ (2023) study on Shakespeare than mean AAP, comparing the two across our poem sample and examining their transfer could provide additional evidence for their relevance in shaping the aesthetic potential of poetic texts.
2.2.3 Emotion Entropy (He)
Emotion Entropy (He) is a measure that quantifies the mixed emotion potential of texts, i.e., a text’s more or less hidden power to elicit associations with several basic emotions like fear, joy, or disgust, or with more sophisticated aesthetic emotions like fascination or beauty. He represents the degree to which a text offers a mix of associations with the six basic primary emotions common to all mammals (see Darwin 1872; Ekman 1992), and also identifies the dominant emotion, i.e. the basic emotion which is most frequently and strongly associated with the text’s words. He is a measure inspired by Shannon’s information entropy measure H. In general, the higher normalized He (large uncertainty), the greater the emotion mix with similar frequencies of occurrence for all associated emotions, e.g. in a line of poetry. Thus, in a uniform distribution where each emotion tag occurs with equal frequency, there is maximum uncertainty about the next emotion tag, resulting in an He value of 1 when normalized. Conversely, for distributions where one outcome occurs with much higher frequency than the others, there is less uncertainty, leading to an entropy value closer to 0 when normalized. Computing the correlation between poetic texts of He, we are interested in the extent to which a translation conserves and transports the same dominant emotion and emotion mix as the original. In order to determine He, we employ SentiArt to compute the theoretical associative strength for each content word in a line with each of the six discrete emotion concepts. Thus, for example, the first word weary in line 1 of Sonnet 27 yields the profile of associated basic emotions shown in the upper part of Table 1. The basic emotion with the highest value is sadness and hence marked as the dominant one for this word. On a scale from –1 to +1, the associative strength with sadness (0.18) is not very high, though. This indicates that when reading the word weary this concept is not the first that comes to mind. Indeed, the five concepts most strongly associated with weary as computed by SentiArt are: ‘tired’ (.66), ‘fatigued’ (.52), ‘wearisome’ (.49), ‘tedious’ (.47), and ‘footsore’ (.46). However, numerous studies in experimental psychology and neuroscience suggest that even weakly, subconsciously or preverbally activated mental representations can influence the way we perceive the world or a text. What is hypothesized here is that if the word weary (partially) activates a basic emotion concept at all then it will likely be sadness.
Table 1
He for basic emotional (above) and aesthetic associations (below) for the target word weary.
| EMOTION | VALUE |
|---|---|
| anger | 0.02 |
| disgust | 0.04 |
| fear | 0.13 |
| joy | 0.13 |
| sadness | 0.18* |
| surprise | –0.05 |
| EMOTION | VALUE |
| sick | 0.45* |
| sickness | 0.26 |
| happy | 0.23 |
| ache | 0.23 |
| misery | 0.21 |
Analysing He only for the six basic emotion terms may appear narrow for poetic text materials that apart from appealing to basic emotions surely also can elicit higher or more complex mixed emotions such as empathy or fascination. We thus complement it by an analysis looking at a wider range of affective-aesthetic concepts that can be associated with the words in a poem. For this, we used the 120 labels of the AAP computation discussed above and identified those concepts most strongly associated with each content word and each line, just as for the six basic emotion terms. For the word weary, as an example, this extended analysis yields the AAP profile shown in the lower part of Table 1. The He feature was then computed across these combined terms.
This computational process was repeated for each content word and each line, and the word-based dominant emotions for each line were counted. Thus, for each line the algorithm notes a series of dominant emotions e.g. [surprise, sadness, fear, fear, fear], and the emotion most often encountered in the entire poem then is identified as the dominant one. In another step, SentiArt computes a poem’s He index based on the distribution of the basic and aesthetic emotions in each line, by taking the average across all lines.
As an extreme example, imagine a line yielded the following distribution of associated basic and aesthetic emotions: [‘sadness’, ‘sadness’, ‘sadness’, ‘sadness’]. The corresponding He would be 0, i.e. theoretically, for a reader it should be quite easy to predict the next (associated) emotion from the one experienced with the preceding content. In contrast, if we find a perfectly heterogeneous distribution, as is actually the case for line 6 from Benediktov’s translation of Sonnet 27 (Benediktov 1884), we get a different picture altogether with: [‘disgust’, ‘sadness’, ‘tenderness’, ‘stink’] yielding an He of 1. Benediktov’s (1884) overall (mean) He of .97 suggests that the translated poem really offers a rather varied mix of associated negative and positive emotions, i.e. readers may experience somewhat of an emotional roller-coaster. Still, the overall dominant emotion with a maximum frequency of occurrence is sadness. What matters here, however, is not so much the absolute value of He, but how it compares to the values of the original and other poems and translations.
Much as the data in Table 1, this example shows that mixed or even contrasting emotions become a psychological possibility already at the single word level, just as had already been indicated by Jung’s results of his famous word association test and many other more recent word association studies (e.g. Hofmann et al. 2018). According to these data the word weary very likely first activates an overall negative concept (sadness or sickness), but this can be ‘water-coloured’ by a positive concept like happiness. Words not only often have multiple (dominant and subdominant) meanings, but they also have multiple basic affective and aesthetic meanings hidden under the dominant one, but with the potential to colour it, or even pop out into the mind’s eye when for instance embedded in an artful context.
2.2.4 Eigensimilarity
Eigensimilarity represents the average semantic similarity between the lines of a stanza, proposed by Jacobs (2018) to evaluate the cohesion of poetry. Following Graesser et al.’s (2004) idea of a family resemblance score, i.e., a metric that computes how similar a sentence is to all other sentences in the text, a higher Eigensimilarity increases the perceived cohesion of a text and thus its potential comprehensibility. To quantify Eigensimilarity, first a semantic vector is computed for each line (using e.g. the German BERT model ‘dbmdz/bert-base-german-uncased’ for German texts), and then the cosine similarity with the vectors for all other lines is computed. The average value for a poem then is its Eigensimilarity value. Concerning reader experience, comprehensibility in terms of high Eigensimilarity alone might not contribute to an increased level of pleasure but be perceived as boring. However, we are not interested in Eigensimilarity per se, but rather in Eigensimilarity correlations between original and translations.
2.2.5 PoS Entropy (Hpos)
The Shannon entropy represents variation in parts of speech (PoS), specifically their ordering (see .py code below). For each line the frequency distribution for all PoS tags is computed (pos_freq) and fed into the code below. Again, much as for the He feature, if a line say had a distribution like [NOUN NOUN NOUN NOUN] its Hpos value would be 0. It would be 1 e.g. for [DET ADJ NOUN VERB]. Jacobs (2023) proposed this measure to test whether a more complex syntactic pattern with greater variation in PoS order, i.e. increased PoS entropy, increases the aesthetic appreciation of Shakespeare sonnets which are full of highly varied syntactic patterns and a huge creativity regarding word order and word type relations. The results of Jacobs’ (2023) analysis indicated that at least for some readers this was the case. PoS annotation was carried out with the simplified universal dependencies PoS tagger from the spacy library for all three languages considered here with a limited number of clear categories, in particular nouns, verbs, adjectives and pronouns.
def pos_entropy(pos_freq):
total_count = sum(pos_freq.values())
entropy = 0
for count in pos_freq.values():
probability = count/total_count
entropy – = probability * math.log2(probability)
max_entropy = math.log2(len(pos_freq))
normalized_entropy = entropy/max_entropy
return normalized_entropy
2.2.6 Proportion of Nouns
Measuring the Proportion of Nouns is based on Kao and Jurafsky’s (2012) finding that the proportion of concrete object words was the best predictor out of 16 text features for distinguishing between amateur and professional poetry (4.1% for professional vs. 1.5% for amateur poetry). This suggests that professional poets follow the maxim of ‘show, don’t tell’ and let images instead of words convey emotions, concepts, and experiences that stick to readers’ minds. Since Dalvean (2015) replicated this result using a much bigger set of 98 features, we include this feature which was dear to many poets such as Keats (1958). All nouns are counted independently of their concreteness values.
2.2.7 Computing correlations
We use SentiArt to compute the contribution of the sampled features to the aesthetic and emotional perception in the original and in the translations. To compare the j-th translation Tj to its original Oi we compute the linewise Pearson’s correlation coefficient ρ for all pairs of Tj, i and Oi. We use correlations instead of comparing absolute values since the values are based on different LLMs. Thus the absolute values cannot be compared, while correlations can. The correlation values range from –1 to +1. A value of +1 implies that the relationship between two poems Pi and Pj is perfectly linear with all values, i.e., linewise AAP scores, lying on a line for which Pi increases as Pj increases. A value of –1 implies a perfect linear relationship where Pi increases while Pj decreases. A value of 0 implies that there is no linear dependency between the lines of two poems. We visualise the results by means of correlation heatmaps.
2.3 Feature verification
The six features employed in our analysis are all established as contributing to the aesthetic and emotional perception of poetry and the measures applied are supported by experimental verification (see Jacobs 2023 for summary). However, they have not yet been applied to the specific poems and translations we investigate here. We therefore first test the reliability of the feature values in predicting the title (authorship) of the poem. In the context of neurocomputational poetics, predictive modeling is a technique that with the help of machine learning methods can be used to predict some target (outcome) variable, such as the authorship of texts or a reader response to texts (e.g. ratings) on the basis of text features as analysed, e.g., by SentiArt (Jacobs 2023). In a second step it can then be used to sort out which features are the most important ones in predicting the outcome. A predictive model typically is a neural network or other machine learning tool, like so-called support vector machines or decision trees, that is trained in a supervised learning procedure on some data (the training set, usually 70–90% of all data) to predict some other data (the test or validation set, usually 10–30%). What is needed for the supervised learning is an adequate training corpus that contains both the input data, e.g. lines of text, and the output or response data, e.g. a label classifying these lines as either ‘literal’ or ‘figurative’, or some human behavioural response to the lines such as a liking rating, or even some neuroimaging measures such as the brain activity elicited by a line of text. Here we applied this technique to predict the title (authorship) of the poems. Lacking any behavioral outcome data (e.g., reader ratings), this analysis was meant as a basic validity check of our approach and to see which features come out as more important than others. If the six features used here were any good, they should allow us to identify the titles/authors of the selected texts. The more accurate the prediction, the more faith we could have in our feature set to be used and extended in future studies.
2.4 Close reading
Our six hidden text features do not exhaust the aesthetic and emotional potential of poetic work. Reader responses are also shaped by foregrounding features such as repetition, chiasms, alliteration, or play of words (Jacobs 2015, 2023). We thus complement our quantitative measures by (qualitative) close reading, applying the Foregrounding Assessment Matrix (FAM) proposed in Gambino et al. (2020). The FAM allows us to identify stylistic devices at the sublexical, lexical, inter- and supra-lexical levels in the domains of phonology, morpho-syntax and rhetoric. It provides a principled way of comparing poems in terms of the application and distribution of stylistic devices, and to experimentally assess related reader responses, as shown by Gambino et al. (2020) and Papp-Zipernovszky et al. (2021) for three Shakespearean sonnets, among them Sonnet 27.
In our study, we use the FAM to compare original and translation in terms of stylistic features, focusing on Ja vas ljubil. In identifying and annotating the original’s poetic devices, we implemented the results of previous poetic analyses (Jakobson 1979; Breschinsky 1974; Chesnokova and Peer 2020) and applied the same criteria and principles in the annotation of translations. Pushkin’s poem is nearly devoid of imagery, such as metaphors and personifications. Its aesthetic and emotional effects are triggered by a specific employment of morphological forms and syntactic arrangements (Jakobson 1979). This high density of grammatical features is expected to cause considerable difficulties for a faithful transfer, in particular into languages with a different morphological and syntactic makeup. In fact, translations of Ja vas ljubil are often considered inadequate, lacking the original’s effects (for an overview see Chesnokova and Peer 2020).
3. Results
3.1 Correlations of aesthetic dimensions
3.1.1 Aesthetic affective potential (AAP)
Tables 2 and 3 report the correlation coefficients for AAP values for the translations of both poems. Correlation coefficients are generally higher for translations of Ja vas ljubil than for translations of Sonnet 27.
Table 2
Linewise Pearson correlation between AAP values of the original of Shakespeare’s Sonnet 27 and its translations.
| TRANSLATION | LANGUAGE | CORRELATION COEFFICIENTS |
|---|---|---|
| Benediktov (1884) | Russian | 0.63 |
| Google Translate (2025d) | Russian | 0.44 |
| Google Translate (2025c) | German | 0.38 |
| Gildemeister (1876) | German | 0.37 |
| Nabokov (1930) | Russian | 0.34 |
| Tieck (1992) | German | 0.34 |
| Bodenstedt (1866) | German | 0.27 |
| Kaljavina (2015) | Russian | 0.25 |
| Kraus (1964) | German | 0.01 |
| Ilʹin (1902) | Russian | –0.02 |
| George (2008) | German | –0.04 |
| Stepanov (1999) | Russian | –0.17 |
Table 3
Linewise Pearson correlation between AAP values of the original of Pushkin’s Ja vas ljubil and its translations.
| TRANSLATION | LANGUAGE | CORRELATION COEFFICIENTS |
|---|---|---|
| Rainbow (2023) | English | 0.86 |
| Shine (2020) | English | 0.85 |
| Wilke (2020) | German | 0.81 |
| Google Translate (2025a) | English | 0.78 |
| Kust (2021) | German | 0.75 |
| Google Translate (2025b) | German | 0.74 |
| Fiedler (1907) | German | 0.74 |
| Kallsen (2022) | English | 0.69 |
| Sedova (2022) | English | 0.59 |
| Jahnke (2019) | German | 0.53 |
| Avrutin (2023) | German | 0.51 |
| Kushanov (2022) | English | 0.49 |
3.1.2 PosNeg Quotient
As expected, PosNeg Quotient and AAP are highly correlated across all poems in our two sets (R2 = .7 for both Shakespeare and Pushkin). However, as can be seen by comparing the data in Tables 2 to 5, notable differences between AAP and PosNeg Quotient emerge. When considering the correlation between the two measures only for the original and its 12 translations, R2 drops to .48 for Pushkin and to .54 for Shakespeare. Thus, both measures are at least partially independent indicators of the affective-aesthetic potential of lines and should be analysed further until more empirical clarity has been gained.
Tables 4 and 5 report the correlation coefficients for PosNeg Quotient for both poems and their translation. One English translation of the Russian original is highly correlated with the Russian original, followed by a German translation of Sonnet 27. Unlike with AAP, there is no clear-cut difference between the poems. Machine translations show moderate to no correlation with the original.
Table 4
Linewise Pearson correlation between PosNeg Quotient values of the original of Shakespeare’s Sonnet 27 and its translations.
| TRANSLATION | LANGUAGE | CORRELATION COEFFICIENTS |
|---|---|---|
| Bodenstedt (1866) | German | 0.45 |
| Google Translate (2025c) | German | 0.34 |
| Gildemeister (1876) | German | 0.31 |
| Nabokov (1930) | Russian | 0.29 |
| Kraus (1964) | German | 0.29 |
| Ilʹin (1902) | Russian | 0.28 |
| Tieck (1992) | German | 0.26 |
| Benediktov (1884) | Russian | 0.19 |
| Google Translate (2025d) | Russian | 0.17 |
| George (2008) | German | –0.19 |
| Kaljavina (2015) | Russian | –0.29 |
| Stepanov (1999) | Russian | –0.34 |
Table 5
Linewise Pearson correlation between PosNeg Quotient values of the original of Pushkin’s Ja vas ljubil and its translations.
| TRANSLATION | LANGUAGE | CORRELATION COEFFICIENTS |
|---|---|---|
| Rain | English | 0.87 |
| Shine (2020) | English | 0.79 |
| Sedova (2022) | English | 0.47 |
| Wilke (2020) | German | 0.32 |
| Kust (2021) | German | 0.26 |
| Google Translate (2025a) | English | 0.13 |
| Jahnke (2019) | German | 0.12 |
| Fiedler (1907) | German | –0.11 |
| Avrutin (2023) | German | –0.21 |
| Kallsen (2022) | English | –0.23 |
| Google Translate (2025b) | German | –0.24 |
| Kushanov (2022) | English | –0.28 |
3.1.3 Emotion Entropy (He)
The data in Table 6 reveal that in terms of emotion entropy, the Russian machine translation is closest to the original Sonnet 27. For Pushkin’s poem (Table 7) it is Kust’s (2021) translation that scores highest. The correlations are not very strong (.4) and only less than 15% of the variance in the originals’ He measures can be accounted for by them. Similarly to PosNeg Quotient, He correlations are generally lower than AAP correlations. However, other translations, e.g. Benediktov (1884), offer even negative linewise correlations (–.3), while this poem’s mean He of .97 across all lines does not really differ much from the original’s mean across lines (.90; see Sect. 2.2).
Table 6
Linewise Pearson correlation between He values of the original of Shakespeare’s Sonnet 27 and its translations.
| TRANSLATION | LANGUAGE | CORRELATION COEFFICIENTS |
|---|---|---|
| Google Translate (2025d) | Russian | 0.38 |
| Bodenstedt (1866) | German | 0.30 |
| Gildemeister (1876) | German | 0.19 |
| Benediktov (1884) | Russian | 0.12 |
| Kraus (1964) | German | 0.04 |
| Kaljavina (2015) | Russian | 0.00 |
| George (2008) | German | –0.06 |
| Nabokov (1930) | Russian | –0.07 |
| Ilʹin (1902) | Russian | –0.07 |
| Tieck (1992) | German | –0.08 |
| Stepanov (1999) | Russian | –0.26 |
| Google Translate (2025c) | German | –0.42 |
Table 7
Linewise Pearson correlation between He values of the original of Pushkin’s Ja vas ljubil and its translations.
| TRANSLATION | LANGUAGE | CORRELATION COEFFICIENTS |
|---|---|---|
| Kust (2021) | German | 0.36 |
| Kallsen (2022) | English | 0.35 |
| Google Translate (2025a) | English | 0.27 |
| Shine (2020) | English | 0.23 |
| Fiedler (1907) | German | 0.15 |
| Avrutin (2023) | German | –0.01 |
| Rainbow (2023) | English | –0.05 |
| Jahnke (2019) | German | –0.09 |
| Wilke (2020) | German | –0.14 |
| Sedova (2022) | English | –0.16 |
| Google Translate (2025b) | German | –0.17 |
| Kushanov (2022) | English | –0.28 |
3.1.4 Eigensimilarity
The Eigensimilarity correlation coefficients between original and translations are shown in Tables 8 and 9. Correlation coefficients are highest for the German machine translation of Sonnet 27 and Fiedler (1907)’s German translation of Ja vas ljubil.
Table 8
Linewise Pearson correlation between Eigensimilarity values of the original of Shakespeare’s Sonnet 27 and its translations.
| TRANSLATION | LANGUAGE | CORRELATION COEFFICIENTS |
|---|---|---|
| Google Translate (2025c) | German | 0.62 |
| Google Translate (2025d) | Russian | 0.41 |
| Benediktov (1884) | Russian | 0.34 |
| Nabokov (1930) | Russian | 0.31 |
| Kaljavina (2015) | Russian | 0.21 |
| Stepanov (1999) | Russian | 0.10 |
| Bodenstedt (1866) | German | 0.08 |
| Kraus (1964) | German | 0.00 |
| Tieck (1992) | German | –0.03 |
| Gildemeister (1876) | German | –0.16 |
| Ilʹin (1902) | Russian | –0.24 |
| George (2008) | German | –0.24 |
Table 9
Linewise Pearson correlation between Eigensimilarity values of the original of Pushkin’s Ja vas ljubil and its translations.
| TRANSLATION | LANGUAGE | CORRELATION COEFFICIENTS |
|---|---|---|
| Fiedler (1907) | German | 0.60 |
| Rainbow (2023) | English | 0.50 |
| Shine (2020) | English | 0.47 |
| Google Translate (2025a) | English | 0.44 |
| Kushanov (2022) | English | 0.43 |
| Jahnke (2019) | German | 0.20 |
| Wilke (2020) | German | 0.18 |
| Sedova (2022) | English | 0.15 |
| Kust (2021) | German | –0.09 |
| Kallsen (2022) | English | –0.12 |
| Avrutin (2023) | German | –0.24 |
| Google Translate (2025b) | German | –0.26 |
3.1.5 PoS Entropy (Hpos)
Mirroring Shakespeare’s creativity in ordering his words is a major challenge for translators of his poetry. The values in Table 10 for Sonnet 27 suggest that in our sample, the Russian Google Translate (2025d) excelled in this respect, featuring the highest Hpos correlation of .76, followed by Gildemeister (1876) with a correlation of .74. As to Pushkin (see Table 11), the highest correlation score of Hpos is achieved by the English machine translation (.89), while the closest human translations achieve a correlation of .77 (both English and German).
Table 10
Linewise Pearson correlation between Hpos values of the original of Shakespeare’s Sonnet 27 and its translations.
| TRANSLATION | LANGUAGE | CORRELATION COEFFICIENTS |
|---|---|---|
| Google Translate (2025d) | Russian | 0.76 |
| Gildemeister (1876) | German | 0.74 |
| Tieck (1992) | German | 0.66 |
| Stepanov (1999) | Russian | 0.60 |
| George (2008) | German | 0.58 |
| Google Translate (2025c) | German | 0.53 |
| Kaljavina (2015) | Russian | 0.49 |
| Kraus (1964) | German | 0.41 |
| Ilʹin (1902) | Russian | 0.25 |
| Bodenstedt (1866) | German | 0.18 |
| Benediktov (1884) | Russian | 0.09 |
| Nabokov (1930) | Russian | –0.11 |
Table 11
Linewise Pearson correlation between Hpos values of the original of Pushkin’s Ja vas ljubil and its translations.
| TRANSLATION | LANGUAGE | CORRELATION COEFFICIENTS |
|---|---|---|
| Google Translate (2025a) | English | 0.89 |
| Jahnke (2019) | German | 0.77 |
| Kallsen (2022) | English | 0.77 |
| Google Translate (2025b) | German | 0.74 |
| Kust (2021) | German | 0.68 |
| Wilke (2020) | German | 0.66 |
| Avrutin (2023) | German | 0.27 |
| Sedova (2022) | English | 0.16 |
| Shine (2020) | English | –0.05 |
| Kushanov (2022) | English | –0.14 |
| Rainbow (2023) | English | –0.36 |
| Fiedler (1907) | German | –0.44 |
3.1.6 Proportion of Nouns
For Sonnet 27 the correlations concerning the Proportion of Nouns are positive for all translations but one, i.e. Nabokov’s (1930) (–.11), see Table 12. Qualitative inspection shows that here, the ordering on the last two lines deviates from the original. This includes a line with four nouns (Lo thus by day my limbs, by night my mind). For Ja vas ljubil, too, all correlations are positive (most 4), with one exception, i.e. the English translation by Kushanov (2022) (–.18, see Table 13). One reason for this might be the translator’s decision to opt for adjectives in line 6 instead of nouns as in the original (jealous vs. revnost’ ‘jealousy’ and shy vs. robost’ ‘timidity’).
Table 12
Linewise Pearson correlation between Proportion of Nouns values of the original of Shakespeare’s Sonnet 27 and its translations.
| TRANSLATION | LANGUAGE | CORRELATION COEFFICIENTS |
|---|---|---|
| Google Translate (2025d) | Russian | 0.76 |
| Gildemeister (1876) | German | 0.74 |
| Tieck (1992) | German | 0.66 |
| Stepanov (1999) | Russian | 0.60 |
| George (2008) | German | 0.58 |
| Google Translate (2025c) | German | 0.53 |
| Kaljavina (2015) | Russian | 0.49 |
| Kraus (1964) | German | 0.41 |
| Ilʹin (1902) | Russian | 0.25 |
| Bodenstedt (1866) | German | 0.18 |
| Benediktov (1884) | Russian | 0.09 |
| Nabokov (1930) | Russian | –0.11 |
Table 13
Linewise Pearson correlation between Proportion of Nouns values of the original of Pushkin’s Ja vas ljubil and its translations.
| TRANSLATION | LANGUAGE | CORRELATION COEFFICIENTS |
|---|---|---|
| Avrutin (2023) | German | 1.00 |
| Google Translate (2025a) | English | 0.89 |
| Rainbow (2023) | English | 0.80 |
| Kust (2021) | German | 0.79 |
| Sedova (2022) | English | 0.75 |
| Jahnke (2019) | German | 0.64 |
| Wilke (2020) | German | 0.64 |
| Google Translate (2025b) | German | 0.57 |
| Shine (2020) | English | 0.50 |
| Kallsen (2022) | English | 0.42 |
| Fiedler (1907) | German | 0.15 |
| Kushanov (2022) | English | –0.18 |
3.1.7 Mean total correlations
Since we are interested in the overall faithfulness of translations to their original, we calculated the mean of the correlations between the six features for both poems and their translations. Mean total correlations per translation range between .50 and –.02 for Sonnet 27 and slightly higher between .57 and .005 for Ja vas ljubil, see Tables 14 and 15. Except for AAP for Pushkin’s Ja vas ljubil (see Section 3.1, Table 3) the features show both positive and negative correlations.
Table 14
Rank order of correlations with the six features for the 12 translations of Shakespeare’s Sonnet 27.
| TITLE | AAP | pnq | He | Hpos | EIGENSIM | NOUNS | mean tot |
|---|---|---|---|---|---|---|---|
| Shakespeare_orig | 1 | 1 | 1 | 1 | 1 | 1 | 1 |
| Google Translate (2025d)_ru | 0.443 | 0.173 | 0.383 | 0.411 | 0.411 | 0.761 | 0.499 |
| Tieck (1992)_de | 0.342 | 0.264 | –0.08 | 0.309 | –0.03 | 0.656 | 0.294 |
| George (2008)_de | –0.04 | –0.19 | –0.06 | 0.722 | –0.24 | 0.583 | 0.249 |
| Google Translate (2025c)_de | 0.38 | 0.336 | –0.42 | –0.07 | 0.617 | 0.531 | 0.219 |
| Nabokov (1930)_ru | 0.344 | 0.293 | –0.07 | 0.235 | 0.312 | –0.11 | 0.206 |
| Gildemeister (1876)_de | 0.368 | 0.308 | 0.189 | –0.16 | –0.16 | 0.741 | 0.188 |
| Kaljavina (2015)_ru | 0.249 | –0.29 | –4e-3 | 0.176 | 0.206 | 0.491 | 0.168 |
| Benediktov (1884)_ru | 0.628 | 0.187 | 0.12 | –0.23 | 0.341 | 0.088 | 0.152 |
| Bodenstedt (1866)_ru | 0.269 | 0.453 | 0.301 | –0.26 | 0.078 | 0.176 | 0.128 |
| Ilʹin (1902)_ru | –0.02 | 0.28 | –0.07 | 0.229 | –0.24 | 0.25 | 0.112 |
| Kraus (1964)_de | 0.009 | 0.286 | 0.037 | –0.2 | 3e-4 | 0.41 | 0.057 |
| Stepanov (1999)_ru | –0.17 | –0.34 | –0.26 | –0.01 | 0.101 | 0.6 | –0.02 |
Table 15
Rank order of correlations with the six features for the 12 translations of Pushkin’s Ja vas ljubil.
| TITLE | AAP | pnq | He | Hpos | EIGENSIM | NOUNS | mean tot |
|---|---|---|---|---|---|---|---|
| Pushkin_orig | 1 | 1 | 1 | 1 | 1 | 1 | 1 |
| Google Translate (2025a)_en | 0.78 | 0.13 | 0.27 | 0.90 | 0.44 | 0.90 | 0.57 |
| Shine (2020)_en | 0.85 | 0.79 | 0.23 | –0.1 | 0.47 | 0.5 | 0.46 |
| Kust (2021)_de | 0.75 | 0.26 | 0.36 | 0.68 | –0.1 | 0.79 | 0.46 |
| Rainbow (2023)_en | 0.86 | 0.87 | –0.1 | –0.4 | 0.5 | 0.8 | 0.44 |
| Wilke (2020)_de | 0.81 | 0.32 | –0.1 | 0.66 | 0.18 | 0.64 | 0.41 |
| Jahnke (2019)_de | 0.53 | 0.12 | –0.1 | 0.77 | 0.2 | 0.64 | 0.36 |
| Sedova (2022)_en | 0.59 | 0.47 | –0.2 | 0.16 | 0.15 | 0.75 | 0.33 |
| Kallsen (2022)_en | 0.69 | –0.2 | 0.35 | 0.77 | –0.1 | 0.42 | 0.31 |
| Google Translate (2025b)_de | 0.74 | –0.2 | –0.2 | 0.74 | –0.3 | 0.57 | 0.23 |
| Avrutin (2023)_de | 0.51 | –0.2 | –0.008 | 0.27 | –0.2 | 1 | 0.22 |
| Fiedler (1907)_de | 0.74 | –0.1 | 0.15 | –0.4 | 0.6 | 0.15 | 0.18 |
| Kushanov (2022)_en | 0.49 | –0.3 | –0.3 | –0.1 | 0.43 | –0.2 | 0.005 |
The most faithful translation for Sonnet 27, the Russian machine translation, surpasses all other translations to a considerable degree, including the German machine translation, which shows a modest score. The most faithful translation for Ja vas ljubil does not have a large lead.
3.2 Title prediction
Title (authorship) prediction to test the reliability of the six features in distinguishing the two poems was carried out following the predictive modeling approach outlined in Jacobs (2023). In short, each line of each poem is transformed into a six-dimensional vector reflecting the average values of our six features. A two-layer Neural Net model was then trained on a subset of the overall 286 lines using a 5-fold cross-validation procedure. JMP® Student Edition 18.2.1 (https://www.jmp.com/en_us/software/predictive-analytics-software.html) was used to run the predictive modeling analyses. For the Neural Net model we used the following parameter set: two hidden layers with 100 and 50 nodes, respectively, and a hyperbolic tan (TanH) activation function. Overall, this six-feature model predicted the correct title of our sample of poems with an accuracy of 99%, as shown in Table 16 with a low misclassification rate (10% of the lines). The Neural Net model allows for the computation of the importance of each of the six features for the overall prediction.
Table 16
Predictive Accuracy of the six feature model in predicting the poem title for each line of both poems.
| MEASURE | VALUE |
|---|---|
| Generalized R2 | 0.99 |
| Entropy R2 | 0.87 |
| Missclassification Rate | 0.10 |
The importance of each feature for this successful prediction is shown in Table 17. Table 18 shows the values for both poems separately. For both poems, Eigensimilarity has the highest effect, followed by Proportion of Nouns. Total effect sizes of these top features are clearly higher for Ja vas ljubil.
Table 17
Importance of each feature for predicting the title of Sonnet 27 and Ja vas ljubil, pooled.
| FEATURE | TOTAL EFFECT |
|---|---|
| Eigensimilarity | 0.70 |
| He | 0.57 |
| Proportion of Nouns | 0.53 |
| PosNeg Quotient | 0.52 |
| AAP mean | 0.48 |
| Hpos | 0.46 |
Table 18
Effects of each feature for predicting Sonnet 27 and Ja vas ljubil.
| FEATURE | SONNET 27 | JA vas ljubil | ||
|---|---|---|---|---|
| MAIN EFFECT | TOTAL EFFECT | MAIN EFFECT | TOTAL EFFECT | |
| Eigensimilarity | 0.072 | 0.634 | 0.082 | 0.774 |
| Proportion of Nouns | 0.038 | 0.565 | 0.051 | 0.701 |
| PosNeg Quotient | 0.033 | 0.546 | 0.020 | 0.458 |
| AAP mean | 0.029 | 0.489 | 0.016 | 0.313 |
| He | 0.019 | 0.433 | 0.016 | 0.425 |
| Hpos | 0.019 | 0.393 | 0.010 | 0.266 |
Clearly, it is the Eigensimilarity (a proxy for cohesion and comprehensibility) that has the largest effect on the prediction for both poems (and thus overall). With total effects around .5, also all other five features are important.
3.3 Poetic devices
In our close-reading investigation, we compared the employment and distribution of poetic devices in Pushkin’s original poem to the German and English human translations with the overall highest and lowest mean total scores (see Table 15), i.e. Kust (2021) and Fiedler (1907) for German, and Shine (2020) and Kushanov (2022) for English. In addition, we examined the English and German machine translations.
For the original poem, the inversion of word order—most prominently in ja vas ljubil lit. ‘I you loved’ instead of unmarked ja ljubil vas ‘I loved you’—constitutes a pervasive structural feature throughout the poem and plays a central role in its syntactic organization (Chesnokova and Peer 2020). The anaphoric repetition of ja vas ljubil ‘I loved you’ is another salient stylistic device. Morphosyntactic parallelism is particularly notable in the second stanza (Chesnokova and Peer 2020), organized in binary patterns across lines 5, 6, and 7. This structures the stanza and enhances its aesthetic symmetry (Breschinsky 1974). The rhythmic structure is further reinforced by a figura etymologica, with derivatives of the root ljub- ‘love’ distributed throughout the poem (Jakobson 1979). Enjambement in the opening and closing lines (lines 1–2, 7–8) frames the poem and creates a sense of cohesion. Overall, the syntactic and morphosyntactic devices employed allow the poem to maintain a register close to everyday language (Chesnokova and Peer 2020), imparting immediacy and emotional intimacy.
The inverted word order of the key phrase of the poem is mirrored in one translation only. This might be due to the syntactic constraints of English and German, which both have a more rigid word order than Russian. In German, the translation by Kust (2021) employs another syntactic deviation that achieves a comparable effect, i.e. subject omission in line 1.
As for rhetorical foregrounding devices, which in Pushkin’s original are limited to one single and commonly used synecdoche, i.e. duša ‘soul’, and a dead metaphor, i.e. ljubov’ ugasla ‘love was dying down’ (Jakobson 1979), all but one (Kushanov 2022) human translations enrich the rhetorical texture with additional metaphorical or personified elements. All translations preserve the morphosyntactic features to a large degree, and all but the machine translations keep the framing enjambements. Regarding the binary structure of the second stanza with its parallelisms, only the English translation by Shine (2020) disrupts the tight parallelism in line 6, rephrasing the alliterative and repetitive sequence to robost’ju, to revnostju tomim into the innovative metaphor in torment such as jealous fear compels. It is also the only translation to preserve the inverted word order of the original’s initial phrase ja vas ljubil ‘I loved you’ by means of you I loved.
Overall, close qualitative examination shows that the human translations scoring highest and lowest in the quantitative predictive modelling analysis represent the foregrounding features of the original poem and their distribution across the poem’s text fairly well; the differences between translations are not substantial. The same is true for the English Google translation, but not for the German one. The latter lacks a consistent representation of the foregrounding features of the original. Both machine translations fail to transfer the original’s end rhyme.
A major flaw of both Google translations is their limited ability to capture and correctly transfer the syntactic structure of the original. This is apparently because their line-wise translation does not recognise enjambements, i.e. dependencies that cross line boundaries. While this failure accidentally creates new metaphors, it leads to evident distortions of sense and meaning and ignores a key feature of the original’s poetic structure.
4. Discussion
In this article, we applied empirically verified methods to test the contribution of selected features to the aesthetic potential of poetry in order to probe into the extent to which translations preserve aspects of the aesthetic and emotional experience a (native) reader of the original is exposed to. To do so, we selected an English (Shakespeare’s Sonnet 27) and a Russian (Pushkin’s Ja vas ljubil) poem and translations thereof into German, English, and Russian. Given the growing importance and increasing quality of AI translations of poetry, we included four Google translations. We evaluated the correlations between originals and translations as a measure of faithfulness by implementing a middle-reading approach. For the computational comparison, we selected six textual features and applied Jacob’s (2023) empirically validated predictive modelling approach, which integrates the recipients’ perspective. The qualitative inspection of foregrounding stylistic devices based on a cross-linguistically and cross-culturally applicable feature matrix integrated the literary experts’ view.
Mean total correlations of the six features considered vary widely across originals and translations, between –.02 and .499 for Sonnet 27, between .005 and .57 for Ja vas ljubil. That is, some translations mirror the aesthetic and emotional potential of the original fairly well, others display mean correlation coefficients close to 0 and don’t seem to represent the potential of the original in a faithful way. Averaging over the feature-specific correlations, translations of course differ in the degree to which they preserve the aesthetic potential along the different dimensions tested. For both poems, a machine translation performs best in the computational comparison, a Russian Google translation of Sonnet 27 and an English Google translation of Ja vas ljubil. The German Google translations of both originals performed considerably worse (Sonnet 27: .219, Ja vas ljubil: .23). This points to substantial differences in language model configuration and quality. We interpret the results to show that while it is possible to make inferences about the aesthetic and emotional potential of a poem based on a translation, not all translations manage to preserve it faithfully. The degree strongly depends on individual decisions made by human translators and the model specifications underlying AI translations.
Mean total correlations for human translations in the computational comparison are notably stronger for Pushkin’s (top 4 translations between .49 and .41) than for Shakespeare’s poem (top 4 translations between .30 and .19). This most certainly relates to the differences in the importance of meaning vs. structure related features: Sonnet 27 is not only longer but also of higher lexical and structural complexity than Ja vas ljubil, which is more repetitive. Translations preserve the aesthetic potential induced by repetitive structures even though they might not exactly mirror the (grammatical) features of the original. We observe a similar effect in title prediction: Eigensimilarity (.77) and Proportion of Nouns (.7) have by far the highest effect for Ja vas ljubil and considerably more than for Sonnet 27 (.63 and .56, resp.). This is likely due to the lexical and structural specifics of the poems as well, as a less complex structure and lower lexical diversity increase similarity between the units of comparison, i.e., lines. These observations hint at the time of composition and its cultural preferences, length, content, and structure of the original co-determining the preservation of the aesthetic and emotional potential in translation.
Feature reliability tests in title (authorship) prediction returned very high prediction accuracy. In feature importance evaluation, Eigensimilarity turned out to have the highest effect (.7), all other features scoring in a range between .46 and .57. This indicates that our limited research-based selection of six features is a promising candidate for future studies using other poetic materials. The fact that Eigensimilarity is the most important feature in predicting the poem to which a line belongs very well reflects the intuition that in principle, content-based similarity between lines of a single poem is greater than the one between lines of two different poems. As for model validation, the results of our exploratory predictive modeling analysis using a neural net are conclusive. Overall, the six-feature model predicts the correct title of our sample of poems with an excellent accuracy of 99%. This result lends support to the suitability of our method to predict the poetic potential of a given poem based on selected features, and its applicability to poems and languages for which we don’t have experimental validation, e.g., underresourced and/or ancient languages and poetic traditions.
The results of the qualitative close-reading evaluation of foregrounding poetic devices in Pushkin’s poem revealed that except for one AI translation, the poetic devices are represented fairly well across translations. At least in this case, translations of linguistically and culturally more distant poetry and verbal art still preserve foregrounding features of the original that contribute to the aesthetic and emotional experience of native readers. In two respects, AI translations stand out: They lack the rhyming patterns of the original and they don’t represent the stylistic device at the syntactic level contained in the enjambements of the original. Further qualitative checks and inspection of content and stylistic devices reveal that our six-feature model is not yet able to reliably capture the aesthetic and emotional potential of a poem based on translations alone. In fact, the translations with the highest mean correlations, i.e. machine translations, score high despite not fully recognizing syntactic structures and intended meaning. We identify one major reason for this discrepancy. Our model operates on a line-by-line basis, assuming that poetic line boundaries coincide with boundaries of linguistic structure and content. The AI translations in our sample share this assumption: they systematically ignore any syntactic and semantic relations that cross line boundaries. But this assumption is obviously inaccurate because many semanto-syntactic units and stylistic devices (e.g., nominal expressions, sentences, clauses, parallelisms, assonances) often cross-cut line boundaries. Our model wrongly rewards AI translations for translating line by line and penalizes human translations that take cross-line dependencies into account. Extensive prompt engineering will likely remedy this deficiency in the case of machine translations (Yao, Kang, and McCosker 2025).
The discrepancy between the high scores in mean total correlations along the six aesthetic dimensions and the failure to transfer a key stylistic device at the syntactic level underscores the relevance of including a principled close-reading approach in a comparative context. As long as models are still confined to line-wise modelling and as long as figurative means are hard to identify and quantify computationally, the middle-reading approach proposed here is the best way to go.
As for the cross-linguistic and cross-cultural evaluation of aesthetic and emotional experience of poetry an important next step will be to validate the results obtained for the translations into Russian with reader responses and, importantly, integrate syntactic and semantic relations and dependencies in the model. This also includes a more flexible way of dealing with units of analysis. In fact, in highly influential theories of poetry and verbal art the (metrically) organised line is the basic unit of analysis (Fabb 2015; Hymes 1994). But this concept is fuzzy and ill-defined, and there is little evidence that it is applicable to, and relevant for all manifestations of poetry and verbal art.
5. Conclusion
Language used in poetry employs content and structure primarily to evoke aesthetic and emotional responses. Often, poetry can be accessed only via translations. This raises the question to which degree translations allow us to access the emotional and aesthetic potential of the original and the reader responses it evokes. We proposed a middle-reading approach to tackle this question, i.e. an approach that combines computational distant reading with philologically informed close reading. Specifically, we quantified the faithfulness of translations to the original along six dimensions of aesthetic space that are in part (for English and German, not yet for Russian) validated with respect to reader responses and in addition qualitatively inspected the preservation of stylistic devices. The results show that it is possible to leverage the potential of a predictive approach for accessing the aesthetic and emotional potential of a poem via translations into languages with LLMs of sufficient size, ideally based on a large number of translations. In the future, our six-feature model should be expanded to include more dimensions of the aesthetic space.
At the same time, close principled examination of translations discloses further features of aesthetic and emotional effects that cannot (yet) be quantified. For the time being, a middle-reading approach seems most promising in assessing the aesthetic and emotional faithfulness of poetry translations and to provide insight into the effects on linguistic and cultural natives.
In a more general perspective, our results open up concrete perspectives for further refining the tools applied and thus provide an important step towards a principled comparative, cross-linguistic and cross-cultural description of the aesthetic and emotional potential of poetry.
Data Accessibility Statement
SentiArt is available at: https://github.com/matinho13/SentiArt
Texts are available at: https://e.pc.cd/x3dy6alK
Acknowledgements
We gratefully acknowledge the feedback and insightful comments provided by two reviewers, which helped us to improve the presentation and clarity of our arguments.
Author contributions
BS: Conceptualisation, Writing: original draft, review, editing;
MR: Conceptualisation, Investigation, Writing: review, editing;
PW: Conceptualisation, Writing: original draft, review, editing;
AJ: Conceptualisation, Formal analysis, Software, Writing: original draft, review, editing.
