1 Context and motivation
English has undergone substantial changes in the way complex events such as resultative and motion expressions are mapped onto the syntactic structure. In terms of Talmy (2000)’s typology, English shows diachronically consistent characteristics of a satellite-framing language in mapping the end state or path onto a grammatical element outside the verb. The inventory of satellites has changed dramatically, however. Old English (OE) and early Middle English (ME) employed an elaborate system of prefixes, which developed out of adverbial preverbs, to mark resultativity and telicity of the predicate (cf. Hiltunen, 1983; Los et al., 2012; Thim, 2012, among others). These prefixes could be inseparable, as in (1) in which the prefix a- moved together with the verb, or separable as in (2) in which the particle ut- is stranded by verb movement. Inseparable prefixes were lost as a productive phenomenon after OE, while the separable prefixes syntactically developed into postverbal particles.
(1)
Ða
the
þry
three
oðre
other
godspelleras
evangelists
a-writon
up-wrote
heora
their
godspell
gospel,
be
by
Cristes
Christ’s
menniscnysse
humanity
“The three other evangelists wrote down their gospel about Christ’s humanity”
(coaelhom, ÆHom & 1:17.7, adapted from Elenbaas, 2007, 114)
(2)
þa
then
sticode
stuck
him
him
mon
one
þa
the
eagan
eyes
ut
out
“then his eyes were gouged out”
(coorosiu,Or 4:5.90.13.1822, Elenbaas, 2007, 131)
These developments took place in the context of extensive language contact with French following the Norman Conquest in 1066. The influx of French vocabulary, including a large number of verbs (Durkin, 2014; Percillier et al., 2024), introduced lexical items originating in a language that, as a verb-framing language, differed typologically from English in the expression of motion and result by encoding these meanings in the verb itself (Troberg and Burnett, 2017; Troberg and Leung, 2021). Compare in this regard the PDE phrasal verb move in to its synonym enter, which was directly copied from French entrer in ME.
The changes in complex event expression thus show interaction across several linguistic domains: morphology, syntax and lexical semantics and are at the same time modulated by language contact effects. There is significant research on each of these respective domains, from various analytical perspectives (space prevents a full discussion, but see Elenbaas, 2007; Fanego, 2012; Huber, 2017; Los et al., 2012; Martín Arista, 2012; McFadden, 2015; Rodríguez-Puente, 2018; van Gelderen, 2018 for some influential and recent work). However, how these phenomena relate to each other is still poorly understood. This is also noted by Thim (2012, 157), who, after an extensive review of earlier literature, comes to the conclusion that “what is missing is a treatment of the different types of word formation in a coherent and systematically related way” to understand how morphology, syntax, semantics and etymology interoperate. Arguably, a reason for this is that up to now there is no good resource available which allows studying these domains comprehensively from a diachronic perspective.
Existing resources and methodologies are often optimized for one of these domains. Historical dictionaries such as the Oxford English Dictionary (OED) (Proffitt, 2015), the Middle English Dictionary (MED) (McSparran et al., 2001), and the Bosworth Toller Anglo-Saxon dictionary (BT) (Bosworth et al., 2010) provide detailed information on etymology and morphological structure, but they cannot be used in a way that allows a quantitative analysis of, for instance, prefix-verb combinations, their relation to the syntactic structure or correlations with syntactically relevant verb classes. Conversely, corpus-based and syntactic studies have frequently focused on individual constructions, such as verb-particle combinations, individual lexical items, such as the behaviour of one verb, or only a token-based approach to an individual prefix. Furthermore, the studies often employ annotation practices and datasets tailored to their specific aims.
As a result, findings are not always easily comparable or interoperable across studies, which complicates broader generalizations about the diachronic development of English event structure. The absence of shared annotation schemes and openly available datasets also makes it difficult to combine lexical-semantic, morphological, and syntactic information systematically or to replicate analyses across larger bodies of data. In practice, this means that (substantial) manual curation and (re)annotation are required before quantitative or multifactorial analyses become possible. A fuller understanding of the historical development of motion and resultative constructions therefore benefits from and in fact requires resources that can relate morphology, lexical semantics, syntax, and etymology within a unified and reusable empirical framework.
This paper will introduce and evaluate the Roots & Results (RoRe) dataset, which was specifically developed to fill this empirical and methodological gap. It is the first exhaustive resource available for the full inventory of English verbs which combines the levels of annotation relevant to study change in the expression of resultativity and motion (prefixation, lexical semantics and etymology), fitting well with other current research which aims to quantitatively assess the syntactic and morphological expression of motion and resultativity in early Indo-European languages, such as Rainsford and Piccione (2025) on French and Italian motion expressions and Farina and McGillivray (2025) on preverbs in Ancient Greek and Latin.
The dataset is fully compatible with the Penn Parsed Corpora of English, the gold standard syntactic corpora in the field of historical linguistics. The original versions of the corpora were not lemmatized, making it difficult to enrich them in a time-efficient way, as a consequence of the significant spelling variation in historical texts. Recently, a lemmatized version of the York-Toronto-Helsinki Parsed Corpus of Old English Prose (YCOE) (Taylor et al., 2024) has become available, as well as a lemmatizer (Percillier, 2018; Trips & Percillier, 2020) dedicated to the lemmatization of the collected parsed corpora for ME, the Penn Parsed Corpus of Middle English 2 (PPCME2) (Kroch et al., 2000), the Parsed Linguistic Atlas of Early Middle English (PLAEME) (Truswell et al., 2018) and the Parsed Corpus of Middle English Poetry (PCMEP) (Zimmermann, 2018). This development allows the corpora to be enriched with additional annotation derived from the lexicographic sources, so that it becomes possible to answer questions relating to changes at the syntax-semantics interface from a contact perspective in a uniform and replicable way.
2 Dataset description
The dataset comprises stand-off lemma-based annotation for all uniquely occurring verb lemmas occurring in the Penn Parsed Corpora of English. The dataset comprises annotation for 4562 OE, 4262 ME and 3839 EModE lemmas. The lemmas are matched wherever possible to entries in the BT, MED, and the OED and subsequently annotated for prefixation, etymology of the prefix and verb and lexical semantics in terms of Levin & Rappaport Hovav’s (1995, 2019) Manner/Result complementarity framework. We here provide a brief overview of our annotation scheme. A more detailed annotation manual is included in the repository.
Repository location
Repository name
OSF.
Creation dates
2023-05-16 - 2025-05-26.
Dataset creators
Tara Struik (University of Mannheim) – conceptualization, methodology, data curation, data annotation, data validation, tool development. Lena Kaltenbach (University of Mannheim) – conceptualization, methodology, data annotation, data validation. Susanne Lang (University of Mannheim) – data curation, data annotation, Carola Trips (University of Mannheim) - funding acquisition.
Language
English.
License
CC-BY 4.
Publication date
2025-09-19.
2.1 Prefix annotation
We first compiled an inventory of prefixes occurring in OE, ME and Early Modern English (EModE) by consulting relevant literature on prefixation in the history of English (Elenbaas, 2007; Hiltunen, 1983; Thim, 2012, among others)1. We took a practical approach to determine whether a lemma is prefixed based on observable co-occurrences of a prefixed and non-prefixed form. This is to avoid overestimating the decompositionality and productivity of prefixes, but this means that there may be cases that were decomposable for a speaker of historical English, but for which no evidence survives.
If a verb is of native origin, we only considered whether separate entries for both the prefixed and non-prefixed variant of the verb existed. For the OE lemma oflecgan, for instance, we find both the lemma oflecgan “to lay down” and lecgan “to lay, set, place” in the BT. The two entries are clearly semantically related, and so we annotated oflecgan as prefixed by of-.
This is more complex for French copies, because of the difficulty to determine whether verbs are copied as simplexes or as derived combinations of prefix and verb. The approach we took is to look at the co-existence of prefixed and non-prefixed forms in both English and French (on the basis of the Tobler-Lommatzsch (Blumenthal & Stein, 2002) and Anglo-Norman dictionaries (Trotter, 2006)). We assume that if a verb exists in a prefixed as well as bare form in French, speakers should at least have been exposed to prefixed and non-prefixed forms of the verb and hence have evidence that the prefixed unit is derived. If there is no bare form in French, we assume that the prefixed form is lexicalised, and that the verb is copied as a simplex unit. If the copied verb only appears in the prefixed form in English, but never in the non-prefixed form, we assume that it is copied as a lexicalised unit. If it appears in both prefixed and non-prefixed form, we assume that it is copied as a derived unit. To illustrate, the ME verb denouncen “to communicate, declare” may be hypothesized to be prefixed by the French de- prefix. However, there is no verb nouncen, which leads us to conclude that denouncen was copied as a simplex and is hence not prefixed.
2.2 Etymological annotation
The dataset includes etymological annotation for all ME and EModE prefixes and bare verb forms based on information in the MED and OED. Copied verbs are very few in OE and for our purposes, as they are mostly from equally s-framing Old Norse, so etymology is not annotated for OE (see the GERSUM project (Dance et al., 2019) for this contact scenario).
ME and EModE lemmas are annotated as either “native”, “French”, “Latin” or “Other”, where the “Other” category is further specified by the source language in square brackets. For example, the ME verb bisemen “to have a certain appearance, seem” is composed of the prefix be-, which is of native origin, and the bare verb semen, which has its roots in early Scandinavian, according to the OED. Therefore, the bare verb has been assigned the label “other [early Scandinavian]”. The annotation thus refines the annotation in Trips & Percillier (2020), who only distinguish between “french” and “nonfrench”, but do not annotate for other etymologies.
Occasionally, the OED indicates that a verb may have multiple origins, making it impossible to assign one etymological source to a verb. For instance, the verb bipilien “to strip a tree of its bark” contains the native prefix be- and the bare verb pilen “to rob, peel”, which is either an earlier Latin or a later Old French copy. These cases were annotated as “French/Latin”. In a handful of cases, the etymology of the bare verb could not be identified. An example is the verb jeer “to speak or call out in derision”, which is of “unknown origin“ according to the OED. It is therefore annotated with a question mark.
2.3 Lexical Semantic annotation
The lexical semantic annotation presented here finds its origin in the seminal Manner/Result distinction by Levin and Rappaport Hovav (1995, 2019). This distinction in the aspectual class of dynamic verbs has long been known to be grammatically relevant (e.g., Fillmore, 1970). It determines which structural event schema is activated, and thus governs the syntactic structures in which a verb may occur. For instance, Present-day English Manner verbs allow object deletion, as in (3a), whereas Result verbs do not, as in (3b).
(3)
a.
John swept (the floor).
b.
John broke *(the vase).
However, while the core meaning of a verb is relatively stable and varies in predictable ways, the structural architecture surrounding event structures may change, as recently shown in van Gelderen (2018). As such, annotating the lexical semantics of verbs is feasible in a diachronically parsed corpus, but also necessary to arrive at a better understanding of the syntax-semantics interface.
We annotated all non-prefixed lemmas (to avoid potential aspectual interference from the prefix) according to their inner aspectual properties. We qualitatively assessed the verb’s lexical semantics using the BT, MED, and OED. Our annotation is based on the observation by LRH that whenever a verb is not Stative, it lexicalizes either Manner or Result in the verbal root (Levin and Rappaport Hovav, 1995, 2019). We thus distinguish three general verb classes in the annotation process: 1) Stative, 2) Manner, and 3) Result. Result verbs are further subdivided into a class of state-encoding Result verbs and Result verbs encoding an eventive change of state.
We took care to determine the category of each lemma based on its canonical use indicated in the dictionaries. However, due to the polysemous nature of words, a lemma may have more than one sense. If the various subsenses converge on one ontological class, we annotated the lemma accordingly. However, when multiple interpretations are possible and can only be distinguished by considering the verb in context, the lexical semantic class of the verb is annotated as “Amb_”, where the values following the underscore indicate the source of the ambiguity2. In other cases, not enough information was available to make an informed decision, although this typically pertains to low frequency items. These verbs are likewise annotated as “Amb”. We annotated 129 out of 1526 OE lemmas (8.5%), 92 out of 2708 ME lemmas (3.4%) and 278 out of 3096 EModE lemmas (9%) as ambiguous.
2.3.1 Manner vs. Result
Manner and Result verbs are both durative and dynamic (Dowty, 1979; Levin and Rappaport Hovav, 2019). The difference lies in the type of dynamic change that is lexicalized in the verbal root: Manner verbs, such as dance, are associated with non-scalar changes: they encode activity or complex movement without implicating progress along an ordered scale. For example, dance involves structured bodily movement but does not inherently specify an endpoint or direction of change; one can dance indefinitely without reaching a result, cf (4a).
(4)
a.
Anna danced.
b.
Anna danced into the room
Any interpretation of scalar change follows from the combination of the Manner root with other sentential material, as in (4b). In contrast, Result verbs lexicalize a specific kind of change on an ordered scale (which can be binary, such as break, cf. Levin and Rappaport Hovav, 2010). For instance, the verb widen in (5) encodes a scalar change, with an ordering from narrower to wider.
(5)
The gap widens.
Unlike Result verbs, Stative verbs, such as love, believe, do not lexicalize a change of state, but lexicalize stable situations that are inherently durative without event-internal dynamicity. Such situations can occur on both an abstract and a concrete level, involving possession, mental and emotional states, dispositions, and habits.
We annotated a verb as a Manner verb if the action denoted by the verb cannot be determined to take place on an ordered scale without reference to other lexical material outside of the verbal root. We annotated a verb as a Result verb if the action denoted by the verb lexicalizes the change of state of an argument without reference to additional linguistic material. We annotated a verb as Stative, if the verb root did not denote dynamicity but encodes a durative, stable situation.
2.3.2 Result types
Recent research shows that the class of Result roots is not homogeneous (Beavers & Koontz-Garboden, 2020) and that more fine-grained distinctions must be made between Result verbs, both within individual languages (Yu et al., 2023) and cross-linguistically (Smith et al., 2026). Leaving theoretical details aside, the crucial distinction lies in whether a Result verb relates an individual to an event of change, which we label Result-cos verbs, or whether a Result verb relates an individual to a state, which we label Result-state verbs. This difference is illustrated in (6) and (7).
(6)
Kunne
Can
a
a
boy
boy
nu
now
breke
break
a
a
spere,
spear,
he
he
shal
shall
be
be
mad
made
a
a
kniht.
knight.
“Once a boy can break a spear, he shall be knighted.”
(MED: c1330 Why werre (Auch)266)
(7)
The
The
wynd
wind
aros,
arose,
the
the
weder
weather
derketh.
darkened.
“The wind arose, the weather darkened.”
(MED: c.1393 Gower CA (Frf 3)8.604)
The ME verb breken “to break” in (6) is classified as an event-focused Result-cos verb. The root meaning of the verb focuses on the event of breaking itself, but not the end-state of the breaking action. This meaning is consistent when we consider all entries in the MED. The ME verb breken has a total of 33 senses and numerous subsenses. All of these senses, however, denote at their core that there is a breaking event, regardless of whether it is a concrete or a more abstract breaking event. Moreover, the eventive nature of breken allows further specification of the result by means of additional resultative elements, such as particles or PPs.
In contrast, the ME verb derken ( “to become dark") in (7) is categorized as a Result-state verb. The verb is used intransitively in (7) with an unaccusative structure. The subject (weder “weather") undergoes a spontaneous change, marking the completion of the darkening process. Formally, this deadjectival verb root encodes the end-state of the event. There are also verbs which are not deadjectival and are transitive, such as the verb destroien “to destroy" in (8), which relates the end state of complete destruction to the argument the toun “the town". How this end-state is achieved is left unspecified.
(8)
…
…
The
the
toun
town
destroyed,
destroyed,
ther
there
was
was
nothyng
nothing
laft.
left
“The town was destroyed, there was nothing left.”
(MED: c.1385 Chaucer CT.Kn.(Manly-Rickert)A.2016)
3 Technical implementation
The lemma-based annotation allows semi-automatic enrichment of the Penn Parsed Corpora of English, making it possible to study diachronic changes in expression of argument and event structure in relation to the lexical semantics of verbs. At the time of writing, only the OE and ME corpora have publicly available lemmatization, although the procedure described here has also been used to enrich a non-publicly available lemmatized version of the EModE corpus. The dataset comes with the RoRe Annotator tool, a Python application which can be run via the command line, which allows users to enrich their versions of the corpora themselves via a user-friendly guided prompt. The RoRe Annotator currently supports annotation of corpora in .psd and .psdx (XML) format.
The annotation procedure is described in Figure 1. After the user has selected the time period (OE, ME or EModE) and the format of the corpus files (.psd or .psdx), the program loops through the files to find POS-tags beginning with V and matches the associated lemma with the corresponding lemma in the RoRe data files. Whenever the match is not unique, the lemma ID will additionally be retrieved, if available3. If this results in a unique match, Manner/Result and Result Type attributes are added to the corpus, and the corresponding annotation is added as values. If no match can be made with a lemma at all, Manner/Result and RType attributes will be added, but they will not have a value, ensuring that they will be returned when the corpus is queried for these attributes. If no unique lemma ID match can be established, the Manner/Result and RType annotation for both lemmas will be added, separated by “|”.

Figure 1
Annotation pipeline.
The annotation procedure for the .psd files relies on the lemmatization format used in the BASICS lemmatizer (Percillier, 2018), which enriches the .psd corpus format with attributes demarcated by “@”. This is illustrated in (9), where (9a) shows the bare corpus node for the verb setting and (9b) shows the enriched corpus nodes with the attributes l for lemma, m for MEDID and e for etymology.
(9)
a.
(VAG settyng)
b.
(VAG settyng@l=setten@m=39654@e=nonfrench@)
The YCOE corpus is lemmatized according to a different format (verb form-lemma), as in (10a). The RoRe Annotator has a built-in function which automatically converts the verb form-lemma format to the BASICS verb form@l=lemma@ format, as in (10b), to ensure clear demarcation of the various annotation levels, as well as consistency across the various corpora to facilitate the querying process.
(10)
a.
(VB^D secene-secan)
b.
(VB^D secene@l=secan@)
Whenever there is a match between the attribute l in the corpus and the RoRe data files, the verb nodes are enriched with the following attributes: pre for prefix, pre_et for prefix etymology, barev for bare verb, barev_et for bare verb etymology, mr for Manner/Result and rtype for Result Type. The fully annotated string for the example in (9) is (11) and the fully annotated string for the example in (10) is (12).
(11)
(VAG settyng@l=setten@m=39654@e=nonfrench@pre=n.a.@pre_et=n.a.
@barev=setten@barev_et=native@mr=Amb_light@rtype=cos@)
(12)
(VB^D secene@l=secan@pre=n.a.@pre_et=@barev=secan@barev_et
@=@mr=Manner@rtype=@)
This annotated structure can be queried using CorpusSearch2 (Randall, 2004) and allows the user to look for specific verbs, verbs of one etymology, prefixed verbs or a particular class of lexical semantics or a combination of these attributes by specifying the level of annotation using the exists function, e.g. (*pre=ge-@* exists) to find all verbs that begin with the prefix ge-.
4 The use of prefixes in the history of English
We illustrate the functionality of our dataset with a case study on changes in the English prefixation system. English used to have a verbal prefix system to mark completive aspect and resultativity, such as ge-, be- and a- (although their function may vary, see Elenbaas, 2007; Hiltunen, 1983; Thim, 2012). Most of these prefixes are lost after early ME, although some survive with marginal productivity, such as be- and for-. The substantial number of French and later in the Renaissance era of Latin verbs have been argued to introduce a novel system of prefixes based on Romance preverbs, such as mis-, dis-/des- and contra-.
This raises the question if verbal prefixation as a morphological mechanism was lost from the language and, more specifically, if the Romance prefixes become productive to the same degree and with the same function as native prefixes. Thim (2012) notes that previous work in this domain is usually only concerned with either the loss of the Germanic aspectual prefixes or the rise of Romance prefixes, but that the two changes are rarely discussed in relation to each other. Thim (2012) presents a type-based analysis of the prefix inventory, which reveals a substantial increase from OE to ME. This leads Thim (2012, 156) to posit that “it can be safely concluded that prefixation as a native word formation type remained productive in Middle English.”
However, Thim (2012) is not specific on his measure of productivity, nor does he consider the combinatorial properties of the prefixes and verbs in relation to their etymologies. This is relevant from a contact perspective, because hybrid formations are an important indicator of morphological copying, rather than only lexical copying (e.g. Gardani, 2018). We follow Dalton-Puffer (1996), who studies the influence of French on English morphology in domains other than verbal prefixation, in considering two types of hybrid forms, illustrated in (13). Especially the second hybrid type (13b) is indicative of morphological copying, as it indicates the reanalysis of a morpheme as a meaningful component.
(13)
a.
Romance base + Germanic affix
b.
Germanic affix + Romance affix
Our dataset allows us to immediately extract all the verbs from the enriched corpora together with their prefix, the etymology of the prefix, the verb stem, and the etymology of the verb stem, which we quantify in terms of types and tokens. We collated the O1 and O2 periods from YCOE (based on manuscript dates) in one period “early OE” (850–950) and the O3 and O4 periods in one period “late OE” (950–1150). Similarly, M1 and M2 from the ME corpora are collated into one period “early ME” (1150–1350) and M3 and M4 in one period “lateME” (1350–1500). The subperiodization for EModE is taken directly from the corpus: E1 (1500–1570), E2 (1570–1640), E3 (1640–1710). We first of all present our findings for prefixation across French and native verbs diachronically, both on type and token level, in Table 1 and Figure 2.
Table 1
Total number of verb types and tokens in the historical English corpora and the percentage of verbs that occur with a prefix.
| TOKENS | % PREFIXED | TYPES | % PREFIXED | |
|---|---|---|---|---|
| early OE | 35591 | 46.6 | 1085 | 69.7 |
| late OE | 160391 | 48.7 | 1344 | 71.1 |
| early ME | 82915 | 17.4 | 2123 | 46.6 |
| late ME | 88505 | 8.4 | 1852 | 29.9 |
| E1 | 55016 | 3.7 | 1863 | 12.2 |
| E2 | 66681 | 3.6 | 2105 | 12.9 |
| E3 | 54524 | 4.4 | 2157 | 13.0 |

Figure 2
Percentage of types and token occurring with a prefix by etymology of the verb from OE to EModE.
The data show that the frequency of prefixation reduces significantly from OE to early ME. Most interesting for our purposes is the period from early ME onwards, when French verbs establish themselves in the language. Prefixed French and native verbs occur at comparable frequencies on a token level, but prefixed French verbs are less frequent on a type level. The frequency of prefixed types drops with both French and native verbs until the early EModE period, when they stabilize. Crucially, French verbs remain more frequently prefixed on a token level. These findings already allow us to refine Thim’s (2012) confident conclusion that prefixation remained productive. While the type inventory of prefixes increases towards Modern English, the verb types that can function as a base for these verbs reduces significantly, suggesting that these prefixes are not productive.
To further gauge the productivity of the various prefixes we introduce the etymology of the prefix as an additional variable and calculate the relative frequencies of native and French prefixes on the subset of prefixed verbs. The data in Figure 3 show a near categorical separation between prefix etymology and the etymology of the verb base: French prefixes combine with French verbs and native prefixes combine with native verbs. There are two exceptions, which we briefly discuss in turn.

Figure 3
Distribution of native and French prefixes across native and Romance verb bases diachronically, on both type and token level.
In early ME we find a substantial number of native prefixes attached to a French base, in terms of types as well as tokens. Closer inspection of the data reveals that the majority of these cases, illustrated in (14), are participles prefixed with ge-/i-, which at this stage is typically analysed as an optional inflectional marker (cf. Elenbaas, 2007; Martín Arista, 2012; McFadden, 2015; Thim, 2012).
(14)
today
today
Þou
you
scalt
shall
ben
be
icrounet.
crowned
biforn
before
Þe
the
king
king
of
of
heuene.
heaven
“today you will be crowned before the king of heaven.” (m2_tr323bt, 288)
We also find some new formations on the basis of the native prefix be-, although this seems to be restricted to a small number of verb types, including begylen “to beguile” on the basis of French guiler, from the noun guile “deceit, trickery”, and betraien “to betray”, on the basis of French traïr “to betray”, potentially stressing the affectedness of the patient associated with the verb. These verbs continue to exist in EModE, although the decompositionality of betraien is lost, as the verb base traien/tray no longer occurs in EModE. At the same time, new formations of a native prefix with a French base emerge, such as misconceive or overcharge. It is clear that the functionality of these prefixes is different: they are not aspectual or resultative markers, but their function is more akin to adverbial modifiers, specifying the way an action is carried out.
The second exception is the combination of French prefixes and native verb bases in EModE. Interestingly, such combinations hardly occur in late ME, despite extensive language contact up until this point. We report one type which occurs only once, suggesting it is based on an analogical formation: enspiren “to inquire” which is derived from the French prefix en- and the native base spiren “to inquire”. The remaining tokens belong to the type reneuen “to renew”, which seems to be the first hybrid form that we can trace in the corpora. The number of hybrid combinations increases in EModE, but the number of types remains relatively low. Most tokens feature the prefix re-, although we also find several new formations based on en-, such as engrave and enrich. These observations corroborate the conclusion in Dalton-Puffer (1996) that the Romance morphology is not productive in ME. At the same time, it shows the relevance of studying contact effects also after language contact has ceased or significantly decreased.
5 Implications/Applications
The present case study allowed us to refine earlier observations about the continued productivity of verbal prefixation in English (cf. Thim, 2012) by distinguishing more systematically between the etymology of both prefix and verb and by tracing their diachronic distribution from both a type and token level, instead of focusing on type frequency of the prefix alone.
This is only one example of the questions that can be addressed on the basis of the RoRe dataset. A further avenue concerns the development of prefixation in interaction with verbs of different ontological classes and in relation to other satellites, such as particles and potentially also preposition phrases. Here, the compatibility with the Penn Parsed corpora is particularly valuable, as this allows the extraction of individual verbs or verb classes in various syntactic structures without the need for extensive manual annotation. In fact, Struik and Kaltenbach (2026) show on the basis of this dataset that especially French Result verbs, which predominantly encode an end state, do not combine with native particles, and our case study suggests that they also do not combine with native prefixes. Together, these observations show that copies from French retain their verb-framing properties.
While the RoRe dataset was developed within the scope of a research project focusing on change in the domains of motion and resultativity in a contact situation, the annotation framework itself was designed to be broadly applicable beyond these domains, for instance, to study the effects of language contact or lexical semantics on other (morpho)syntactic phenomena. Furthermore, the lemma-based annotation pipeline is readily extendible to any lemma-based phenomenon, for instance, suffixation as an immediate counterpart to prefixation or other semantic features of verbs, such as verb classes (as in Levin, 1993, cf. Percillier, 2018). As such, the RoRe dataset illustrates the value of combining parsed corpora with detailed linguistic annotation in a transparent and reusable way to study complex questions regarding diachronic syntactic and semantic change.
Notes
[1] This overview is available in the repository and also contains general information on the semantics of the prefixes. The semantic annotation is not part of the annotated data, because the precise meaning of the prefixes is elusive and may vary per verb. We leave this to future work.
Acknowledgements
This work has greatly benefited from in-depth discussions with members of the SILPAC research unit regarding both linguistic and technical issues. We are particularly grateful to Carola Trips, Achim Stein, Michelle Troberg, Thomas Rainsford, Mariapaola Piccione, and Michael Percillier. We also acknowledge the feedback and questions we received from the participants at the 46th International Computer Archive of Modern and Medieval English (ICAME) conference held at Vilnius University, 17–21 June 2025, where this work was presented.
Author Contributions
Tara Struik – Conceptualization, methodology, data curation, validation, formal analysis, project administration, software, writing – original draft/review & editing; Lena Kaltenbach – Conceptualization, methodology, writing – original draft/review & editing; Susanne Lang – Data curation, writing - original draft.
