1 Context and motivation
Initially, we planned to conduct a philological study comparing the Armenian, Georgian, and Ancient Greek versions of the Book of Ezra, to confirm the proposed relationship that the Book of Ezra in the Oshki Bible was translated via Armenian from Greek rather than directly from Greek.1 The idea was that we would establish a checklist of linguistic features extracted from the Book of Ezra that are indicative of Armenian as the source language rather than Ancient Greek, and then apply this checklist to determine whether other Old Georgian Biblical translations derive from Armenian or from other source languages.
In the process, we experimented with different methodologies and eventually decided to build a digital corpus annotated according to the Universal Dependencies (UD) guidelines, in order to facilitate statistical and distributional analyses of morphosyntactic features in comparison with the original Armenian source text. In the broader context, we believe that publishing our corpus as an open dataset in the era of large language models (LLMs) and artificial intelligence (AI) will benefit not only our own research and the wider scholarly community, but also all those interested in pre-modern Georgian texts and Georgian cultural heritage.
1.1 Related Work
When we started the project in November 2025, we were not aware of standard UD guidelines for Old Georgian. However, as the reviewers have kindly pointed out to us, in the latest UD release (May 2026), Irina Lobzhanidze published an Old Georgian treebank based on the GLC corpus and provided the annotation guidelines she developed (Lobzhanidze et al., 2026).2 Besides this, there are two Modern Georgian UD-annotated corpora: one built by the Georgian National Corpus (GNC),3 developed by Paul Meurer,4 and one built by the Georgian Language Corpus (GLC),5 developed by Irina Lobzhanidze. We decided to build an Old Georgian UD-annotated corpus according to a set of new guidelines specifically designed for Old Georgian, which represents a separate stage of the Georgian language, morphosyntactically distinct from Modern Georgian, rather than merely a variety of Modern Georgian.6
This project follows in the footsteps of similar works on building UD-annotated corpora for ancient languages, such as Eckhoff et al. (2018) on Latin, Greek, Gothic, Classical Armenian, and Old Church Slavonic, Yasuoka et al. (2022) on Classical Chinese, Zeldes and Abrams (2018) on Coptic, Doyle and McCrae (2025) on Old Irish, and Hellwig et al. (2020) on Vedic Sanskrit.7
2 Dataset description
Repository location https://zenodo.org/records/20928957; https://github.com/UniversalDependencies/UD_Old_Georgian-OGLauRo
Repository name Zenodo, Github
Object name UD_Old_Georgian-OGLauRo
Format names and versions .conllu
Creation dates 2026-01-01
Dataset creators Chia-Wei Lin, Université de Lausanne, Switzerland (role: data annotation and methodological elaboration) and Diego Luinetti, Università degli Studi Guglielmo Marconi, Roma, Italy (role: data annotation and methodological elaboration).
Language Old Georgian, English.
License CC BY-NC-SA 4.0
Publication date 2026-04-10
3 Method
We construct the treebank through a human-in-the-loop annotation workflow that combines manual annotation by the authors with iterative few-shot prompting of a LLM (see Figure 1). The procedure is designed to let trained linguists scale annotation for a low-resource historical language without requiring programming expertise.

Figure 1
Old Georgian treebank annotation workflow.
3.1 Seed annotations
We first manually annotated the opening verses of the text following our annotation guidelines for Old Georgian (described below), producing a small set of gold examples in CoNLL-U format.
3.2 Iterative few-shot prompting
We then prompted Google’s Gemini8 to draft annotations for the next verse, conditioning on the gold examples produced so far, and two external resources retrieved at prompt time: (1) the UD_Georgian-GNC treebank9 for Modern Georgian morphosyntax, and Ilia Abuladze’s (1973) Dictionary of the Old Georgian Language (Materials) (ძველი ქართული ენის ლექსიკონი (მასალები)), digitized by the National Library of Georgia.10 We manually corrected each draft, and the corrected verse was appended to the in-context example pool used for subsequent verses. This loop was repeated for the entire Book of Ezra. After approximately 20 iterations, the model’s draft annotations stabilized in the sense that the proportion requiring manual correction dropped significantly. We attribute this to the growing pool of in-domain examples in the prompt, which progressively disambiguates guideline edge cases that are not covered by the Modern Georgian resources (more in detail, see Section 4.8).
3.3 Final proofreading
After manual annotation and correction of the LLM output, we finalized the dataset by proofreading the consistency of our annotations and discussing the noteworthy issues in our annotation presented in this paper.
4 Results and discussion
After a brief sketch of the Old Georgian language in Section 4.1, we discuss the challenges of developing UD-conformed guidelines tailored for Old Georgian in Sections 4.2, 4.3, 4.6, and 4.7, before turning to the limitations of LLM-based annotation in Section 4.8.
4.1 A sketch of the Old Georgian language
Georgian belongs to the Kartvelian (or South Caucasian) language families, one of the autochthonous language family of the Caucasus, alongside other minorities languages such as Mingrelian, Laz, and Svan. According to Tuite (2004), literary Georgian can be classified into four main stages:
Early Old Georgian (5th–8th centuries)
Classical Old Georgian (9th–11th centuries)
Middle Georgian (12th–18th centuries)
Modern Georgian (18th–20th centuries)
In terms of morphology, Georgian is a highly agglutinative language. Linguists distinguish up to nine nominal cases in Old Georgian (Fähnrich, 2012, 35; Imnaishvili, 1957; Sarjveladze, 1997, 23): nominative, absolutive, ergative, dative, genitive, directional, instrumental, ablative, and vocative. (see Section 4.3).
The verbal morphology of Georgian is characterized by polypersonalism (a single verb can simultaneously encode subject, direct object, and indirect object) and by a distinctive system known as the “screeve” (Georgian მწკრივი mc̣ḳrivi), which encodes tense, aspect, and mood (TAM; see Section 4.6). The distance between Old Georgian and Modern Georgian is smaller than, for instance, that between Ancient Greek and Modern Greek or Latin and Romance varieties, but comparable to that between Old English and Modern English or Classical Japanese and Modern Japanese. The principal differences lie in the following aspects:11
Writing systems: Old Georgian uses ასომთავრული Asomtavruli and ნუსხური Nusxuri scripts, whereas Modern Georgian uses მხედრული Mxedruli;
Orthography and phonology: Old Georgian preserved the following graphemes that have since become obsolete: ჲ ⟨Y⟩, ჱ ⟨ē⟩, ჳ ⟨wi⟩, ჴ ⟨q⟩, ჵ ⟨ō⟩;
Nominal morphology: Old Georgian preserved Suffixaufnahme (see Section 4.3.1) and two plural declensions: (1) in -ნი -ni (nominative case) and -თა -ta (oblique cases; see Section 4.3); (2) in -ებ- -eb-. Modern Georgian has predominantly plurals in -ებ- -eb-. Plurals in -ნი -ni- /-თა -ta are found mostly in formal context and archaising styles;12
Verbal morphology: preverbs in Old Georgian are primarily spatial or directional markers, whereas in Modern Georgian they have been grammaticalized to express temporal, aspectual and modal values. Furthermore, nominal plurals in -ნი -ni are cross-referenced by a verbal suffix -ნ- -n- in Old Georgian (see Section 4.6).
Morphosyntax: series III in the verbal conjugation was primarily resultative in Old Georgian; in Modern Georgian, series III has shifted toward expressing evidentiality or inferentiality.
4.2 Parts of Speech and features
We use the following UPOS tags from the UD guidelines: ADJ, ADP, ADV, CCONJ, NOUN, NUM, PRON, PROPN, PUNCT, SCONJ, VERB. For XPOS, we adapt the existing GNC corpus annotation for Old Georgian. More specific details are provided in the sections below.
4.3 Nouns and adjectives
In Old Georgian there is no formal distinction between nouns and adjectives, since they both share the same set of case markers (see also Hardie & Daraselia, 2026). Admittedly, a morphological divergence can be observed in superlative forms, which are generally available for adjectives but not for nouns (Shanidze, 1980, 50, 62–63). Crucially, however, no different labels are adopted by Fähnrich and Sardshweladse (2005). In our annotation, we distinguished them based on their semantic features: nouns express entities and adjectives denote qualities (Sarjveladze, 1997, 53). NOUNs and ADJs possess the Case and Number features. The XPOS field is N_[case]_[number] for nouns and A_[case]_[number] for adjectives.
As mentioned, linguists distinguish up to nine nominal cases for Old Georgian. Among these, we chose to annotate eight. The directional case is formally indistinguishable from the genitive, so we annotated them the same way. Although some scholars categorize the nominative and absolutive as variants of a single case, we kept them separate to reflect their different linguistic functions (see Berikashvili & Lobzhanidze, 2022).
Case feature has the following eight values: Nom for nominative (ი i, ჲ y), Dat for dative (ს s, სა sa), Erg for ergative (მან man), Gen for genitive (ის is, ისა isa), Ins for instrumental (ით it, ითა ita), Abs for absolutive (bare stem), Ess for adverbial (ად ad, დ d) and Voc for vocative (ო o).13 The UD annotation system lacks a specific value for adverbial case, which is very peculiar to Kartvelian languages. UD_Georgian-GNC an UD_Old_Georgian-GLC resolved this by choosing Ess (essive) from among the existing values. For consistency, we made the same annotation choice, even though the adverbial case originally had rather an allative function, which is still used in Old Georgian, as shown in example (1).
(1)
რომელნი
romel-n-i
გამოვიდოდეს
gamo-vid-od-es
სამნასარის
samnasar-is
თანა
tana
ბაბილონით
Babilon-it
იერუსალემდ
Ierusalem-d
rel-pl-nom pvb-come_out-ipf-s3pl Samnasar-gen with Babylon-ins Jerusalem-adv
‘(those) who came out of Babylon to Jerusalem together with Samnasar’ (1Ezra 1:11)
The Number feature takes the values Sing for singular and Plur for plural. However, Old Georgian presents an additional complication: in the plural, the ergative, dative, genitive, and instrumental cases share the syncretic form ta (თა). In order to preserve its formal polysemy, we noted ta-plurals as Obl in the XPOS field. Consequently, the annotation of Case in FEATS for ta-plurals is purely based on syntactic relations, such as the case governed by the verb or by an adposition. For instance, in example (2) ადგილთა adgilta is annotated as N_Obl_Pl in XPOS and as Case=Dat|Number=Plur in FEATS.
(2)
მიესმოდა
mi-e-sm-od-a
ჴმაჲ
qma-y
იგი
igi
შორიელთა
šoriel-ta
ადგილთა
adgil-ta
pvb-io3-hear-ipf-s3sg voice-nom det.nom distant-obl.pl place-obl.pl
‘A loud voice was heard in distant places’ (1Ezra 3:13)

We also annotated as Number=Plur eb-plurals, even though they trigger neither verbal agreement nor noun phrase agreement in Old Georgian.
The Number feature is not included for nominals in the adverbial case, since they only occur in the singular.
The UPOS tag PROPN is used for proper names, since they partially differ from other nominals (see Vogt, 1968; Tschenkeli, 1956). Notably, they tend to be uninflected where a nominative, ergative or dative case would be required, as can be seen in example (3), where ისუ Isu is uninflected while the adjective იოსედეკეანი iosedeḳeani referring to it stands in the nominative. If the case is not expressed, the Case and Number features are omitted. Additionally, proper names feature the NameType field, with the values Prs for given names and Geo for geographical names.
(3)
და
da
აღდგა
aγ-dg-a
ისუ
Isu
იოსედეკეანი
iosedeḳean-i
and pvb-stand_up-s3sg.aor Isu of_Iosedek-nom
‘and Isu son of Iosedek stood up’ (1Ezra 3:2)
4.3.1 Suffixaufnahme
Suffixaufnahme is certainly one of the hallmarks of Old Georgian. It consists of the repetition of case and number endings of the head noun onto other nominals governed by it, and it mainly functions as a strategy for delimiting noun phrase constituency.14 For instance, in example (4), the modifier უფლისასა upl-isa-sa has the genitive ending ისა isa and the dative ending სა sa stacked onto it, replicating the dative ending of the head noun ტაძარსა ṭaʒarsa.
(4)
წინაშე
c̣inaše
ტაძარსა
ṭaʒar-sa
მას
mas
უფლისასა
upl-isa-sa
‘in front of the temple of the Lord’ (1Ezra 10:1)
before temple-dat det.dat Lord-gen-dat

The UD annotation system lacks a native way to account for Suffixaufnahme. The UD_Old_Georgian-GLC treebank inconsistently annotates Suffixaufnahme, and the dedicated Case[stack] field it uses is not sufficient to grasp the full morphological profile of the construction, as it conflates the case of the genitive flag with that of the Suffixaufnahme proper, and provides no dedicated field for the latter’s number.15 A similar case-stacking phenomenon also occurs in Naga-Suansu (Sino-Tibetan), which also happens to have a treebank, annotated by Jessica K. Ivani and Kira Tulchynska.16 The authors opted for combined values for Case, such as ErgTop or GenAbl. However, Naga-Suansu case-stacking differs from Old Georgian Suffixaufnahme, since it does not involve repetition of endings. We thus prefer to separate the case marker of the word from the duplicated ending into two distinct FEATS fields: the already-mentioned Case and Number for the word’s proper endings and Case[sauf] and Number[sauf] for the duplicated endings. Moving back to example (4), uplisasa is annotated as N_Gen_Sg_Dat_Sg and Case=Gen|Case[sauf]=Dat|Number=Sing|Number[sauf]=Sing. This approach facilitates the separation of the word’s proper features from the phrase features during automatic data extraction from the treebank (for instance, with the udpipe package in R).
Moreover, it is versatile enough to account for another related phenomenon of case stacking called hypostasis or ellipsis.17 It consists of a noun phrase lacking a head noun, with double inflection on a genitive: the stacked case marker acts as a pronoun standing in the case required by syntax, governing the genitive.18 For instance, in example (5), the genitive noun phrase სახლისა უფლისაჲ saxl-isa upl-isa-y lacks a nominal head and features an extra nominative case marker, that encapsulates the whole noun phrase as the object of the verb უბრძანა ubrʒana.
(5)
კჳროს
ḳwiros
მეფემან
mepe-man
უბრძანა
u-brʒan-a
სახლისა
saxl-isa
უფლისაჲ
upl-isa-y
Cyrus king-erg io3-command-s3sg.aor house-gen Lord-gen-nom
‘Cyrus the king deliberated on the house of the Lord’ (1Ezra 6:3)

We rarely encounter complex cases of Suffixaufnahme with triple case marking. Following the same principle, we annotate such cases with the extra FEATS fields Case[dsauf] and Number[dsauf]. For instance, in example (6), იუდაჲსთაჲ iuda-ys-ta-y governed by the genitive მამათმთავართაჲ mamatmtavar-ta-y, in turn governed by the nominative მფლობელნი mplobeln-i, is annotated as:
Case=Gen|Case[sauf]=Gen|Case[dsauf]=Nom|Number=Sing|Number[sauf]=Pl|Number[dsauf]=Sing
(6)
და
da
აღდგეს
aγ-dg-es
მფლობელნი
mplobel-n-i
იგი
igi
მამათმთავართაჲ
mamatmtavar-ta-y
მათ
mat
იუდაჲსთაჲ
iuda-ys-ta-y
and pvb-stand-aor.3pl chief-pl-nom det.nom patriarch-obl.pl-nom det.obl.pl Iuda-gen.sg-obl.pl-nom
‘and the chiefs of the patriarch of Iuda’ (1Ezra 1:5)

4.4 Pronouns
Our treatment of pronouns generally follows that of UD_Georgian-GNC, except in the case of იგი igi, ესე ese, and ეგე ege. In Modern Georgian, these are either demonstrative or personal pronouns, but in Old Georgian, they can also function as postposed determinative articles.19 To resolve this ambiguity, we label all forms of these three pronouns as PRON in UPOS and Pron in XPOS. When they function as pronouns, we assign them their corresponding syntactic function in DEPREL (see example 7); when they function as articles, we label them as det in DEPREL (see example 8).
(7)
დასაჯეთ
da-sa-e-t
იგი
igi
გინა
gina
თუ
tu
სიკუდილად
siḳudil-ad
გინა
gina
თუ
tu
ტანჯვად
ṭanǯvad
pvb-condemn-ipv-s2pl 3sg.nom or if death-adv or if torture-adv
‘condemn him either to death or to torture’ (1Ezra 7:26)

(8)
არა
ara
განეშოვრა
gan-e-šovr-a
ერი
er-i
იგი
igi
ისრაელსაჲ
israel-sa-y
negpvb-ver-separate-aor.s3sg people-nom.sg det.nom.sg Israel-dat.sg-nom.sg
‘the people of Israel did not depart’ (1Ezra 9:1)

4.5 Adpositions
Old Georgian features both prepositions and postpositions, with a preference for postpositions. To account for this, the ADP UPOS corresponds either to Pre for prepositions or to Post for postpositions in the XPOS. The FEATS field includes AdpType, with the values Pre and Post, and Case, cross-referencing the case of the nominal governed by the adposition. For instance, tana in example (9) was annotated as AdpType=Post|Case=Gen.
(9)
რომელნი
romel-n-i
გამოვიდოდეს
gamo-vid-od-es
სამნასარის
samnasar-is
თანა
tana
ბაბილონით
Babilon-it
იერუსალემდ
Ierusalem-d
rel-pl-nom pvb-come_out-ipf-s3pl Samnasar-gen with Babylon-ins Jerusalem-adv
‘(those) who came out of Babylon to Jerusalem together with Samnasar’ (1Ezra 1:11)

Some postpositions, such as გან gan ‘from’ and ურთ urt ‘with’ are enclitic. Since they can attach in principle either to the head noun or to its determiners and modifiers, we decide to analyze them as multi-word tokens, following the established practice in UD_Georgian-GNC. When the postposition attaches to a determiner instead of the head noun, as shown by მისგან mis=gan in example (10), its syntactic dependency is still on the head noun.
(10)
სახლისა
saxl-isa
მისგან
mis=gan
უფლისა
upl-isa
house-gen det.dat=from Lord-gen
‘from the house of the Lord’ (1Ezra 6:5)

Finally, there are some instances where a postposition hosts Suffixaufnahme. To account for these cases, we added a Case[sauf] and a Number[sauf] feature to the FEATS field, in the exact same way we did for nominals. For instance, in example (11), the prepositional phrase მთავართა მათგანნი mtavar-ta mat=gan-n-i gets the nominative Suffixaufnahme from its nominal head ათორმეტნი atormeṭ-n-i and the postposition განნი gan-n-i is annotated as:
AdpType=Post|Case=Gen|Case[sauf]=Nom|Number[sauf]=Plur
(11)
და
da
წარვავლინეთ
c̣ar-v-a-vlin-e-t
მთავართა
mtavar-ta
მათგანნი
mat=gan=n-i
მღდელთაჲსა
mγdel-ta-ysa
ათორმეტნი
atormeṭ-n-i
and pvb-s1-ver-send-aor-s.pl chief-obl.pl det.obl.pl=from=pl-nom priest-obl.pl-gen twelve-pl-nom
‘and we sent forth twelve of the chiefs of the priests.’ (1Ezra 8:24)

4.6 Verbs
Verbs are arguably the morphologically most complex part of speech in Old Georgian. The verbal chain has a total of 10 slots, plus two additional slots reserved for clitics (Boeder, 2005; Hewitt, 1995; Lobzhanidze, 2022, 61–66; Tuite, 2004, 960–963). See Figure 2.
slot 1 hosts up to two directional preverbs;
slot 2 hosts a number of adverbials and indefinite pronouns that can trigger tmesis;
slot 3 hosts direct and indirect object markers;
slot 4 hosts subject 1st and 2nd person subject markers;
slot 5 hosts the version vowel (cf. Section 4.6.2);
slot 6 hosts the verbal root;
slot 7 hosts passive and causative markers;
slot 8 hosts direct object plural marker;
slot 9 hosts the series marker;
slot 10 hosts tense and mood markers;
slot 11 hosts 3rd person subject markers or 1st and 2nd person plural subject markers;
slot 12 hosts other clitics.

Figure 2
Old Georgian verbal chain.
Since not all morphological information that is relevant to traditional Georgian grammar can be annotated in the UD FEATS field of VERB, following the practice established in UD_Georgian-GNC, we use the XPOS field to include it. The other information is distributed among the native features of UD as shown below in Sections 4.6.1, 4.6.2 and 4.6.3. The XPOS field is structured as follows:
V_[voice]_[mood]_[tense]_[preverb]_[version]_[S:3Sg]_[O:3Sg]_[IO:3Sg]
The slots [voice], [preverb] and [version] carry information that is not annotated in FEATS. The [voice] slot has Act for active verbs or Pass for passive verbs, according to Schanidse’s (1982) classification. Since these labels are purely formal, they cannot be included among FEATS.20 The [mood] slot, either Subj for subjunctive or Imp for imperative, is left empty for indicative. The values of the [tense] slot mirror those of the corresponding FEATS field, discussed in Section 4.6.1. The [preverb] slot is filled with Pv if there is a preverb; otherwise it is left empty. The [version] is discussed in Section 4.6.2. The last three slots are dedicated to personal indexation.
The FEATS field for VERB includes the following features: Aspect, Mood, Number[io], Number[obj], Number[subj], Person[io], Person[obj], Person[subj], Tense, and VerbForm.
4.6.1 Tense and mood
In the Georgian grammatical tradition, tenses and moods are combined into screeves (მწკრივი mc̣ḳrivi), which are in turn organized into three series (რიგი rigi), as shown in Table 1.21
Table 1
Old Georgian screeves.
| INDICATIVE | NON-INDICATIVE | ||
|---|---|---|---|
| PRESENT | PAST | ||
| Series I | Present | Imperfect | Subjunctive I |
| Present Iterative | Imperfect Iterative | Imperative I | |
| Series II | Aorist | Subjunctive II | |
| Aorist Iterative | Mixed Subjunctive | ||
| Mixed Iterative | Imperative II | ||
| Series III | Perfect | Pluperfect | Subjunctive III |
| Perfect Iterative | |||
It is widely accepted that the distinction between Series I and Series II screeves is aspectual, with the former being imperfective and the latter perfective (Sarjveladze, 1997, 79; Schanidse, 1982, 77). However, given the complex combinations of TAM values expressed by screeves, for practical purposes of annotation and for consistency with the UD_Georgian-GNC treebank, we decided to use mainly the fields Tense and Mood, reserving the Aspect field for Iteratives. The field Tense natively supports values corresponding to some of the screeves: Pres for Present, Imp for Imperfect, Past for Aorist, Pf for Perfect, and Pqp for Pluperfect. The tags Pres, Aor and Pf in combination with moods other than the indicative purely signal the distinction between Series I, II and III, respectively. Iterative screeves were annotated combining the Tense feature with the Aspect feature, which is skipped for other screeves. For instance, in example (12) დავდევით davdevit was annotated as:
Aspect=Iter|Mood=Ind|Number[subj]=Plur|Person[subj]=1|Tense=Past|VerbForm=Fin
(12)
და
da
დავდევით
da-v-dev-i-t
მუნ
mun
სავანე
savane
სამ
sam
დღჱ
dγe-y
and pvb-s1-place-aor.iter-s.pl there abode.abs three.abs day-nom
‘and we camped there for three days’ (1Ezra 8:15)
The Mood feature has the values Ind for Indicative, Subj for Subjunctive, and Imp for Imperative. In order to distinguish between series I, II and III, we decided to use the Tense feature, with Pres for series I, Aor for series II and Perf for series III. To account for Mixed Subjunctive, we employ again the Aspect feature with the Imp (for Imperfective) value. Correspondences between screeves, XPOS and FEATS are given in Table 2.
Table 2
Screeves correspondences.
| SCREEVE | XPOS | FEATS |
|---|---|---|
| Present | Pres | Mood=Ind|Tense=Pres |
| Present Iterative | Iter1 | Aspect=Iter|Mood=Ind|Tense=Pres |
| Imperfect | Imp | Mood=Ind|Tense=Imp |
| Imperfect Iterative | ImpIter | Aspect=Iter|Mood=Ind|Tense=Imp |
| Subjunctive I | Pres_Subj | Mood=Subj|Tense=Pres |
| Imperative I | Pres_Imp | Mood=Imp|Tense=Pres |
| Aorist | Aor | Mood=Ind|Tense=Past |
| Aorist Iterative | Iter2 | Aspect=Iter|Mood=Ind|Tense=Past |
| Subjunctive II | Aor_Subj | Mood=Subj|Tense=Past |
| Mixed Subjunctive | SubjM | Aspect=Imp|Mood=Subj|Tense=Past |
| Imperative II | Aor_Imp | Mood=Imp|Tense=Past |
| Perfect | Perf | Mood=Ind|Tense=Pf |
| Perfect Iterative | Iter3 | Aspect=Iter|Mood=Ind|Tense=Perf |
| Pluperfect | Plup | Mood=Ind|Tense=Pqp |
| Subjunctive III | Perf_Subj | Mood=Subj|Tense=Perf |
4.6.2 Version vowels
Georgian verbs have a unique grammatical category traditionally called “version” (Georgian ქცევა kceva) by linguists, with four possible values: ა a, ე e, ი i, უ u. Nonetheless, no standardized morphosyntactic classification of the different version vowels with regard to their functions has yet been established. Schanidse (1982, 92–95) classifies the version vowels according to their beneficiary:
saarviso version (საარვისო): ა a or ∅, where there is no beneficiary.
sataviso version (სათავისო): ი i, where the beneficiary is the agent.
sasxviso version (სასხვისო): ი i or უ u, where the beneficiary is different from the agent.
Tuite (2004, 961) classifies the version vowels according to their respective syntactical functions:
subjective version: ი i, where the action is performed for the benefit of the agent themselves, or directed toward a direct object linked to the agent.
objective version: ი i or უ u, where the version vowel indicates the presence of an indirect object.
superessive version: ა a, where the version vowel indicates an indirect object upon which the action is accomplished.
neutral version: ა a or ∅.
Instead of grouping multiple version vowels into one group, Fähnrich (2012, 140–141) indicates the functions of each individual version vowel:
ა a marks transitivity of the verb or an indirect object.
ი i marks reflexivity or an indirect object in the 1st or 2nd person.
უ u marks an indirect object in the 3rd person.
ე e marks an indirect object.
In the XPOS of UD_Georgian-GNC, the authors use the annotation LV (Locative version) for ა a, SV (Subjective version) for ი i, OV (Objective version) for უ u.22
Although the functions of version vowels have not altered much from Old Georgian to Modern Georgian, in our UD annotation for Old Georgian, we propose to simplify the annotation by labelling a-version as AV, i-version as IV, u-version as UV, and e-version as EV for XPOS purely based on the morphological traits of the versions as they appear in a verb. This has the advantage of avoiding the need to formulate detailed and complex morphosyntactic rules to accommodate the classifications described above into the UD framework, since they entail several ambivalent cases. For example, ა a may be interpreted either as locative/superessive or neutral depending on the context, and ი i may be interpreted either as objective version or subjective version depending on the context. We argue that using labels AV, IV, UV, and EV based on the superficial appearance of the version vowels would be the simplest way to avoid the ambiguity of morphosyntactic classifications. The annotation of version vowels is summarized in Table 3.
Table 3
Version vowel annotation.
| EXAMPLE | VERSION | XPOS | TRANSLATION |
|---|---|---|---|
| აკურთხევდეს aḳurtxevdes (1Ezra 3:11) | ა a | V_Act_Imp_AV_S:3Pl | ‘they blessed’ |
| დაიდვა daidva (1Ezra 3:11) | ი i | V_Pass_Aor_IV_S:3Sg | ‘they put’ |
| მიუგო miugo (1Ezra 10:2) | უ u | V_Act_Aor_Pv_UV_S:3Sg_IO:3Sg | ‘he answered’ |
| ესხა esxa (1Ezra 10:2) | ე e | V_Act_Aor_EV_S:3Sg_IO:3Sg | ‘they had’ |
Nevertheless, our approach still makes it possible to infer the traditional classification of versions by combining the formal label in the XPOS field of the verb with the values of the DEPREL field of its arguments. For instance, in examples (13) and (14), all the verbs are annotated as IV in the XPOS field. However, in example (13), the verbs ვიმარხევდით vimarxevdit and ვითხოვდით vitxovdit correspond to the sataviso (subjective) version, since both lack iobj as their arguments. On the other hand, the verb მიგიგოთ migigot in example (14), corresponds to the sasxviso (objective) version, since it has შენ šen as its iobj.
(13)
და
da
ვიმარხევდით
v-i-marx-ev-di-t
და
da
ვითხოვდით
v-i-txov-di-t
ღმრთისაგან
mrt-isa=gan
ჩუენისა
čuen-isa
ამის
am-is
ყოვლისათჳს
q̇ovl-isa=twis
and s1-sv-fast-tm-ipf-s.pl and s1-sv-ask-ipf-s.pl God-gen=from our-gen this-gen everything-gen=for
‘and we fasted and sought from our God for all this’ (1Ezra 8:23)

(14)
და
da
აწ
ac̣
რაჲ-მე
ra-y=me
მიგიგოთ
mi-gi-g-o-t
შენ
šen
წინაშე
c̣inaše
შენსა
šen-sa
and now what-nom=ptcl pvb-io2-answer-aor.subj-s.pl 2sg before 2sg-dat
‘and now what shall we answer you in your presence?’ (1Ezra 9:10)

4.6.3 Personal indexation
As stated in Section 4.1, Old Georgian verbs allow for polypersonalism and indexes the subject, direct object and indirect object. Personal indexation is distributed among slots 3, 4, 5, 8, and 11 of the verbal chain (see Figure 2). While the subject is always directly inferable from the morphology, direct and indirect object can be zero-marked due to morphophonotactic adjustments, which concern especially 3rd person singular objects. In such instances, we decide to exclusively annotate the information available from morphology. For instance, in example (15), მისცეს misces has a 3rd plural subject signaled by -ეს -es, a 3rd person indirect object, with unspecified number, signaled by the index -ს- -s-, and no object marking, even though it governs the direct object წიგნი c̣igni. For this reason it is annotated as:
Mood=Ind|Number[subj]=Plur|Person[iobj]=3|Person[subj]=3|Tense=Past|VerbForm=Fin
(15)
და
da
მისცეს
mi-s-c-es
აღწერილი
aγ-c̣er-il-i
იგი
igi
წიგნი
c̣ign-i
ჭურჭრისაჲ
č̣urč̣r-isa-y
მის
mis
and pvb-io3-send-s3pl.aor pvb-write-pst.ptcp-nom det.nom letter-nom treasure-gen-nom det.gen
‘and they gave them the written inventory of the treasure’ (1Ezra 8:36)

Moreover, Old Georgian features an aspect-based split alignment system, with each series distinguished by its specific argument flagging and indexation pattern (Sarjveladze, 1997, 160–161; Schanidse, 1982, 66, 171–173). As for argument flagging, Series I screeves have Nom subjects, and Dat direct and indirect objects; Series II screeves have Erg subjects, Nom direct objects and Dat indirect objects; Series III screeves have Dat subjects and Nom direct objects. As for argument indexation, Series I and Series II screeves both cross-reference the subject with the ვ v set of affixes, the direct object with the მ m set, and the indirect object with the მი mi set; Series III screeves, on the other hand, cross-reference the subject with the მი mi set of affixes and the direct object with the ვ v set: this phenomenon is known as “inversion” (Imnaishvili and Imnaishvili, 1996, 210, 289).23 For instance, in example (16), აღგჳსრულებიეს aγgwisrulebies is thus annotated as:
Mood=Ind|Number[obj]=Sing|Number[subj]=Plur|Person[obj]=3|Person[subj]=1|Tense=Pf|VerbForm=Fin
(16)
და
da
მიერითგან
mierit=gan
ვაშენებთ
v-a-šen-eb-t
და
da
არღა
ar=γa
აღგჳსრულებიეს
aγ-gw-i-srul-eb-ies
and thereafter s1-ver-build-tm-s.pl and neg=ptc pvb-io1pl-complete-tm-s3sg.pf
‘and since then we are building it and we have not completed it yet’ (1Ezra 5:16)

Notably, the -გ -g 2nd person object marker denotes both singular and plural objects. Due to this morphological ambiguity, we decide to skip the annotation of the Number[obj] feature for verbs displaying that marker. For instance, გაუწყებთ gauc̣q̇ebt is annotated as:
Mood=Ind|Number[obj]=Plur|Number[subj]=Plur|Person[subj]=1|Tense=Pres|VerbForm=Fin
(17)
აწ
ac̣
გაუწყებთ
g-a-uc̣q̇-eb-t
თქუენ
tkuen
მეფესა
mepe-sa
now o2-ver-inform-tm-s1pl 2pl king-dat
‘now we inform you, the king’ (1Ezra 4:16)

4.7 Dependency Relations
In general, we follow UD_Georgian-GNC for the annotation of DEPREL. Here we discuss the following two special cases: constructions with the verb ყოფა q̇opa (4.7.1), and masdars (4.7.2).
4.7.1 Constructions with the auxiliary and the copula
In Old Georgian, the verb ყოფა q̇opa can occur as an auxiliary with participles or as a copula in noun predicates. In such cases, we annotated it as AUX in the UPOS and as aux or cop in the DEPREL field, respectively, as shown in examples (18–19).
(18)
და
da
ყვეს
q̇v-es
დღესასწაული
dγesasc̣aul-i
ტაძრობისაჲ,
ṭaʒrob-isa-y,
ვითარცა
vitarca
წერილ
c̣eril
არს
ar-s
and do-aor.s3pl festival-nom temple-gen.sg-nom.sg like write-pst.ptcp.abs be-pres.s3sg
‘and they celebrated the festival of the temple as it is written’ (1Ezra 3:4)

(19)
და
da
ესე
ese
არს
ar-s
წელი
c̣el-i
ბრძანებისაჲ
brʒaneb-isa-y
and this.nom be-s3sg.prs year-nom order-gen-nom
‘and this is the year of the order’ (1Ezra 7:11)

4.7.2 Masdar
The verbal noun in Georgian has both verbal and nominal characteristics. For this reason, linguists often refer to it as the “masdar” (Georgian მასდარი masdari), a term borrowed from the Arabic grammatical tradition via Semitic linguistics. A masdar in Old Georgian exhibits nominal character in that it can be declined like a full-fledged noun and take nominal dependents:
(20)
აწ
ac̣
ამიერითგან
amier-it=gan
მიუშჳთ
mi-u-šw-it
შენებაჲ
šen-eb-a-y
სახლისა
saxl-isa
მის
mis
ღმრთისაჲ
mrt-isa-y
მთავართა
mtavar-ta
მათ
mat
now this-ins=from pvb-io3-allow-s2pl.ipv build-tm-masd-nom house-gen det.gen God-gen-nom leader-obl.pl obl.pl
‘Now, from this point forward, allow the building of that house of God by the leaders’ (1Ezra 6:7)
A masdar in Old Georgian exhibits verbal character in that, when inflected in the adverbial case, it functions like the complementary infinitive of Indo-European languages:
(21)
და
da
ნუცა
nu=ca
ეძიებ
e-ziʒ-eb
ყოფად
q-op-ad
მშჳდობისა
mšwidob-isa
მათ
mat
თანა
tana
ყოველთავე
qove-ta=ve
ჟამთა
žam-ta
and now=ptc ver-seach-tm.prs.s2sg do.tm-masd.adv peace-gen.sg dem.obl.pl with all-obl.pl=ptc time-obl.pl
‘And you shall not seek to make peace with them at any time.’ (1Ezra 9:12)
This gives rise to a notational complication when a masdar takes multiple objects, as in the following example:
(22)
და
da
დაიდვა
da-i-dv-a
ეზრა
Ezra
გულსა
gul-sa
თჳსსა
twis-sa
(…)
(…)
სწავლად
sc̣avl-ad
ისრაელისა
Israel-isa
ბრძანებანი
brʒaneba-n-i
მისნი
mis-n-i
and pvb-ver-set-s3sg.aor Ezra.abs heart-dat refl.gen-dat teach-masd.adv Israel-gen statute-pl-nom det.gen-pl-nom
‘And Ezra set in her heart (…) to teach Israel his statutes.’ (1Ezra 7:10)

In the example above (22), the masdar სწავლად sc̣avlad ‘to teach’ governs two arguments: the indirect object ისრაელისა israelisa ‘Israel’ and the direct object ბრძანებანი brʒanebani ‘statutes’.
In UD_Georgian-GNC, masdar is treated as NOUN both in the UPOS and XPOS. Given the verbal nature of masdar, we propose to annotate a masdar as VERB in the UPOS and V_Masd in the XPOS. Notably, in Old Georgian some masdars are highly lexicalized and refer to specific objects, as in the case of ბრძანებანი brʒanebani in example (22). In such cases, Fähnrich and Sardshweladse (2005) includes double lemmas for them: ბრძანება brʒaneba ‘to order’ vs. ბრძანებაჲ brʒanebay ‘statute’. We therefore decide to annotate lexicalized masdars as NOUNs.
4.8 LLM-assisted annotation
4.8.1 Qualitative analysis
The initial draft annotations of the subsequent verses prepared by the LLM (see Section 3.2) presents issues with the recognition of historical graphemes (ჲ ⟨Y⟩, ჱ ⟨ē⟩, ჳ ⟨wi⟩, ჴ ⟨q⟩, ჵ ⟨ō⟩), which were automatically normalized (ი ⟨i⟩, ე ⟨e⟩, ვი ⟨vi⟩, ხ ⟨x⟩, ო ⟨o⟩). We therefore specify in the prompt to preserve the Old Georgian graphemes.
A second issue concerns our annotation of Suffixaufnahme (see Section 4.3.1), which is not automatically recognized and annotated by the LLM.
Another issue concerns the annotation of personal indexes. Apparently, the LLM initially annotates personal indexes based on the arguments expressed in the syntax rather than on morphological markers. This is especially prominent for 3rd person direct objects in series II verbal forms and for 3rd person indirect objects. In the first case, although series II forms have zero marked direct objects in the singular, the LLM initially annotates them as O:3Sg in XPOS and Person[obj]=3|Number[obj]=Plur in FEATS (see example 23). In the second case, the LLM fails to annotate 3rd person indirect objects marked on the verb when they are not expressed in the syntax.
(23)
რაჲთა
rayta
უშჱნო
u-šēn-o
მას
mas
ტაძარი
ṭaʒar-i
so_that io3-build-s2.subj det.dat temple-nom
‘so that you build a temple to Him’(1Ezra 1:2)

For უშჱნო u-šēn-o, the following is the LLM output: (XPOS): V_Act_Aor_Pv_S:3Sg_O:3Sg_IO:1Sg(FEATS): Mood=Ind|Number[iobj]=Sing|Number[obj]=Sing|Number[subj]=Sing|Person[iobj]=1|Person[obj]=3|Person[subj]=3|Tense=Past|VerbForm=Fin
This is the corrected output according to our guidelines:(XPOS): V_Act_Aor_Pv_S:3Sg_IO:1Sg(FEATS): Mood=Ind|Number[iobj]=Sing|Number[subj]=Sing|Person[iobj]=1|Person[subj]=3|Tense=Aor|VerbForm=Fin
The above-mentioned issues are resolved after including a sufficient number of correctly annotated examples in the prompt.
4.8.2 Quantitative evaluation
To evaluate the precision of our LLM-assisted annotation, we conducted the following experiment: Ezra 1–8 was used as the training set, and two different LLMs (Claude Opus 4.7 and Gemini-3.1-Pro-Preview) were prompted to annotate Ezra 9–10. We then ran the official UD eval.py on their outputs as predictions, with the manually corrected Ezra 9–10 serving as the gold standard test set. The LLM outputs contained dependency cycles and were missing some multi-word token range rows, which prevented eval.py from executing properly. To resolve this, we prompted the LLMs again to break the dependency cycles and to reattach one of the affected tokens to the nearest plausible head prior to running the evaluator, as well as to add the multi-word-token rows, without consulting the gold standard test set.24
The results are summarized in Table 4. In all the columns of UD annotation, Claude (Opus 4.7) outperforms Gemini (3.1-Pro-Preview), and achieves at least 80% accuracy. Noteworthy is that Gemini scores only 63.92% in the XPOS column, which implies that Gemini did not learn the pattern of our guidelines specifically designed for Old Georgian XPOS.
4.9 UDPipe training
We train a UDpipe model in R (package udpipe version 0.8.16) and in Python (UDPipe 1) using Ezra 1–8 as the training set, Ezra 9 as the test set and Ezra 10 as the development set. The accuracy achieved by the model is summarized in Tables 5 and 6. Overall, the output of UDPipe trained on our corpus is substantially inferior to the LLM’s output, especially in the syntactic annotation (UAS and LAS scores). This is probably due to the limited size of our training data.
5 Implications and Applications
This dataset is a UD treebank for the Book of Ezra in Old Georgian and can serve as the gold standard for the future expansion of the LLM-assisted annotations of the Old Georgian Oshki Bible. The dataset can be used as training data for language-specific UDPipe models and for downstream NLP tasks such as dependency parsing, PoS tagging, and lemmatization. For linguistic studies, the treebank provides data for the statistical and quantitative analysis of various morphosyntactic phenomena, such as verbal agreement patterns. Moreover, the dataset can be used for comparative studies with other UD treebanks of Bible translations, such as those in the PROIEL project, and for diachronic research tracing morphosyntactic change from Old Georgian to Modern Georgian. The successful integration of a LLM into the annotation workflow demonstrates that similar approaches can be adopted by linguists and domain experts working on other low-resource languages, helping to facilitate and accelerate the digitization and creation of digital corpora for minority and historical languages. Such efforts contribute to democratizing knowledge of understudied languages and extending language technology beyond high-resource modern languages such as English and Chinese.
Notes
[2] https://github.com/UniversalDependencies/UD_Old_Georgian-GLC; https://universaldependencies.org/oge/index.html.
[6] For instance, we would like to fully integrate the Suffixaufnahme phenomenon in the UD annotation, which is an important typological feature of Old Georgian. However, this is left out in the UD_Old_Georgian-GLC treebank. See Section 4.3.1.
[7] Other treebanks on pre-modern languages can be found at https://universaldependencies.org/: Akkadian, Ancient Hebrew, Classical Sanskrit, Hieroglyphic Egyptian, Hittite, Middle French, Old East Slavic, Old English, Old French, Old Occitan, Old Turkish, Ottoman Turkish.
[8] Model: gemini-3.1-pro-preview, accessed March 2026. The full prompt is available as Ezra_prompt.md in our dataset release. We chose Gemini over other models (Claude, ChatGPT) based on a zero-shot comparison; we find that Gemini outperforms the alternatives on Natural Language Processing (NLP) tasks involving low-resource languages. For benchmarks supporting this, see Maheshwari et al. (2026), Ojo et al. (2025), Adelani et al. (2025).
[12] cf. Fähnrich (2012, 592–593) and Hewitt (1995, 33–35); we thank the anonymous reviewer for their helpful comment on this point.
[13] Notably, dative, genitive and instrumental can have an extension marker -ა -a, which is not a case value. We thank the anonymous reviewer for pointing this out.
[14] cf. Boeder (1995); Plank (1995); Schanidse (1982, §276).
[15] In UD_Old_Georgian-GLC, sometimes the native FEATS field Case is used, but sometimes it refers to the genitive flag (line 145, მღდელობისასა mγdelobisasa Case=Gen; the Suffixaufnahme itself is in dative), and sometimes to the Suffixaufnahme itself (line 164, ეპისკოპოსობისაჲთა eṗisḳoṗosobisayta Case=Ins; the genitive flag is not annotated). Sometimes it uses a dedicated field Case[stack] (line 197, მოხუცებულისასა moxucebulisasa Case=Dat|Case[stack]=Gen). However, Case[stack] is always annotated as Gen, thus referring to the genitive flag – which should be annotated in the Case field – rather than the Suffixaufnahme itself. Moreover, UD_Old_Georgian-GLC lacks a separate field dedicated to the number of the Suffixaufnahme, which can differ from the number of the flag (line 2926, სიტყუათასა siṭq̇uatasa Case=Dat|Case[stack]=Gen|Number=Plur, but the flag -თა -ta is genitive plural and the Suffixaufnahme -სა -sa is dative singular).
[17] cf. Boeder (1995, 186).
[19] See Fähnrich (2012, 113–114); Schanidse (1982, 51-52); Tuite (2004, 955).
[20] Broadly speaking, active verbs correspond to transitive verbs; passive verbs include true passive and unaccusative verbs. The class of middle verbs, corresponding to unergatives, was grouped together with actives.
[21] See Imnaishvili and Imnaishvili (1996, 78–267).
[23] As noted by Cole et al. (1980, 738ff.), in Old Georgian, the dative argument of Series III screeves exhibits syntactic subject properties, even if it does not trigger plural agreement with the verb, as it happens in Modern Georgian.
Acknowledgements
We would like to thank Dr. Paul Meurer for discussing the project with us and sharing his experience from the GNC project, and Dr. Jesse Wichers Schreur for putting us in contact with Dr. Meurer. We are also grateful to Prof. Dr. Ingo Strauch and Prof. Luca Alfieri for their guidance and support. Furthermore, we would like to thank the anonymous reviewers for their insightful comments and suggestions, which helped to improve the quality of this work. Finally, we thank Dr. Adam Bremer-McCollum and the other friends of our Old Georgian reading group for their insightful discussions and our weekly readings of Old Georgian texts.
Author Contributions
The paper is the result of the joint work of both authors, who contributed equally. For academic purposes, the responsibilities are distributed as follows: Chia-Wei Lin is responsible for Sections 1, 2, 3, 4.1, 4.8, 5; Diego Luinetti is responsible for Sections 4.2-7.
