(1) Context and motivation
(1.1) Context: microdiachrony and Lex-Br-Ius corpus
Among corpus linguistics, sociolinguistics, diachronic studies, lexicology and others, the term microdiachrony has been increasingly used in the last ten years indicating an empirical method of analysis that identifies lexical, morphosyntactic but also phonetic and phonological changes in a short period of time using a corpus (Labov, 1994; Sankoff, 2006; Mihatsch, 2020).
Legal language is an umbrella term that includes a series of very specialized languages, such as the one used in laws, in courts, but also in depositions or by lawyers and police in their daily work. Linguists have found a major field of research in this kind of language, as it allows for a panoramic view of its characteristics, given that it is a very specialized genre (Carapinha, 2018), and also to compare it interlinguistically, especially for translation purposes (Biel, 2009; Giampieri, 2018). Numerous study groups around the world have compiled corpora for the study of legal language (Pontrandolfo, 2012; Lorz, 2019; Goźdź-Roszkowski, 2011, 2012, 2021; Goźdź-Roszkowski and Pontrandolfo, 2014, 2015), including in Brazil, where the interest is mostly directed to lexical research and the creation of glossaries and dictionaries (Krieger et al., 1998; 2006, 2008; Maciel, 2001) or for translation purposes (Carvalho, 2006, 2014). Another field in large expansion is Natural Language Processing, that collects data and develops models for legislative and judicial usage (Martim et al., 2018). Despite this growth, when we started our research, we found a lack of corpora specifically designed to describe the characteristics of Brazilian Portuguese legal language (Sousa & Fabro, 2019; Ferrari & Cunha, 2022).
The lex-Br-Ius corpus (Ferrari & Marques, 2022) contains 754 legal text (3,296,993 tokens) manually extracted from Brazilian federal statutory laws in force on the compilation date, collected in their entireness at Portal da Legislação (Legislation Portal) (https://www4.planalto.gov.br/legislacao/), the governmental website that hosts and continually updates Brazilian federal laws.
Data collection and processing occurred between January 2022 to January 2024: the compilation first checked each of the 13,833 texts in the portal discarding the ones that were no longer in effect. Then, since as representativeness criterium was chosen the employment of each law in the legal and judicial world (Barbera & Onesti, 2009; Onesti, 2011), every law use frequency was checked in the Jusbrasil portal (https://www.jusbrasil.com.br/), the biggest and most important portal of Brazilian law-related content, and the ones with over 10,000 citations were selected.
As Table 1 above shows, the corpus is divided in six major categories that follow the Portal da Legislação (Legislation Portal) classification. 13 of 17 Códigos (Codes) and 11 of 18 Estatutos (Statutes) were selected for their importance and wider scope in society. Sampling Emendas à Constituição (Amendments to the Constitution), Leis Complementares (Complementary laws), and Leis Ordinárias (Ordinary laws) was more difficult because they have a more specific scope and only a very limited subset met our criteria.
Table 1
The LEX-BR-Ius Corpus.
| SUB-CORPUS | AVAILABLE TEXTS | SELECTED TEXTS | WORDS | WORDS (SMALLEST TEXT) | WORDS (LONGEST TEXT) | AVERAGE |
|---|---|---|---|---|---|---|
| Códigos | 17 | 13 | 711,952 | 12,225 | 167,416 | 54,765.50 |
| Constituição | 1 | 1 | 97,082 | N/A | N/A | N/A |
| Emendas à Constituição | 114 | 48 | 50,241 | 43 | 12,015 | 1,046.69 |
| Estatutos | 18 | 11 | 125,356 | 2,370 | 35,241 | 11,396.00 |
| Leis Complementares | 192 | 43 | 182,768 | 90 | 35,837 | 4,250.42 |
| Leis ordinárias | 13,491 | 638 | 2,129,594 | 37 | 86,611 | 3,337.92 |
| TOTAL | 13,833 | 754 | 3,296,993 | 37 | 167,416 | 3,372.67 |
Law’s dimensions vary from a few dozen words (e.g. EC 90_15.09.2015,1 with 56 words) to hundreds of thousands of words (e.g. LO10.406_10.01.2002,2 with 101,629 words). As can be seen, balance was not achieved in this corpus. This was a deliberate methodological decision since, to preserve laws’ internal organization, textual variation, representativeness, and structure (Sinclair, 2004; Biber, 1993), they were maintained complete (Ferrari & Marques, 2022).
Processing included, for each law, the creation of a folder containing different files resulting from: (a) the semi-automatically cleaning all extratextual information and the creation of a file of raw text in UTF8 format named after the number and date of the law; (b) the manual filtering of metadata in a separate file in XML format; (c) the manual tagging of each text following the Modest XML proposition of Hardie (2014) with a tag set especially created for the project (Marques, 2023). This tag set is meant to signal specific text boundaries and is particularly important for this research, as it also indicates, when available, the precise date of the amendment.
We will now briefly describe two previous studies using the LEX-BR-Ius corpus that mapped some of its general linguistic characteristics and manually verified possible microdiachronic features in a minor sample. A Multidimensional Analysis comparing LEX-BR-Ius with the Brazilian Register Variation Corpus (CBVR) (Marques & Ferrari, 2023; Marques et al., 2025) confirmed its literacy, specialized language, high informational density, elevated accuracy, and elaborated syntactic structures. The level of argumentation doesn’t seem dominant as law’s main purpose is establishing and informing rules and not discussing them, a stage that occurs during the legislative process. The corpus also scored very highly in the future time orientation, as this feature establishes the correct application and its consequences in concrete cases of the laws.
A microdiachronic preliminary study (Ferrari et al., 2026) focused on two specific laws to manually analyze their verbal forms in order to have a panoramic view of the distinctive features of this genre. Among other statistically ambiguous traits, a considerable decrease was found in the use of the reflexive voice, and partially of the passive voice. Within the passive voice, an increase in the use of the ser (to be) auxiliary was observed regarding the clitic reflexive passive se, likely due to a rearrangement of pronominal reflexive particles in Brazilian Portuguese (Castilho, 2010). Regarding subordination, especially in adverbial clauses of manner, condition and cause, an increase in the use of infinitives and a decrease in the use of gerunds were detected. Since the different use of these moods and tenses seems to be more diatopic (Brazilian Portuguese vs. European Portuguese) than diachronic (Móia & Viotti, 2004), the results were not conclusive on the cause of this difference. The research also confirmed a greater use of the future tense and a decline in the use of the present tense, the same result shown in the Marques et al. (2025) study. According to the Manual de Redação da Presidência da República (2018) (the Style Manual of the Presidency of the Republic), the present tense is preferred to state the law’s reason for being, its main objectives, and the permanent rights and duties it creates. The future tense, in turn, is used for actions that will be performed once, are dependent on a future fact, or are not the law’s central purpose. One hypothesis for this change is that recent modifications have specified the applications of laws and their consequences after promulgation, thereby demanding the use of the future tense.
(1.2) Motivation: the CLT pilot study
In the present research we focused on the Consolidação das Leis do Trabalho (CLT: Decreto-Lei nº 5.452/1943), the Brazilian labor law. The choice of this specific statute is due to its time span, its great number of alterations and its dimension, ideal for a computational pilot study.
CLT origin dates to 1943, under the civil dictatorship of Getúlio Vargas, when a series of labor measures were unified into a single, highly detailed law that, at the time of its enactment, comprised 921 articles. As will be seen in details in the Method section, a great number of changes occurred during a new democratic period known as the Populist Republic (1946–1964), but it was during the military dictatorship (1964–1985) and the subsequent period of Brazil’s redemocratization (1985–present), specifically in 2017 and 2019, under the labor law reforms of President Temer and President Bolsonaro, that the most significant changes were made.
This pilot study partially followed the above-mentioned periodization, dividing data into four intervals: (a) 1943–1962: from the enactment of the statute to the last modification under the democratic regime; (b) 1964–1986: from the beginning of the military dictatorship until the ultimate amendment before the new democratic Constitution; (c) 1988–2016: remembering that the 1988 Constitution marks the start of the current democratic regime, with different presidents, until the impeachment of President Roussef; (d) 2017–2022, including the massive neoliberal labor revisions.
We remember that Table 2 shows the number of changes made for period.
Table 2
Number of Constitutional Amendments across legislative periods.
| PERIOD | NUMBER OF AMENDMENTS |
|---|---|
| 1943–1962 | 5 |
| 1964–1986 | 55 |
| 1988–2016 | 30 |
| 2017–2022 | 48 |
[i] Note: Source: the authors. Periods correspond to distinct political/legislative eras as defined in the original analysis.
The main goal of this study is to investigate complexity-related features during time in the CLT. In face of recent plain language movements arguing to more comprehensible texts in legal language (Benson, 1985; Garner, 2013; Adler, 2012; Hartig & Lu, 2014), our research question was: “does Brazilian Portuguese laws language simplified or became less complex in the last 80 years?”. To do so, we consulted a bibliography on the subject focusing on legal language and defined a set of characteristics described as typical of this kind of text, as follows.: Word Order Inversion (VS), Nominalizations, Subordinate Clauses, Passive Constructions, Gerunds & Participles, Relative Pronouns and Conjunctions, Appositions and Parentheticals, Prepositional Phrases, Complex Verb Forms (Bhatia, 1993; Almut, 2010; Chovanec, 2013; Fanego & Rodríguez-Puente, 2019; Busso, 2022a, 2022b).
Using the Modest XML tags that mark legislative amendments, we divided the CLT corpus into the four periods analyzed in this study. After removing XML markup and legal formatting conventions, we obtained continuous text suitable for grammatical analysis. We then used UDPipe to annotate the texts and extract the complexity features described above, normalizing the measures to enable comparison across periods.
(1.3) Content of this work
In the following sections, after the corpus technical specification, we will present a step-by-step of the procedures adopted in order to clean and obtain the modified files per year and the grammatical complexity analysis, complemented by the statistical analysis. We then briefly discuss the results and future applications.
(2) Dataset description
Repository location – https://doi.org/10.5281/zenodo.19698539
Repository name – Zenodo
Object name – LEX-BR-Ius Corpus
Format names and versions – .txt, .xml, .docx
Creation dates – January 2022 to January 2024
Dataset creators – Lúcia de Almeida Ferrari (UFMG, editor, data collector), Carolina Godoi de Faria Marques (UFMG, editor, data collector); Nayane Araujo (data curator); Luisa Ramos de Oliveira (data curator), Lara Siewerdt Cotting (annotator), Aline Barbosa Ricardo (annotator).
Language – Brazilian Portuguese
License – CC BY 4.0
Publication date – The corpus was published on Zenodo on 2026-04-21.
(3) Method
(3.1) Text clean-up and normalization
We first normalized the compiled legal text to remove elements that are typical of legal drafting and XML markup but not directly relevant to our grammatical analyses.3 Using a series of regular-expression–based scripts, we:
deleted structural tags and layout markers (titles, chapter headings, etc.) that remained from the XML compilation;
removed article and paragraph identifiers (e.g. “Art. X”, paragraph signs, item numbering with Roman or Arabic numerals);
stripped out editorial and legislative notes in parentheses, such as references to the law or decree that introduced a particular wording, indications of vetoes, or remarks about changes in redaction; and
standardized or eliminated stray punctuation and special characters introduced during XML processing.
The goal of this stage was to obtain continuous running prose that preserves the legal content while minimizing noise from formatting conventions and metadata.
(3.2) Splitting into diachronic subtexts and manual verification
The cleaned text was then divided into a baseline text and a set of diachronic subtexts based on the year associated with each modification. Because the XML markup already distinguished between original text and modified passages and encoded the year of modification, we used these tags to automatically extract all segments corresponding to a given year and group them into separate text files. In this way, we created one text representing the original 1943 version of the code and 67 additional subtexts containing only the amending text, each corresponding to a specific year between the 1940s and 2022. This yields a total of 68 comparable text files (one baseline plus 67 yearly amendment corpora), with an irregular distribution across years that reflects the actual legislative activity.
After the automatic split, every yearly subtext and the 1943 baseline were checked manually against the compiled source text. A research assistant compared the content of each file to the original XML to confirm that all modified segments had been correctly assigned to their year, that no passages were missing or duplicated, and that the original-only text did not contain any text that should have been classified as a later modification. This manual verification also allowed us to identify and correct any lingering XML or tokenization issues that were not resolved by the automated scripts.
(3.3) Automatic annotation with UDPipe and quality control
Once the diachronic texts had been cleaned and manually verified, we automatically annotated all files using UDPipe (Straka & Straková, 2017). We employed the pre-trained UDPipe model for the UD_Portuguese-Bosque treebank (Rademaker et al., 2017), which is based primarily on European Portuguese, to perform sentence segmentation, tokenization, part-of-speech tagging, lemmatization, morphological analysis, and dependency parsing in a single pipeline.
In practice, we accessed UDPipe via the Claude large language model, using it as an interface to call UDPipe with a fixed Portuguese-Bosque model and to return the output in standard Universal Dependencies (CoNLL-U) format. Each of the 68 text files (the 1943 baseline plus all yearly amendment texts) was processed with identical model settings, ensuring comparability across years.
For each token in each sentence, the annotation provides its surface form, lemma, universal part-of-speech tag, language-specific tag (where applicable), a bundle of morphological features (such as number, gender, tense, mood, and person), the index of its syntactic head, and the label of the dependency relation linking it to that head (Figure 1). Together, these annotations constitute a syntactically and morphologically enriched diachronic text of the labor code.

Figure 1
Sample of annotation of a sentence from the subtexts 1943.
To assess the quality of the automatic annotation, we conducted manual spot-checks on a sample of sentences from different periods of the corpus, with particular attention to constructions central to our research questions (e.g. verbal tense and voice, subject position, pronominal subjects, and subordination).
(3.4) Grammatical complexity analysis
After UDPipe annotation, we quantified grammatical complexity in the diachronic texts using a custom Python program that operates directly on the CoNLL-U output (Nivre et al., 2020). The program reads one annotated file per year and aggregates a set of syntactic and morphological indicators of complexity across time. For each year, it computes both raw counts and normalized frequencies (per 100 tokens), as well as measures of sentence length and depth of subordination.
The script first parses each CoNLL-U file sentence by sentence, using the sentence boundaries and token-level information provided by UDPipe. Multi-word tokens (e.g. contractions represented by token ranges such as “6–7”) are skipped, so that all counts are based on syntactic tokens only. For every sentence, the program extracts the token ID, lemma, part-of-speech tags, morphological features, dependency head, and dependency relation, and stores them in memory for subsequent calculations.
On this basis, the script computes a series of complexity-related features in each sentence:
Word order inversion (VS order). Using the dependency relations that mark subjects (nsubj: nominal subject in an active construction; nsubj:pass: nominal subject in a passive construction; and csubj: clausal subject), the program compares the position (token index) of each subject to the position of its governing verb. Instances in which the verb precedes the subject (verb–subject order) are counted as non-canonical word order and treated as one type of syntactic complexity.
Nominalizations. To capture the prevalence of deverbal and abstract nouns, the program scans all tokens tagged as nouns and checks their lemmas for a set of frequent Portuguese nominalization suffixes (e.g. -ção, -mento, -ncia, -dade, -eza, -ismo, -agem, -ura, -ância). Each noun lemma ending in one of these patterns contributes one nominalization count.
Subordination. Different types of subordinate clauses (relative, adverbial, complement, etc.) are detected via their Universal Dependencies labels. Tokens whose dependency relation indicates clausal dependency, e.g. acl (adjectival clause), acl:relcl (relative clause), advcl (adverbial clause), ccomp (clausal complement), xcomp (open complement – no subject), csubj (clausal subject – active), csubj:pass (clausal subject – passive), are counted as instances of subordination. This yields a sentence-level measure of how many subordinate clauses are present.
Passive constructions. The script identifies both analytic and synthetic passives. Analytic passives are captured via the presence of passive subjects (nsubj:pass4), while synthetic passives with the clitic “se” are detected when “se” appears as a passive expletive (expl:pass5). Each such occurrence is counted as one passive construction.
Gerunds and participles. To measure the use of non-finite verb forms, the program examines the morphological features of all verbs and auxiliaries. Tokens marked as gerunds (VerbForm = Ger) or as participles (VerbForm = Part) are counted.
Relative pronouns and subordinating conjunctions. Relative pronouns were identified using the Universal Dependencies morphological feature (PronType = Rel), which marks relative clause constructions directly in the UD annotation scheme. Subordinating conjunctions were identified through the (SCONJ) part-of-speech tag and a small set of legal Portuguese subordinators (e.g. “embora”, “posto que”, “visto que”). Both categories were treated as markers of subordination.
Appositions and parentheticals. Using the dependency labels appos and parataxis, the program counts appositive structures and other parenthetical segments, which increase local structural complexity and density of information.
Prepositional Phrases. Following the Universal Dependencies framework, the script identifies prepositional phrases through case dependency relations linking adpositions (ADP) to their nominal heads. In UD annotation, the noun functions as the syntactic head of the phrase, while the preposition is attached as a case dependent. The program therefore detects nominal heads that contain adpositional dependents and reconstructs the corresponding phrase for qualitative inspection (e.g. “de qualquer procedência”, “das cidades”). This measure serves as a proxy for lexical and structural density in legal language.
Complex verb forms. The program focuses on morphologically marked verb forms associated with legal style and syntactic complexity. Tokens annotated with subjunctive mood (Mood = Sub), conditional mood (Mood = Cnd), or future tense (Tense = Fut) in verbs and auxiliaries are counted as instances of complex verb forms. This operationalization captures the prevalence of morphologically marked verbal constructions characteristic of normative and hypothetical legal discourse.
For each sentence, the program also records basic structural metrics:
Sentence length, measured as the number of syntactic tokens per sentence.
Depth of subordination, defined as the maximum depth of embedding in the dependency tree. To compute this, the script builds the dependency tree for each sentence and iteratively calculates, for every token, the distance (in number of arcs) from the root. The maximum value across all tokens in the sentence corresponds to the deepest level of syntactic embedding.
All these sentence-level counts are accumulated by year. For each year, the program keeps track of the total number of sentences, the total number of tokens, and the sum of each complexity feature, as well as the maximum sentence length and maximum subordination depth observed in that year. After processing all years, the script derives normalized metrics:
average sentence length (tokens per sentence);
average subordination depth and maximum subordination depth;
for each complexity feature, the average number of occurrences per sentence; and
for each feature, the frequency per 100 tokens.
These normalized measures allow direct comparison of grammatical complexity features across years. Our final analysis consists of lists for each year, presenting the main structural statistics (number of sentences, tokens, average sentence length) and the per-sentence and per-token rates of each complexity indicator, as well as a summary of temporal trends (e.g. relative increases or decreases between the earliest and latest years).
(4) Results and discussion
Table 3 presents the basic text statistics for each of the four analytical periods. The corpus comprises 67 modified legal files distributed unevenly across periods, ranging from 5 files in 2017–2022 to 26 in 1988–2016. Despite having the fewest files, the most recent period contains a comparatively high number of sentences (1,361) and tokens (45,360). The Type-Token Ratio (TTR), that measures the lexical diversity in a text, varies across periods, peaking in 1988–2016 (0.2087) and reaching its lowest value in 2017–2022 (0.0803). Although these differences should be interpreted with caution, as TTR is sensitive to text size, since longer texts tend to have lower TTR because as tokens increase, new types appear less frequently, apparently the last period measures indicate a lower lexical diversity within the texts. We advise that the uneven distribution of files across periods reflects the availability of documents for each era rather than a deliberate sampling decision.
Table 3
Period Statistics per Period.
| PERIOD STATISTICS: FILES, TOKENS, TYPES, TTR | |||||
|---|---|---|---|---|---|
| PERIOD | FILES (MODIF.) | SENTENCES | TOKENS | TYPES (APPROX.) | TTR (APPROX.) |
| 1943–1962 | 15 | 1044 | 29822 | 3995 | 0.134 |
| 1964–1986 | 21 | 1847 | 58560 | 8211 | 0.1402 |
| 1988–2016 | 26 | 783 | 29332 | 6121 | 0.2087 |
| 2017–2022 | 5 | 1361 | 45360 | 3643 | 0.0803 |
[i] Note: TTR = Type-Token Ratio.
(4.1) Overall Results of Complexity Features
Table 4 reports the mean frequency of each syntactic complexity feature per 100 tokens across the four periods. Prepositional Phrases are by far the most frequent feature throughout the entire text (ranging from 16.7 to 17.9), followed by Subordinate Clauses and Nominalizations (both averaging approximately 5.5 per 100 tokens), and Gerunds & Participles (approximately 5.0). At the lower end, Word Order Inversion, Passive Constructions, and Complex Verb Forms occur less than once per 100 tokens in all periods. Notably, most features remain relatively stable across periods, with no dramatic shifts in raw frequency. This surface stability contrasts with the composite PCA results reported below, which suggest that the combined pattern of these features does shift meaningfully in the most recent period, underscoring the value of a multivariate approach over examining individual features in isolation.
Table 4
Complexity Features Normalized Frequency per Period.
| COMPLEXITY FEATURES — MEAN PER 100 TOKENS | ||||
|---|---|---|---|---|
| FEATURE | 1943–1962 | 1964–1986 | 1988–2016 | 2017–2022 |
| Word Order Inversion (VS) | 0.268 | 0.360 | 0.361 | 0.373 |
| Nominalizations | 5.771 | 5.932 | 5.615 | 5.721 |
| Subordinate Clauses | 5.858 | 5.605 | 5.117 | 5.258 |
| Passive Constructions | 3.947 | 3.967 | 3.948 | 3.746 |
| Gerunds & Participles | 4.973 | 4.964 | 5.097 | 5.077 |
| Relative Pronouns & Conj. | 0.899 | 0.750 | 0.610 | 0.701 |
| Appositions & Parentheticals | 1.462 | 1.679 | 2.063 | 1.748 |
| Prepositional Phrases | 16.873 | 16.665 | 17.674 | 17.866 |
| Complex Verb Forms | 1.724 | 1.984 | 1.841 | 1.724 |
[i] Note: Values represent normalized frequency rates per 100 tokens, extracted from CoNLL-U dependency parsed output using UDPipe (Straka & Straková, 2017) with the Portuguese Bosque model (Rademaker et al., 2017).
While the frequency data presented in Table 4 provides a useful overview of individual syntactic features, it does not readily reveal whether an overall shift in syntactic complexity occurred across periods. Examining nine features independently risks missing the broader structural pattern that emerges from their combination: features may move together, compensate for one another, or contribute jointly to an underlying dimension of complexity that no single measure can capture on its own. Therefore, we chose to run a PCA to identify the most relevant combinations of complexity features according to our texts.
(4.2) Principal Component Analysis (PCA)
Principal Component Analysis (PCA) is a multivariate technique designed to reduce the dimensionality of a dataset of interrelated variables while retaining as much of the original variation as possible (Jolliffe, 2002, p. 1). It does so by transforming the original variables into a smaller set of uncorrelated derived variables — the principal components (PCs) — ordered so that the first few components capture most of the total variance. Each PC is a linear combination of the original variables, derived from the eigenvectors of the covariance matrix, with the variance explained by each component equal to its corresponding eigenvalue (Jolliffe, 2002, pp. 2–6). When the original variables are substantially intercorrelated, as is common with linguistic features, the first component typically accounts for a large share of the total variance, providing a parsimonious composite index of the underlying shared dimension.
Principal Component Analysis (PCA) addresses these limitations by reducing the nine syntactic features to a smaller set of orthogonal components that capture the dominant axes of variation in the data. In particular, PC1 can be interpreted as a general syntactic complexity index, summarizing the shared variance across features into a single comparable score for each year. This allows for a more robust and nuanced diachronic comparison than raw or normalized frequencies alone would permit and forms the basis of the analysis reported in the present section.
Principal Component Analysis (PCA) was conducted on the nine standardized syntactic complexity features at the year level (N = 67 years). Components were retained following the Kaiser criterion, which specifies that only components with eigenvalues (λ) greater than 1.0 should be retained, as these account for more variance than a single original variable (Kaiser, 1960; Jolliffe, 2002). The analysis yielded a dominant first component (PC1; eigenvalue = 4.464), which alone accounted for 49.6% of the total variance, followed by a smaller second component (PC2; eigenvalue = 1.812, 20.14% of variance explained). Together, the first two retained components accounted for 69.74% of the total variance observed in the dataset, indicating that the complexity features share substantial underlying structure while still capturing partially distinct dimensions of syntactic variation. Although PC3 approached the Kaiser threshold (eigenvalue = 0.897), it did not exceed λ > 1 and was therefore not retained. The scree plot confirms this structure, showing a pronounced decline after the first two components and a gradual flattening of the remaining eigenvalues, suggesting that subsequent components capture progressively less meaningful variation.
To evaluate the internal consistency of the proposed complexity index, Cronbach’s alpha was calculated for the full nine-feature set. The initial model produced good reliability at the year level (α = 0.851), indicating substantial shared variance among the features. A backward-elimination procedure was then applied to identify the subset of features that maximized internal consistency. The iterative removal process showed that Word Order Inversion and Complex Verb Forms contributed least to the coherence of the composite, as their removal increased overall reliability. The final four-feature subset — consisting of Nominalizations, Subordinate Clauses, Gerunds & Participles, and Prepositional Phrases — yielded excellent reliability (α = 0.909). This optimized subset was therefore adopted as the primary composite syntactic complexity index employed in the diachronic analyses.
PC1 showed broadly similar loadings across all features, suggesting it captures a common underlying dimension shared by the indicators (interpretable as an overall syntactic complexity/density factor). Because PCA component signs are arbitrary, we oriented the component for interpretability (higher scores indicating higher complexity), regardless of directions: negative or positive.
To construct a parsimonious composite, we assessed internal consistency (Cronbach’s α) across the feature set. The full nine-feature scale showed good reliability at the year level (α = 0.851). We then applied backward elimination to maximize α, yielding a four-feature subset—Nominalizations (PC1 = –0.402), Subordinate Clauses (PC1 = –0.418), Gerunds & Participles (PC1 = –0.409), and Prepositional Phrases (PC1 = –0.413)—with improved reliability (α = 0.909). Features whose removal increased α were interpreted as contributing less consistently to the common scale (Figure 2).6

Figure 2
Scree Plot of Principal Components for Syntactic Complexity Features.
Note. Eigenvalues (λ) for each principal component extracted from a year-level PCA of nine syntactic complexity features. The red dashed line indicates the Kaiser criterion threshold (λ = 1), above which components are considered meaningful (Kaiser, 1960; Jolliffe, 2002, pp. 2–6). Three components meet this criterion, collectively explaining the majority of variance in the data.
Table 5 reports mean PC1 composite scores by period for both the full feature set (9 features) and the optimized subset (4 features: Nominalizations, Subordinate Clauses, Gerunds & Participles, and Prepositional Phrases). Both models converge on the same substantive pattern: the earlier periods (1943–1962, 1964–1986, and 1988–2016) display mean PC1 scores near or below the standardized mean, whereas the most recent period (2017–2022) shows a marked positive increase in syntactic complexity (all-features model: mean PC1 = 1.476) (Table 6). The optimized subset also yields lower within-period variability, suggesting that the removal of weaker or redundant features produces a cleaner and more internally consistent complexity signal. Given its higher reliability (Cronbach’s α = 0.909), the optimized four-feature composite was adopted as the primary syntactic complexity index for subsequent analyses, while the full-feature model is retained as a robustness check.
Table 5
Principal Component Loadings for the Syntactic Complexity Features.
| FEATURE | PC1 | PC2 |
|---|---|---|
| Word Order Inversion (VS) | –0.086 | –0.483 |
| Nominalizations | –0.402 | 0.213 |
| Subordinate Clauses | –0.418 | 0.004 |
| Passive Constructions | –0.345 | –0.354 |
| Gerunds & Participles | –0.409 | –0.202 |
| Relative Pronouns & Conj. | –0.284 | 0.413 |
| Appositions & Parentheticals | –0.269 | 0.176 |
| Prepositional Phrases | –0.413 | 0.265 |
| Complex Verb Forms | –0.215 | –0.533 |
[i] Note. PCA loadings for syntactic features across the first three principal components. Values represent standardized loadings; sign indicates direction and magnitude reflects contribution to each component. PC1 captures a broad shared variance across most features, while PC2 reflects more localized structural contrasts among subsets of constructions.
Table 6
PC1 Complexity Composite by Period — Best Subset: Nominalizations, Subordinate Clauses, Gerunds & Participles, and Prepositional Phrases.
| PERIOD | n YEARS | MEAN PC1 | SD | MIN | MAX |
|---|---|---|---|---|---|
| 1943–1962 | 15 | 0.067 | 1.223 | –2.140 | 2.291 |
| 1964–1986 | 21 | –0.088 | 1.522 | –2.802 | 2.879 |
| 1988–2016 | 26 | –0.251 | 2.230 | –6.534 | 2.806 |
| 2017–2022 | 5 | 1.476 | 0.922 | 0.317 | 2.790 |
[i] Note: Best subset comprises four features: Nominalizations, Subordinate Clauses, Gerunds & Participles, and Prepositional Phrases (α = 0.909).
Finally, visual inspection of the 95% confidence intervals suggests that the 2017–2022 period (M = 1.476) differs substantially from all earlier periods, with non-overlapping confidence intervals, indicating a potentially significant increase in PC1 scores. The three earlier periods (1943–1962, 1964–1986, and 1988–2016) show largely overlapping confidence intervals, suggesting no significant differences among them (Figure 3). However, given the small sample size of the 2017–2022 period (n = 5), these results should be interpreted with caution.

Figure 3
Distribution of PC1 Scores Across Historical Periods in Brazilian Labor Law.
Note. Each point represents the period mean; thick bars indicate 95% confidence intervals (CI), calculated as mean ± 1.96 × SE, where SE = SD/√n; thin bars indicate the observed min–max range. The red dashed line marks PC1 = 0, representing the grand mean of the standardized scores. Higher PC1 values indicate greater overall syntactic complexity.
(4.3) Discussion
The results reveal a counterintuitive pattern: while individual syntactic features remained largely stable in their raw frequencies across nearly eight decades, a meaningful shift in overall syntactic complexity emerges when features are examined in combination. This discrepancy between surface-level stability and multivariate change underscores the analytical value of composite indices over single-feature approaches in diachronic corpus research. The PC1 composite index shows that the three earlier periods (1943–1962, 1964–1986, 1988–2016) cluster close to zero, while the most recent period (2017–2022) displays a pronounced increase (mean PC1 = 1.476), driven by a combination of nominalization, subordination, non-finite verb forms, and prepositional density. The fact that documents from this period are also substantially longer on average suggests a broader densification of legal language rather than an increase in any single construction. The stability observed across the three earlier periods is also noteworthy: despite spanning different political contexts, the syntactic profile of Brazilian constitutional amendments remained remarkably consistent, suggesting that legal drafting conventions were resistant to political change until recently. Some limitations should be noted: the corpus is unevenly distributed across periods, no formal annotation reliability metrics were computed, and TTR differences across periods cannot be fully disentangled from corpus-size effects. Nevertheless, the convergence of full-feature and best-subset PCA results, together with the high internal consistency of the composite index (α = 0.909), supports the robustness of the main finding.
(4.4) Limitations
Although we performed a manual spot-check of the automatic annotations and corrected a small number of errors, we did not conduct a comprehensive validation study. In particular, we did not create a gold-standard reference set or compute formal reliability/accuracy metrics. Further studies should address this problem by adopting or creating such a reference. Annotation errors may therefore remain in the data and could introduce noise into the downstream analyses; results should be interpreted with this limitation in mind.
Additionally, we acknowledge that this study does not engage with the literature on computational readability assessment for Brazilian Portuguese, including the PorSimples project (Aluisio & Gasperin, 2010) and related work, nor does it employ more recent dependency parsers such as PortParser (Lopes & Pardo, 2024). These choices reflect the scope and nature of the study, whose primary objective was not readability prediction but the diachronic analysis of grammatical complexity in legal texts. Future research could build on the present work by incorporating Brazilian Portuguese readability frameworks and evaluating the robustness of the findings using these or newer parsing models.
(5) Applications
As discussed above, the Portuguese Bosque model was not trained on legal documents, so some degree of data skewing is possible. Therefore, it is advisable to perform a validation step before conducting new analyses on a larger portion of the LEX-BR-Ius corpus. Because the texts vary greatly in size, we suggest that future investigations focus on specific subcorpora, such as Códigos or Estatutos. Both the corpus and the scripts are publicly available, and we encourage other researchers to contribute to the analysis of microdiachronic characteristics of Brazilian legal language. The reuse of the corpus is also possible for lexical and morphosyntactic analysis or as raw data for training NPL modeling.
Notes
[9] More details on the scripts used and steps taken can be found on this GitHub repository.
AI Declaration
The authors declare the use of a Generative AI (GenAI) tool in the preparation of this manuscript, in accordance with Ubiquity Press’s policy on GenAI use.
Claude (Anthropic), a publicly available GenAI tool, was used to assist with the implementation of the automated syntactic annotation pipeline in conjunction with UDPipe and to provide language-editing support, including suggestions related to grammar, vocabulary, and clarity. The authors acknowledge that content submitted to publicly available GenAI platforms may be incorporated into future model outputs.
All research design, data collection, corpus preparation, annotation decisions, tagging and categorization of linguistic features, data analysis, interpretation of results, and final editorial decisions were carried out by the authors. GenAI tools were not used to generate, alter, or manipulate research data or research findings.
No GenAI tool has been listed as an author. The authors accept full responsibility for the content, accuracy, and integrity of this work.
Acknowledgements
We thank all the LEX-BR-Ius team, especially Carolina Godoi de Faria Marques for her precious contributions on the conception of the corpus and revision of its data. We also thank Ana Clara Scigliano Valerio Modesto for manually checking the compiled source text of this study.
Author Contributions
Lúcia de Almeida Ferrari was responsible for the conceptualization, funding acquisition, supervision, methodological decisions, analysis of results and writing of sections 1, 2, 4 and 5.
Luciana Dias de Macedo was responsible for data curation, methodology, formal analysis of results and writing of sections 3, 4 and 5.
