Table 1
The LEX-BR-Ius Corpus.
| SUB-CORPUS | AVAILABLE TEXTS | SELECTED TEXTS | WORDS | WORDS (SMALLEST TEXT) | WORDS (LONGEST TEXT) | AVERAGE |
|---|---|---|---|---|---|---|
| Códigos | 17 | 13 | 711,952 | 12,225 | 167,416 | 54,765.50 |
| Constituição | 1 | 1 | 97,082 | N/A | N/A | N/A |
| Emendas à Constituição | 114 | 48 | 50,241 | 43 | 12,015 | 1,046.69 |
| Estatutos | 18 | 11 | 125,356 | 2,370 | 35,241 | 11,396.00 |
| Leis Complementares | 192 | 43 | 182,768 | 90 | 35,837 | 4,250.42 |
| Leis ordinárias | 13,491 | 638 | 2,129,594 | 37 | 86,611 | 3,337.92 |
| TOTAL | 13,833 | 754 | 3,296,993 | 37 | 167,416 | 3,372.67 |
Table 2
Number of Constitutional Amendments across legislative periods.
| PERIOD | NUMBER OF AMENDMENTS |
|---|---|
| 1943–1962 | 5 |
| 1964–1986 | 55 |
| 1988–2016 | 30 |
| 2017–2022 | 48 |
[i] Note: Source: the authors. Periods correspond to distinct political/legislative eras as defined in the original analysis.

Figure 1
Sample of annotation of a sentence from the subtexts 1943.
Table 3
Period Statistics per Period.
| PERIOD STATISTICS: FILES, TOKENS, TYPES, TTR | |||||
|---|---|---|---|---|---|
| PERIOD | FILES (MODIF.) | SENTENCES | TOKENS | TYPES (APPROX.) | TTR (APPROX.) |
| 1943–1962 | 15 | 1044 | 29822 | 3995 | 0.134 |
| 1964–1986 | 21 | 1847 | 58560 | 8211 | 0.1402 |
| 1988–2016 | 26 | 783 | 29332 | 6121 | 0.2087 |
| 2017–2022 | 5 | 1361 | 45360 | 3643 | 0.0803 |
[i] Note: TTR = Type-Token Ratio.
Table 4
Complexity Features Normalized Frequency per Period.
| COMPLEXITY FEATURES — MEAN PER 100 TOKENS | ||||
|---|---|---|---|---|
| FEATURE | 1943–1962 | 1964–1986 | 1988–2016 | 2017–2022 |
| Word Order Inversion (VS) | 0.268 | 0.360 | 0.361 | 0.373 |
| Nominalizations | 5.771 | 5.932 | 5.615 | 5.721 |
| Subordinate Clauses | 5.858 | 5.605 | 5.117 | 5.258 |
| Passive Constructions | 3.947 | 3.967 | 3.948 | 3.746 |
| Gerunds & Participles | 4.973 | 4.964 | 5.097 | 5.077 |
| Relative Pronouns & Conj. | 0.899 | 0.750 | 0.610 | 0.701 |
| Appositions & Parentheticals | 1.462 | 1.679 | 2.063 | 1.748 |
| Prepositional Phrases | 16.873 | 16.665 | 17.674 | 17.866 |
| Complex Verb Forms | 1.724 | 1.984 | 1.841 | 1.724 |
[i] Note: Values represent normalized frequency rates per 100 tokens, extracted from CoNLL-U dependency parsed output using UDPipe (Straka & Straková, 2017) with the Portuguese Bosque model (Rademaker et al., 2017).

Figure 2
Scree Plot of Principal Components for Syntactic Complexity Features.
Note. Eigenvalues (λ) for each principal component extracted from a year-level PCA of nine syntactic complexity features. The red dashed line indicates the Kaiser criterion threshold (λ = 1), above which components are considered meaningful (Kaiser, 1960; Jolliffe, 2002, pp. 2–6). Three components meet this criterion, collectively explaining the majority of variance in the data.
Table 5
Principal Component Loadings for the Syntactic Complexity Features.
| FEATURE | PC1 | PC2 |
|---|---|---|
| Word Order Inversion (VS) | –0.086 | –0.483 |
| Nominalizations | –0.402 | 0.213 |
| Subordinate Clauses | –0.418 | 0.004 |
| Passive Constructions | –0.345 | –0.354 |
| Gerunds & Participles | –0.409 | –0.202 |
| Relative Pronouns & Conj. | –0.284 | 0.413 |
| Appositions & Parentheticals | –0.269 | 0.176 |
| Prepositional Phrases | –0.413 | 0.265 |
| Complex Verb Forms | –0.215 | –0.533 |
[i] Note. PCA loadings for syntactic features across the first three principal components. Values represent standardized loadings; sign indicates direction and magnitude reflects contribution to each component. PC1 captures a broad shared variance across most features, while PC2 reflects more localized structural contrasts among subsets of constructions.
Table 6
PC1 Complexity Composite by Period — Best Subset: Nominalizations, Subordinate Clauses, Gerunds & Participles, and Prepositional Phrases.
| PERIOD | n YEARS | MEAN PC1 | SD | MIN | MAX |
|---|---|---|---|---|---|
| 1943–1962 | 15 | 0.067 | 1.223 | –2.140 | 2.291 |
| 1964–1986 | 21 | –0.088 | 1.522 | –2.802 | 2.879 |
| 1988–2016 | 26 | –0.251 | 2.230 | –6.534 | 2.806 |
| 2017–2022 | 5 | 1.476 | 0.922 | 0.317 | 2.790 |
[i] Note: Best subset comprises four features: Nominalizations, Subordinate Clauses, Gerunds & Participles, and Prepositional Phrases (α = 0.909).

Figure 3
Distribution of PC1 Scores Across Historical Periods in Brazilian Labor Law.
Note. Each point represents the period mean; thick bars indicate 95% confidence intervals (CI), calculated as mean ± 1.96 × SE, where SE = SD/√n; thin bars indicate the observed min–max range. The red dashed line marks PC1 = 0, representing the grand mean of the standardized scores. Higher PC1 values indicate greater overall syntactic complexity.
