Skip to main content
Have a personal or library account? Click to login
A Pilot Corpus Study on Microdiachronic Complexity Features in Brazilian Labor Law Cover

A Pilot Corpus Study on Microdiachronic Complexity Features in Brazilian Labor Law

Open Access
|Jul 2026

Figures & Tables

Table 1

The LEX-BR-Ius Corpus.

SUB-CORPUSAVAILABLE TEXTSSELECTED TEXTSWORDSWORDS (SMALLEST TEXT)WORDS (LONGEST TEXT)AVERAGE
Códigos1713711,95212,225167,41654,765.50
Constituição1197,082N/AN/AN/A
Emendas à Constituição1144850,2414312,0151,046.69
Estatutos1811125,3562,37035,24111,396.00
Leis Complementares19243182,7689035,8374,250.42
Leis ordinárias13,4916382,129,5943786,6113,337.92
TOTAL13,8337543,296,99337167,4163,372.67

[i] Note: Adapted from Marques et al. (2025) and Ferrari et al. (2026). N/A indicates not applicable because only one text was available in that sub-corpus. Averages are calculated per selected text.

Table 2

Number of Constitutional Amendments across legislative periods.

PERIODNUMBER OF AMENDMENTS
1943–19625
1964–198655
1988–201630
2017–202248

[i] Note: Source: the authors. Periods correspond to distinct political/legislative eras as defined in the original analysis.

Figure 1

Sample of annotation of a sentence from the subtexts 1943.

Table 3

Period Statistics per Period.

PERIOD STATISTICS: FILES, TOKENS, TYPES, TTR
PERIODFILES (MODIF.)SENTENCESTOKENSTYPES (APPROX.)TTR (APPROX.)
1943–19621510442982239950.134
1964–19862118475856082110.1402
1988–2016267832933261210.2087
2017–2022513614536036430.0803

[i] Note: TTR = Type-Token Ratio.

Table 4

Complexity Features Normalized Frequency per Period.

COMPLEXITY FEATURES — MEAN PER 100 TOKENS
FEATURE1943–19621964–19861988–20162017–2022
Word Order Inversion (VS)0.2680.3600.3610.373
Nominalizations5.7715.9325.6155.721
Subordinate Clauses5.8585.6055.1175.258
Passive Constructions3.9473.9673.9483.746
Gerunds & Participles4.9734.9645.0975.077
Relative Pronouns & Conj.0.8990.7500.6100.701
Appositions & Parentheticals1.4621.6792.0631.748
Prepositional Phrases16.87316.66517.67417.866
Complex Verb Forms1.7241.9841.8411.724

[i] Note: Values represent normalized frequency rates per 100 tokens, extracted from CoNLL-U dependency parsed output using UDPipe (Straka & Straková, 2017) with the Portuguese Bosque model (Rademaker et al., 2017).

Figure 2

Scree Plot of Principal Components for Syntactic Complexity Features.

Note. Eigenvalues (λ) for each principal component extracted from a year-level PCA of nine syntactic complexity features. The red dashed line indicates the Kaiser criterion threshold (λ = 1), above which components are considered meaningful (Kaiser, 1960; Jolliffe, 2002, pp. 2–6). Three components meet this criterion, collectively explaining the majority of variance in the data.

Table 5

Principal Component Loadings for the Syntactic Complexity Features.

FEATUREPC1PC2
Word Order Inversion (VS)–0.086–0.483
Nominalizations–0.4020.213
Subordinate Clauses–0.4180.004
Passive Constructions–0.345–0.354
Gerunds & Participles–0.409–0.202
Relative Pronouns & Conj.–0.2840.413
Appositions & Parentheticals–0.2690.176
Prepositional Phrases–0.4130.265
Complex Verb Forms–0.215–0.533

[i] Note. PCA loadings for syntactic features across the first three principal components. Values represent standardized loadings; sign indicates direction and magnitude reflects contribution to each component. PC1 captures a broad shared variance across most features, while PC2 reflects more localized structural contrasts among subsets of constructions.

Table 6

PC1 Complexity Composite by Period — Best Subset: Nominalizations, Subordinate Clauses, Gerunds & Participles, and Prepositional Phrases.

PERIODn YEARSMEAN PC1SDMINMAX
1943–1962150.0671.223–2.1402.291
1964–198621–0.0881.522–2.8022.879
1988–201626–0.2512.230–6.5342.806
2017–202251.4760.9220.3172.790

[i] Note: Best subset comprises four features: Nominalizations, Subordinate Clauses, Gerunds & Participles, and Prepositional Phrases (α = 0.909).

Figure 3

Distribution of PC1 Scores Across Historical Periods in Brazilian Labor Law.

Note. Each point represents the period mean; thick bars indicate 95% confidence intervals (CI), calculated as mean ± 1.96 × SE, where SE = SD/√n; thin bars indicate the observed min–max range. The red dashed line marks PC1 = 0, representing the grand mean of the standardized scores. Higher PC1 values indicate greater overall syntactic complexity.

DOI: https://doi.org/10.5334/johd.578 | Journal eISSN: 2059-481X
Language: English
Page range: 91 - 91
Submitted on: Apr 25, 2026
Accepted on: Jun 4, 2026
Published on: Jul 3, 2026
Published by: Ubiquity Press
In partnership with: Paradigm Publishing Services

© 2026 Lúcia de Almeida Ferrari, Luciana Dias de Macedo, published by Ubiquity Press
This work is licensed under the Creative Commons Attribution 4.0 License.