Abstract
The FR-MIGR-TWIT Corpus 2.0 is a French-language module within the broader MIGR-TWIT Corpus, a multilingual micro-diachronic dataset designed to study political discourse on migration in Europe. It constitutes an openly available resource for quantitative diachronic linguistics, covering a continuous period from 1 January 2011 to 30 June 2022 and enabling the analysis of linguistic and discursive change in digital political communication. Structured into two subcorpora, FR-R and FR-L, corresponding respectively to French political actors and organizations associated with the right-wing and left-wing continua, the corpus comprises 17,397 tweets and retweets containing at least one occurrence of a form derived from the Latin root migr-. It therefore makes it possible to study lexical productivity, distributional tendencies, and patterns of discursive framing and categorization over time. A major contribution of version 2.0 lies in its multi-layer annotation scheme. Annotated properties include surface form and lemma, syntactic function, semantic role, modification constructions, and list/parallelism structures. This annotation model is designed to make observable how migration-related entities are linguistically categorized, positioned, and evaluated in political discourse over time on Twitter/X. The corpus supports both qualitative and quantitative analyses of ideological variation, diachronic change, and categorization processes in social media discourse, and is intended as a reusable resource for corpus linguistics, discourse studies, and future annotation-based or computational research.
© 2026 Sangwan Jeon, Paola Pietrandrea, published by Ubiquity Press
This work is licensed under the Creative Commons Attribution 4.0 License.
