(1) Context and motivation
Within the broader MIGR-TWIT Corpus,1 whose aim is to document and study migration-related political tweets both in France and in the UK, the corpus presented here, namely FR-MIGR-TWIT 2.0, is an open-source resource designed to represent the micro-diachronic evolution of discourse on immigration in France on Twitter/X between 2011 and 2022. It captures a form of short-term diachrony whose dynamics are accelerated by the properties of computer-mediated communication (CMC), where discourse is persistent, searchable, and reproducible.
As immigration issues have recurrently occupied a salient place in the political agenda of parties situated on the right and far-right continuum in France, the first half of the 2010s was marked by two main contextual factors: technological, with the growing use of Twitter in political communication; and geopolitical, with the increased salience of migration in Europe around 2015.
From a discourse-analytic point of view, Twitter/X has occupied a particularly important position. As a microblogging platform, it combines message brevity with hypertextual affordances, enabling posts to circulate through hashtags and user mentions. The brevity and relative decontextualization of political tweets may also contribute to what Hart (2013) describes as the “imperfect and biased way our mechanisms of information processing work”; more radically, Ott (2017) argues that Twitter “demands simplicity, promotes impulsivity, [and] fosters incivility.” From the perspective of previous work in political discourse analysis, migration-related lexical forms are highly ideologically charged in French political discourse (Charaudeau, 2005: 76), and appear to be particularly salient in political positioning and contestation. This raises the theoretical interest of documenting migration-related tweets posted and reposted by a broad range of political actors on Twitter/X, and of investigating whether and how the migr- derived terms are differently used across two macro-orientations and over time, as digital data make it possible to document such variation.
(1.1) Corpus rationale
The rationale underlying FR-MIGR-TWIT is therefore twofold. On the one hand, the corpus is designed from a discourse-analytic perspective, in that it documents migration-related political discourse in a contemporary CMC environment marked by rapid circulation and platform-specific forms of textual organization (hashtagged word sequences, emojis, multimedia-linked textuality, relation between actors, etc.). (see Paveau, 2013; Longhi, 2013, 2016, among others). In this respect, the FR-MIGR-TWIT corpus contributes to updating the empirical basis of the field by integrating a new type of linguistic material drawn from CMC.
On the other hand, the annotated corpus structure makes it possible to examine how recurrent migr- lexical forms may display different frequency patterns, syntactic distributions, semantic role assignments, modification patterns, as well as their involvement in list and parallelism structures across political orientations and over time, thereby capturing both ideological differentiation and short-term diachronic evolution.
(1.2) Corpus design
The corpus contains 17,397 tweets posted or reposted by 39 Twitter/X accounts associated with French political actors and parties. It focuses on their (re)posts between 1 January 2011 and 30 June 2022 that contain at least one occurrence of a lexical form derived from the root migr-, including both neologisms such as immigrationniste (‘immigrationist’) and non-spaced word sequences containing migr- derivatives. These linguistic forms, yielding approximately 20,000 occurrences, constitute the core annotation units and are referred to as the migr-lexicon. Table 1 presents the 10 most frequent items in the migr-lexicon across the corpus.
Table 1
Frequency ranking of the 10 most frequent migr-lexicon items in the FR-MIGR-TWIT corpus.
| MIGR-LEXICON | ENGLISH TRANSLATION | NUMBER OF OCCURRENCES | FREQUENCY (%) |
|---|---|---|---|
| immigration | immigration | 7,954 | 39.81 |
| migrant | migrant | 6,642 | 33.25 |
| migratoire | migratory | 2,857 | 14.30 |
| immigré | immigrant (noun or adjective) | 676 | 3.38 |
| migration | migration | 608 | 3.04 |
| #pjlasileimmigration | #pjtheasylumimmigration | 201 | 1.01 |
| #loiasileimmigration | #lawasylumimmigration | 138 | .69 |
| #débatimmigration | #debateimmigration | 81 | .41 |
| immigrationniste | immigrationnist | 78 | .39 |
| anti-migrant | anti-migrant | 74 | .37 |
| TOTAL | 19,309 | 96.65 |
The selection of accounts was based on explicit criteria aimed at ensuring political representativeness, breadth across the political spectrum, and sufficient discursive productivity, both on the CMC platform itself through (re)posting activity and in terms of media visibility: (i) participation in presidential elections; (ii) representation at the European level, that is, in the European Parliament; (iii) sufficient discursive productivity on Twitter/X, both in general and with respect to (re)posts containing migr-lexicon (henceforth migr-tweets), during at least part of the period between 2011 and 2022; and (iv) affiliation with structured political parties. The first two criteria were intended in particular to ensure political representativeness and media visibility. Regarding the discursive productivity of migr-tweets, their proportion among all (re)tweets posted by a given account was taken into consideration. For the selected accounts, the total number of all (re)tweets was retrieved using the Twitter API v2 Full-archive counts endpoint (Jeon, 2023, 2025a, 2025b).
Based on the traditional French left–right political spectrum, as well as on the labels under which the affiliated parties of political actors are themselves positioned, two subcorpora were constructed, which together enable systematic comparison across political orientations:
FR-R-MIGR-TWIT groups together actors and parties situated on the center-right to far-right continuum (n = 16) and comprises 11,761 tweets.
FR-L-MIGR-TWIT groups together actors and parties situated on the left to far-left continuum (n = 23) and comprises 5,636 tweets.
Table 2 presents the Twitter/X accounts associated with the actors and parties included in FR-R-MIGR-TWIT, while Table 3 presents those included in FR-L-MIGR-TWIT. The representative type of each account and the proportion of migr-tweets are also provided. The FR-R/FR-L distinction should be understood as a deliberately coarse-grained analytical partition intended to capture broad tendencies across the French left–right political divide. The aim is to identify and quantify systematic tendencies between these two macro-orientations. The methodological implications and limitations of this partition are discussed further in section (4).
Table 2
Twitter/X accounts included in the FR-R-MIGR-TWIT subcorpus.
| POLITICAL ACTOR OR PARTY | TWITTER/X ACCOUNT | REPRESENTATIVE TYPE | PERCENTAGE OF (RE)POSTS CONTAINING MIGR-LEXICON |
|---|---|---|---|
| Christian Estrosi | @cestrosi | Individual | 3.27 |
| Emmanuel Macron | @EmmanuelMacron | Individual | .71 |
| Éric Ciotti | @Eciotti | Individual | .40 |
| Éric Zemmour | @ZemmourEric | Individual | 4.06 |
| Florian Philippot | @f_philippot | Individual | 2.04 |
| Jordan Bardella | @J_Bardella | Individual | 6.17 |
| Marine Le Pen | @MLP_officiel | Individual | 6.27 |
| Marion Maréchal | @MarionMarechal | Individual | 5.76 |
| Michel Barnier | @MichelBarnier | Individual | 1.68 |
| Nicolas Bay | @NicolasBay_ | Individual | 10.68 |
| Nicolas Dupont-Aignan | @dupontaignan | Individual | 2.73 |
| Philippe Meunier | @Meunier_Ph | Individual | 2.59 |
| Rassemblement National | @RNational_off | Organization | 8.13 |
| Valérie Boyer | @valerieboyer13 | Individual | 1.84 |
| Valérie Pécresse | @vpecresse | Individual | 1.12 |
| Xavier Bertrand | @xavierbertrand | Individual | .79 |
Table 3
Twitter/X accounts included in the FR-L-MIGR-TWIT subcorpus.
| POLITICAL ACTOR OR PARTY | TWITTER/X ACCOUNT | REPRESENTATIVE TYPE | PERCENTAGE OF (RE)POSTS CONTAINING MIGR-LEXICON |
|---|---|---|---|
| Adrien Quatennens | @Aquatennens | Individual | 2.25 |
| Alexis Corbière | @alexiscorbiere | Individual | 1.23 |
| Anne Hidalgo | @Anne_Hidalgo | Individual | 2.45 |
| Arnaud Montebourg | @montebourg | Individual | .14 |
| Benoît Hamon | @benoithamon | Individual | 1.85 |
| Christiane Taubira | @ChTaubira | Individual | .32 |
| Clémentine Autain | @Clem_Autain | Individual | 1.53 |
| Danièle Obono | @Deputee_Obono | Individual | 4.86 |
| Esther Benbassa | @EstherBenbassa | Individual | 5.81 |
| Europe Ecologie-Les Verts | @EELV | Organization | 5.05 |
| François Hollande | @fhollande | Individual | .41 |
| François Ruffin | @francois_ruffin | Individual | .33 |
| Gauche Républicaine et Socialiste | @gauche_rs | Organization | .52 |
| Génération.s | @generationsmvt | Organization | 5.95 |
| Jean-Luc Mélenchon | @jlmelenchon | Individual | .69 |
| La France Insoumise | @franceinsoumise | Organization | 2.13 |
| Manon Aubry | @manonaubryfr | Individual | 1.27 |
| Nathalie Arthaud | @n_arthaud | Organization | 5.78 |
| Parti Radical de Gauche | @partiradicalg | Organization | .38 |
| Parti Socialiste | @partisocialiste | Organization | .79 |
| Philippe Poutou | @philippepoutou | Individual | 1.43 |
| Raphael Glucksmann | @rglucks1 | Individual | 2.45 |
| Yannick Jadot | @yjadot | Individual | 2.28 |
(1.3) Annotation structure and theoretical grounding
The annotation structure of FR-MIGR-TWIT is built around the migr-lexicon, which serves as the central annotation unit in the corpus. The annotation layers were established to operationalize theoretically motivated dimensions of variation in migration-related political discourse. The main analytical dimensions of the annotation, conducted manually in ANALEC (see section (3.2)), are summarized in Table 4.
Table 4
Annotation dimensions, analytical purposes, and labels.
| ANNOTATION DIMENSION | ANALYTICAL PURPOSE |
|---|---|
| Lexical identification | Identifies the surface form of the annotated migr- occurrence |
| Lemmatization | Normalizes inflected or surface forms under a shared lexical entry |
| Syntactic function | Captures the syntactic function of the annotated occurrence. The inventory includes core syntactic labels such as adjunct (AD), dependent (DEP), direct object (OBJ), oblique object (OBL), predicate (PRED), root (ROOT), syntactic subject (SUB) and verb (VERB), as well as platform-specific labels such as tweet-final and tweet-initial hashtags or mentions (tw.final.hashtag; tw.final.mention; tw.initial.hashtag; tw.initial.mention) and non-hyperlinked tweet-initial positions (tw.initial.nonhypertext). |
| Semantic role | Captures the semantic role assigned to the annotated occurrence. The inventory includes traditional labels such as Agent, Content, Experiencer, Force, Goal, Instrument, Result, Source, and Theme, as well as adapted labels including Beneficiary (Theme.Beneficiary.neg) and Maleficiary (Theme.Beneficiary.pos). It also includes genre-specific labels corresponding to discursive and semiotic functions specific to political tweets, such as Content, Speaker, and discursive topic (Topic), and Place (including subtypes such as Place.Media and Place.PoliticalEvent). |
| Modification | Captures whether the annotated item modifies another noun or is itself modified, and records the lexical identity of the modifier or non-migr- head noun |
| List and parallelism structures | Captures whether the annotated item appears in list or parallelism structures, and records their extent through the number of conjuncts to the left and/or right of the migr- occurrence |
| Tweet metadata | Links each annotation to contextual tweet-level information (date, username, retweet origin) |
The model combines (i) syntactic annotation adapted from the Rhapsodie project (Lacheret et al., 2014; Lacheret et al., 2019), (ii) semantic role annotation based on a reconfigured inventory adapted from Jurafsky and Martin (2019), Tellier (2004), and Saeed (2016), and (iii) annotation of modification, as well as list and parallelism structures, which are relevant for analysing categorization and discursive framing in political discourse.
(1.3.1) Syntactic function
The syntactic function inventory developed for the annotation model is presented in Table 4. The typology developed in the Rhapsodie project was used as the starting point for two main reasons. First, a substantial number of political tweets included in the corpus originate from spoken contexts, such as political campaign meetings, television interviews, or radio broadcasts (Jeon, 2025a). Although these utterances become decontextualized through their circulation on Twitter/X, they also tend to be recontextualized within the CMC environment, insofar as a tweet itself may function as a discursive unit. Second, the Rhapsodie typology offers a terminology that foregrounds syntactic dependency relations rather than traditional labels such as “complement of noun”. For instance, in répartition des 141 migrants ‘distribution of 141 migrants’, the occurrence migrants occupies a syntactically dependent position rather than a salient agentive position. From a discourse-analytic perspective, such configurations are particularly relevant because migration-related entities are often represented not as discourse topics or agentive participants, but rather as socially passive entities associated with institutional or other agentive actors such as the EU or political opponents, as illustrated in example (tw4).
Thus, this layer makes it possible to examine how migr-lexicon occurrences are syntactically integrated into utterances, for instance as subjects, objects, oblique or dependent complements. Specific values were also added for tweet-specific configurations, such as tweet-initial or tweet-final hashtags and mentions. Examples (tw1)–(tw3) illustrate three frequent values in the corpus, namely sub, obj, and dep.
(tw1) [SUB] [Des migrants] bloquent le tunnel, attaquent les forces de l’ordre : une “richesse culturelle exceptionnelle” dixit la maire UMP de Calais?… (R_@f_philippot, 2015-10-03).
English translation: [Migrants] are blocking the tunnel and attacking the police: an “extraordinary cultural enrichment”, according to the UMP Mayer of Calais?…
(tw2) [OBJ] RT @david_rachline: “Pour assurer la sécurité des Français, il faut stopper [l’immigration], sortir de Schengen, retrouver nos frontières nationales.” @BFMTV (R_@RNational_off, 2017-10-16).
English translation: RT @david_rachline: «To ensure the safety of the French, we must stop [immigration], leave Schengen, and restore our national borders.” @BFMTV
(tw3) [DEP] Tout est dans le « à nouveau » : ces évacuations de campements de [migrants] sont sans fin car nous n’avons pas de frontières nationales. \n Situation exécrable pour tout le monde. #GrandeSynthe (R_@f_philippot, 2018-09-06).
English translation: Everything lies in the phrase “once again”: these evacuations of [migrant] camps are endless because we do not have national borders. An appalling situation for everyone. #GrandeSynthe
(1.3.2) Semantic role
Semantic role annotation is designed to capture how migr-lexicon referents are positioned with respect to represented events. The semantic role inventory contains 16 labels, as presented in Table 4, including traditional labels such as Agent, Theme, Force, and Experiencer, as well as discourse-semiotic functions such as discursive topic (Topic). In this model, Theme is derived from traditional semantic role literature (Tellier, 2004; Saeed, 2016; Jurafsky & Martin, 2019), whereas Topic is based on Van Dijk’s discourse-semantic perspective on macrostructures and topical organization (1998: 206). Within this inventory, analytical attention was paid to the contrast between Theme, Theme.Beneficiary.pos, and Theme.Beneficiary.neg. Here, Theme refers to an affected or undergoing participant without explicit evaluative orientation; Theme.Beneficiary.pos to a subtype of Theme construed as benefiting from, or as deserving, support, protection, or solidarity; and Theme.Beneficiary.neg to a subtype of Theme construed as a maleficiary of the event or as a participant in an event that is itself axiologically rejected. Examples (tw4)–(tw6) illustrate these three values. Van Leeuwen (1996) emphasized the “lack of biuniqueness of language” in relation to the representation of social actors, arguing that sociological categories such as ‘agents’ or ‘patients’ do not systematically correspond to linguistic categories. By combining syntactic function and semantic role annotation, the present model seeks to address this lack of biuniqueness. In this respect, the distinction between Theme, Theme.Beneficiary.pos, and Theme.Beneficiary.neg makes it possible to observe both de-agentivization and axiological polarization in the representation of migration-related entities through linguistic material.
(tw4) [Theme] Au lieu de se féliciter de la répartition [des 141 migrants] de l’#Aquarius dealé avec l’Allemagne, le Luxembourg, le Portugal et l’Espagne, @EmmanuelMacron ferait mieux d’impulser une politique migratoire européenne ferme, seule gage d’humanité. (R_@ECiotti, 2018-08-14)
English translation: Instead of congratulating himself on the distribution of [the 141 migrants] from the #Aquarius, bargained over with Germany, Luxembourg, Portugal, and Spain, @EmmanuelMacron would do better to promote a firm European migration policy, the only true guarantee of humanity.
(tw5) [Theme.Beneficiary.neg] RT @MLP_officiel: Tout mon soutien aux forces de l’ordre confrontées à la violence de 300 clandestins armés à #Calais. Il n’y a plus de place pour la naïveté : il faut expulser [ces migrants] illégaux qui n’ont rien de «réfugiés» et qui s’en prennent à nos policiers! MLP (R_@RNational_off, 2021-06-03)
English translation: RT @MLP_officiel: My full support goes to the police facing the violence of 300-armed clandestine people in #Calais. There is no longer any place for naivety: we must expel [these illegal migrants], who are nothing like “refugees” and who are attacking our police officers! MLP
(tw6) [Theme.Beneficiary.pos] @MichelMagniez: #HeninBeaumont – Une conseillère municipale #EELV aide [des migrants] à pouvoir manger : le maire #FN la qualifie de «honte de la commune». (L_@yjadot; 2016-08-21).
English translation: @MichelMagniez: #HeninBeaumont – An #EELV municipal councillor is helping [migrants] get something to eat: the #FN mayor qualifies her as “a disgrace to the town”.
Additional values were introduced for tweet-specific cases, especially when migr- derivatives occur as hashtags or mentions functioning as discourse topics. Thus, a tweet-final hashtag such as #Migrants may be annotated as Topic, as in example (tw7).
(tw7) [Topic] Je répondais aux questions de la chaîne américaine @CNN : https://t.co/OCnrlVWGBM #AttentatsParis [#Migrants] (R_@MLP_officiel, 2015-11-20).
English translation: I was answering questions from the American TV channel @CNN: https://t.co/OCnrlVWGBM #ParisAttacks [#Migrants]
(1.3.3) Modification
This layer records the form, lemma, and relative position of modifiers associated with migr-lexicon occurrences, as well as the lemma of the non-migr head noun when the migr- derivative itself functions as a modifier. It makes it possible to examine how migration-related entities are lexically framed. Examples (tw8) and (tw9) illustrate contrasting lexical framings of the same referential group, namely the 629 migrants aboard the Aquarius, represented respectively as ‘endangered migrants’ and ‘clandestine migrants’.
(tw8) RT @Gemenne: Le nombre d’arrivées en #Italie n’a jamais été aussi bas depuis 2013, divisé par 3 par rapport à 2017. Mais #Salvini a décidé de faire un coup politique sur le dos de 629 #migrants en danger. C’est ça, l’extrême-droite. #Aquarius (via @ispionline @emmevilla). (L_@yjadot, 2018-06-12)
English translation: The number of arrivals in #Italy has never been so low since 2013, down threefold compared with 2017. But #Salvini has decided to make a political move at the expense of 629 endangered #migrants. That is the far right. #Aquarius (via @ispionline @emmevilla).
(tw9) Une fois en #Espagne, les 629 migrants clandestins pourront se rendre dans n’importe quel pays européen. Sans contrôle de nos frontières, nous continuerons de subir le chaos migratoire. Finissons-en avec #Schengen! \n #Aquarius #chiudiamoiporti (R_@NicolasBay_, 2018-06-12)
English translation: Once in #Spain, the 629 clandestine migrants will be able to travel to any European country. Without control of our borders, we will continue to suffer migratory chaos. Let us put an end to #Schengen! #Aquarius #chiudiamoiporti
A recurrent configuration in the corpus is the relational adjective migratoire ‘migratory’ modifying a non-migr noun. In such cases, the annotation scheme records the relation between the migr- derivative and its head noun, making it possible to investigate how migration-related reference is construed within nominal phrase structure. This configuration is particularly relevant for examining ideological representations, in line with Van Dijk’s socio-cognitive approach to discourse (Van Dijk, 1998, 2006). Examples (tw10) and (tw11) illustrate this pattern with submersion migratoire ‘migratory submersion’ and présence migratoire ‘migratory presence’.
(tw10) [Force] “L’immigration est massive dans notre pays. La submersion [migratoire] que nous subissons n’est pas un fantasme.” #RTLMatin (R_@MLP_officiel, 2017-04-18)
English translation: «Immigration is massive in our country. The [migratory] submersion that we are experiencing is not a fantasy.” #RTLMatin
(tw11) [Agent] RT @GilbertCollard: Grande -Synthe (Nord): M. Le Pen accuse les élus “d’organiser la présence” [migratoire] : tout simplement le courage de dire la vérité! (R_@NicolasBay_, 2017-01-25)
English translation: RT @GilbertCollard: Grande -Synthe (Nord): M. Le Pen accuses elected officials of “organizing [migratory] presence”: simply the courage to tell the truth!
(1.3.4) List and parallelism structures
Particular attention is given to list and parallelism structures, which are encoded as formal configurations grouping lexical items into coordinated or paratactic sequences. In line with previous work on list constructions (Blanche-Benveniste, 1990; Kahane & Pietrandrea, 2012), these constructions make it possible to identify cases where migr-lexicon occurrences are inserted into ad hoc categories that suggest shared properties and contribute to contextual categorization at a discursive level. Pietrandrea and Battaglia (2022: 154), for instance, show that in French right-wing political tweets such configurations may group migration-related terms with other negatively evaluated social phenomena. Example (tw12) illustrates this pattern. The annotation scheme records whether the occurrence appears in a list or parallelism structure and specifies the extent of the construction through the number of conjuncts occurring to the left and/or right of the annotated item.
(tw12) #Chômage, #Insécurité, #Migrants #révisionconstitutionnelle, 2012–2016 : #Hollande Président «la chienlit» permanente. (R_@Meunier_Ph, 2016-01-26)
English translation: #Unemployment, #Insecurity, #Migrants #constitutionalrevision, 2012–2016: President #Hollande “the permanent disorder”.
(2) Dataset description
Repository location
https://doi.org/10.5281/zenodo.17828433
The dataset is also archived in Ortolang (https://www.ortolang.fr/market/corpora/fr-migr-twit-corpus-20).
Repository name
FR-MIGR-TWIT Corpus 2.0
Object name
FR-MIGR-TWIT_2.0.zip
Format names and versions
The corpus is distributed in two complementary versions. A fully annotated version, including the complete annotation layers and all associated metadata, is provided in CSV (UTF-8) and XML formats. These files are named FRMIGRTWIT_v2.csv and FRMIGRTWIT_v2.xml, respectively. A standalone corpus-text version containing essential tweet-level metadata, such as author username and identifier, tweet identifier, and tweet URL, is provided in TEI XML format within the tei/subfolder and is distributed as yearly files covering 2011–2022, following the filename pattern FRMIGRTWIT_YYYY_text.tei.xml, where YYYY denotes the year. The repository also includes a Python script, query_frmigrtwit.py, which takes FRMIGRTWIT_v2.csv as input and enables the generation of concordances from simple queries (Jeon & Pietrandrea, 2025).
Creation dates
The FR-R module was created between 2021-09-01 and 2022-11-01, including the completion of truncated retweets. The FR-L module was created between 2023-03-01 and 2023-07-31, likewise including the completion of truncated retweets. These two modules were subsequently compiled into FR-MIGR-TWIT 2.0 and underwent manual annotation until 2024-05-31, followed by an inter-annotator agreement phase completed on 2025-04-30.
Dataset creators
Sangwan Jeon, Paola Pietrandrea, Elena Battaglia.
Dataset annotator
Sangwan Jeon
Language
French-language data; variable names in English (e.g. data__text, data__created_at, data__author_id, etc.)
License
CC BY-NC-SA 4.0.
Publication date
2025-12-18.
(3) Method
(3.1) Data collection and corpus construction
Data collection was conducted in three successive steps. First, tweets containing at least one occurrence of migr- derivatives were collected from the 39 accounts selected according to the criteria described in subsection (1.2). This initial phase aimed to retrieve the textual content of relevant tweets together with their associated metadata. To this end, data were collected through the Full-archive search endpoint of the Twitter API v2 (Academic Research track). To facilitate collaborative querying and the reuse of request settings, API calls were executed through the Postman interface. Table 5 presents the lexical query terms used in the Full-archive search endpoint, and Table 6 summarizes the tweets.fields parameters retrieved for corpus construction and metadata archiving.
Table 5
Query terms used for migr-tweet extraction.
| KEYWORDS | ENGLISH TRANSLATION |
|---|---|
| immigration, immigrations | immigration, immigrations |
| migrant, migrants | migrant, migrants |
| immigré, immigrés, immigrée, immigrées, immigrant, immigrants, immigrante, immigrantes | immigrant |
| migr | migr- string (broad retrieval key) |
| migration, migrations | migration, migrations |
| migratoire, migratoires | migratory |
Table 6
Metadata settings for tweet extraction.
| TWEETS.FIELDS |
|---|
| author_id; context_annotations; conversation_id; created_at; entities; geo; id; in_reply_to_user_id; lang; possibly_sensitive; public_metrics; referenced_tweets; reply_settings; source; text; withheld |
The second phase relied on the Full-archive Tweet counts endpoint and aimed to record the overall volume of tweets and retweets published by the selected accounts. This made it possible to operationalize the discursive productivity of (re)tweets containing migr- derivatives by measuring their proportion relative to the total posting activity of each account. The third phase consisted in manually completing truncated retweets by tracking the URL of the original post retweeted. Owing to technical limitations in the retrieved retweet text, some retweets exceeding 140 characters were returned in truncated form, restricted to the first 140 characters. These retweets were therefore manually completed in order to restore the full textual content of the post.
(3.2) Annotation model and implementation
Annotation was performed manually using ANALEC, a linguistic annotation and analysis tool developed at the CNRS research unit Lattice (Landragin, Poibeau & Victorri, 2012), which made it possible to implement a fine-grained multi-layer annotation scheme. The analytical dimensions introduced in subsection (1.3) are implemented in the annotation scheme through the field names summarized in Table 7. Figure 1 displays the annotation structures implemented in ANALEC, including the drop-down menus used to assign syntactic function and semantic role labels.
Table 7
Annotation dimensions implemented in ANALEC: field names of the annotation unit MIGR-LEXICON.
| ANALYTICAL DIMENSION | ANNOTATION FIELD NAME(S) ASSOCIATED WITH THE ANNOTATION UNIT MIGR-LEXICON IN ANALEC |
|---|---|
| Lexical identification | #forme# |
| Lemmatization | lemma |
| Syntactic function | func_syn |
| Semantic role | role_sem |
| Modification | modification, lemma_modif_l1…/lemma_noun-1 |
| List and parallelism structures | list_par, rel-list_par, length-1, #forme#_migr-list_par |
| Tweet metadata | rt, username, date |

Figure 1
Annotation structures in ANALEC.
In this model, the central annotation unit is the migr-lexicon occurrence. Each annotated occurrence is annotated through a set of fields corresponding to these dimensions, including lexical identification, lemmatization, syntactic function, semantic role, modification, list and parallelism structures, and tweet-level metadata (i.e., rt, username, date). In addition to the intrinsic properties of the migr-lexicon occurrence, the annotation also includes information linked to structurally related units, such as list or parallelism structures (list_parallelism), conjuncts (conjunct), modifier lemmas (lemma_modif), or the lemma of a modified noun (lemma_noun-1), as well as tweet-level metadata manually integrated into the annotation files. As an illustration, Figure 2 presents the annotated occurrence migr-lexicon-9075 in ANALEC interface, and Figure 3 illustrates the corresponding XML encoding.

Figure 2
Annotation unit MIGR-LEXICON-9075 in the ANALEC interface.

Figure 3
XML encoding of annotation unit MIGR-LEXICON-9075.
(3.3) Annotation campaign and inter-annotator agreement
The annotation campaign was carried out by trained annotators (final-year undergraduate students in linguistics at Université de Lille) following a shared protocol, a training phase, and successive calibration steps. Inter-annotator agreement was evaluated on a subset of the corpus against a reference annotation, and Table 8 reports the comparison with annotator A for illustration. This reference-based approach was preferred to pairwise comparison between annotators because the aim was to assess the extent to which trainee annotations converged with an expert annotation already established and used in the research underlying the corpus (Jeon, 2025a). To ensure feasibility, the annotators each worked on a randomly selected sample of tweets, preceded by a training set. Agreement was calculated by aligning annotated migr-lexicon occurrences with the reference annotation and comparing the annotated properties using exact agreement and Cohen’s kappa. Overall, agreement reached satisfactory levels for lexical and morphosyntactic layers, while semantic role annotation yielded lower but still interpretable scores. Disagreements were analyzed and used to refine the annotation guidelines.
Table 8
Inter-annotator agreement results: exact agreement and Cohen’s kappa.
| EXACT AGREEMENT | KAPPA | |
|---|---|---|
| LEMMA | 1.0 | 1.0 |
| #forme# | .932 | .927 |
| ROLE_SEM | .729 | .606 |
| FUNC_SYN | .822 | .671 |
| LIST/PAR | .904 | .788 |
| modification | .957 | .925 |
| LEMMA_MODIF_L1 | .965 | .656 |
| LEMMA_MODIF_R1 | .951 | .846 |
| LEMMA_MODIF_R2 | .986 | .716 |
| LEMMA_MODIF-R3 | 1.0 | |
| LEMMA_MODIF-R4 | 1.0 | |
| LEMMA_MODIF-R5 | 1.0 | |
| LEMMA_NOUN-1 | .938 | .819 |
(4) Methodological challenges
(4.1) Subcorpus design and keyword selection
The corpus design distinguishes two political subcorpora, FR-R and FR-L, in order to enable comparison between contrasting communities of practice. One methodological challenge, however, lies in the quantitative imbalance between these two subcorpora, with right-wing and far-right actors contributing to a substantially larger number of migration-related tweets and migr-lexicon occurrences than left-wing and far-left actors. This imbalance partly reflects a corpus design constraint, but it also corresponds to a meaningful empirical property of the dataset: migration is not equally salient across political communities, and the right-wing continuum shows a stronger and more sustained investment in migration-related references over the period considered. In this sense, the asymmetry should be interpreted as a characteristic of the political discourse under study. These differences are also supported by statistical testing conducted on the distribution of migr-lexicon occurrences, as well as on the proportion of tweets containing migr-lexicon occurrences (Jeon, 2025a, 2025b).
A further point which was raised during the review process was that the disparity between left- and right-wing discourse might stem from differences in lexical choice, and it was suggested that the feasibility of including migration-related terms not derived from the migr-lexicon, such as demandeur d’asile ‘asylum seekers’ or réfugiés ‘refugees’, be examined. Indeed, these terms were tested during the collection phase, during which preliminary observations suggest that this disparity cannot be explained solely by the use of alternative migration-related lexical items, notably because, across the 39 selected accounts, tweets containing these terms without any co-occurring migr- derivative remained rare in the French context. In other words, non-migr- migration-related terms frequently co-occurred with migr-lexicon occurrences rather than functioning as exclusive alternatives to them. These observations contributed to centering the corpus on the migr-lexicon, at least in the French context. Interestingly, this pattern differs from observations made in the UK context, where non-migr- migration-related terminology appears more autonomous.
Another remark concerned the representativeness of the selected political accounts. In particular, some major French center-right political parties are not represented through their official accounts, notably @lesRepublicains and @Renaissance, representing respectively Les Républicains (formerly UMP) and Renaissance (inherited from En Marche).
It should be noted that the construction of FR-R corpus, initially developed as a pilot corpus of political migr-tweets, followed an iterative process. During the initial account selection stage, migration-related discourse was found to be particularly concentrated among individual accounts associated with, or originating from, the Rassemblement National (formerly Front National), including @J_Bardella, @MLP_officiel, @MarionMarechal, @NicolasBay_, and @f_philippot. Frequent retweeting practices between these accounts and the organizational account @RNational_off were also observed, which lead to the inclusion of the latter in the corpus. This decision is further supported by the proportion rates reported in the subsequent analysis:2 @RNational_off displayed a substantially higher proportion of migr-tweets (8.13%) (see Table 2) than @lesRepublicains (.66%; 505/76,800) and @Renaissance (.20%; 44/22,200).
Nevertheless, several accounts presenting lower proportion rates were also retained in order to capture variation across political affiliations. Given the macro-comparative scope of the present study, the absence of these accounts is unlikely to substantially affect the overall analytical tendencies observed in the corpus. Future research could nevertheless integrate migr-tweets from @lesRepublicains and @Renaissance, particularly in the context of account-level quantitative and qualitative analyses.
(4.2) Political categorization
Another methodological challenge concerns the political positioning of certain political actors whose placement on the left–right axis is not always stable or self-evident. This is particularly the case for actors such as Emmanuel Macron, who has long claimed to stand outside the traditional left–right divide. It should therefore be emphasized that, in the present study, the FR-R/FR-L distinction is not intended as a definitive ideological classification of political actors. Rather, it functions as a deliberately coarse-grained analytical partition designed to capture broad tendencies over time in the discursive treatment of migration-related entities and issues across the French political field. The annotation design nonetheless makes it possible to conduct finer-grained analyses at the account level, which remain beyond the scope of the study but constitute an important direction for future research.
(4.3) Inclusion of retweets
Another methodological decision concerns the inclusion of retweets in the corpus. From a purely text-based perspective, retweets may appear as repeated material. However, in the context of CMC, retweeting constitutes a platform-specific socio-technical discursive practice, embedded in the communicative ecology of Twitter/X (Crystal, 2011: 40–41; Paveau, 2013; Longhi, 2016: 112–113). From the perspective of political discourse analysis, retweets are therefore treated as observable traces of recirculation, reinforcement, and public uptake. They provide evidence of lexical recurrence, ideological alignment, and selective endorsement. Their inclusion is thus methodologically justified insofar as the present study is concerned not only with linguistic forms within (re)tweets as discourse units in their own right, but also with their circulation as decontextualized discourse segments, as well as with their persistence, visibility, and repeated activation in digital political communication.
(5) Results and discussion
The analysis presented here focuses on the distributional behavior of the most frequent lexemes in the corpus, namely immigration ‘immigration’ and migrant ‘migrant’. The annotated data point to two main findings: first, a progressive referential reconfiguration of the migr-lexicon, observable through semantic roles and syntactic positioning; second, a contrastive use of list and parallelism structures as sites of diachronic ideological categorization across the left- and right-wing subcorpora (FR-L and FR-R).
(5.1) Referential reconfiguration and syntactic positioning
A first result concerns the semantic and syntactic reconfiguration of migrant and immigration over time. For immigration, the right-wing subcorpus shows a significant restructuring of semantic-role distributions (χ² (6) ≈ 43.13; p < .001) (Jeon & Pietrandrea, in press): although Topic remains dominant overall, it declines after 2015, while more dynamic and negatively oriented roles such as Force and Theme.Beneficiary.neg become more prominent. Rather than functioning primarily as a discursive topic label, immigration increasingly appears in configurations that construe it as a source of pressure, disturbance, or harmful consequences, as illustrated in (tw13):
(tw13) [Force] RT @MLP_officiel: «Notre pays est confronté à une explosion de [l’immigration], de la délinquance et l’on voit s’imposer partout un climat de grande violence criminelle, terroriste ou sociale. La délinquance silencieuse mais épouvantable à vivre gagne nos campagnes!» #HauteMarne #OnArrive (R_@RNational_off, 2019-02-23).
English translation: Our country is facing an explosion of [immigration] and delinquency, and we can see a climate of serious criminal, terrorist, and social violence imposing itself everywhere. A silent form of delinquency, yet dreadful to live with, is spreading into our rural areas.
In the left-wing subcorpus, distributional change is also significant (χ²(6) ≈ 24.7; p < .01), but much weaker: Topic remains predominant throughout the period, and the other roles, including Agent, Force, and Theme.Beneficiary.pos, remain rare (<5%) and do not alter the overall configuration. This suggests that immigration remains predominantly framed as a topic of political debate rather than as an axiologically polarized entity.
For migrant, the right-wing subcorpus shows a shift away from relatively descriptive uses, associated with Theme attribution, toward more impersonal, causative, and axiologically charged construals, associated with Force attribution, as illustrated in (tw14) and (tw15).
(tw14) L’UE va donc regarder à la jumelle les bateaux de migrants…Pour l’action, ce sera sans elle ni l’RPS, sponsors historiques de l’immigration (R_@f_philippot, 2015-06-22)
English translation: The EU will therefore watch the boats of migrants through binoculars…As for action, it will be without the EU or the RPS, historical sponsors of immigration.
(tw15) “Le port de #Calais est entrain de s’effondrer, notamment à cause de l’afflux de #migrants.” @F3nord @F3Picardie; (R_@MLP_officiel, 2015-12-02)
English translation: The port of #Calais is collapsing, notably because of the influx of #migrants.
In both of such cases, the lemma, occurring in the plural form, tends to appear in syntactically dependent positions, as in les bateaux de migrants ‘the boats carrying migrants’ (literally, ‘the boats of migrants’) or l’afflux de #migrants ‘the influx of #migrants’. Such configurations reduce the salience of its intrinsic [+human] features. This recurrent distribution contributes to a redistribution of agency and supports the representation of migration-related entities as affected, framed, or backgrounded rather than volitionally acting participants or social actors. In line with an anonymous reviewer’s suggestion, expressions such as explosion de l’immigration ‘explosion of immigration’, afflux de migrants ‘influx of migrants’, or submersion migratoire ‘migratory submersion’ reflect broader metaphorical framings of migration-related entities in anti-immigration discourse (Hart, 2013; Taylor, 2021), in which immigration and migrants are construed as overwhelming natural forces requiring containment.
In the left-wing corpus, by contrast, migrant remains predominantly associated with vulnerable human referents (Theme.beneficiary.pos, as illustrated earlier in example (tw6)). These patterns show that the combined annotation of semantic role, modification, and syntactic function makes it possible to quantify the referential reconfiguration of the migr-lexicon across political orientation (Jeon, 2025a; Jeon & Pietrandrea, in press).
(5.2) List structures as sites of diachronic ideological categorization
A second result concerns list and parallelism structures, which provide a privileged site for observing how migration-related lexemes are aligned with other semantic domains and integrated into broader ideological frames. As noted earlier (cf. (1.3.4)), these configurations correspond to coordinated or paratactic configurations in which several lexical items are aligned within the same syntactic level of clustering. Following previous work on list constructions (Blanche-Benveniste, 1990, 2011; Rabatel, 2011; Kahane & Pietrandrea, 2012; Masini et al., 2018), they are analyzed here as discursive devices that temporarily group heterogeneous items into a common interpretive frame, resulting in construction of ad hoc category, that is, context-dependent groupings whose coherence is not based on stable lexical taxonomy, but on pragmatic and discursive inference.
This functional dimension is especially relevant in political tweets, where the juxtaposition of migration-related lexemes with other terms may suggest shared properties, shared evaluative orientations, or common ideological relevance. In other words, the list does not merely enumerate, but contributes to contextual categorization by inviting the reader to infer a relation of association between otherwise heterogeneous referents. To compare these environments quantitatively, conjunct lexemes were assigned to a set of thematic categories designed to capture recurrent discursive domains in which migrant and immigration are embedded. Table 9 summarizes the category codes and provides lexical indicators for each category. Their proportional distributions across the two subcorpora are shown in Figure 4 and Figure 5.
Table 9
Classification of lexical items associated with migr-lexicon items in lists and parallelisms.
| CATEGORY CODE | LEMMA_CONJUNCT | ENGLISH TRANSLATION |
|---|---|---|
| Security_judicial | insécurité, terrorisme, terroriste, islamisme, criminel, criminalité, délinquant, délinquance, clandestin, trafic, trafiquant, justice, police, violence | insecurity, terrorism, terrorist, Islamism, criminal, criminality, delinquent, delinquency, clandestine migrant(s), trafficking, trafficker, justice, police, violence |
| Economic_institutional | csg, retraite, impôt, impot, fiscalité, fraude, fraude sociale, pouvoir d’achat, concurrence, concurrence déloyale, traité de libre-échange, schengen, union européenne, europe, bruxelles, état, institution | CSG (general social contribution), retirement/pension, tax, taxation, fraud, welfare fraud/social fraud, purchasing power, competition, unfair competition, free trade agreement, Schengen, European Union, Europe, Brussels, state, institution |
| Identity | communautarisme, laïcité, identité, priorité nationale, nation, culture | secularism, identity, national priority, nation, culture |
| Humanitarian_asylum | réfugié, asile, demandeur d’asile, exilé, intégration, accueil, solidarité | refugee, asylum, asylum seeker, exile/exiled person, integration, reception/welcome, solidarity |
| Social_vulnerability | femme, lgbtqi, pauvre, sans-abri, rom, enfant | woman, LGBTQI, poor person, homeless person, Roma person, child |
| Geography | calais, ceuta, aquarius, paris, nice, italie, france | Calais, Ceuta, Aquarius, Paris, Nice, Italy, France |

Figure 4
Annual evolution of the thematic distribution of conjuncts associated with immigration in the FR-L and FR-R subcorpora.

Figure 5
Annual evolution of the thematic distribution of conjuncts associated with migrant in the FR-L and FR-R subcorpora.
In the right-wing subcorpus, the conjunct patterns associated with migrant show a clear short-term diachronic evolution. During the 2015–2016 period, corresponding to the so-called “migration crisis”, its most frequent conjuncts belong mainly to the security/judicial and geography domains. These environments frequently include emblematic place names such as Calais, Ceuta, and Aquarius, which foreground border pressure and crisis management, as illustrated in (tw16):
(tw16) N’attendons pas un nouveau drame! La loi doit s’appliquer, les clandestins doivent être reconduits. {#Calais | #Migrants | #Eurotunnel}; (R_@ECiotti, 2015-07-29)
English translation: Let us not wait for another tragedy! The law must be enforced, and clandestine migrants must be sent back. {#Calais | #Migrants | #Eurotunnel}.
From 2018 onward, and more clearly in 2019, the lexeme becomes increasingly embedded in list constructions involving economic and institutional terms, indicating a progressive extension from migration-specific reference to broader political and socio-economic framing. This shift is consistent with the wider political context, including debates surrounding the 2018 Marakech Pact and the 2019 European elections. In some later examples, list structures also operate through contrastive recategorization, as when réfugiés ‘refugees’ are explicitly opposed to migrants économiques ‘economic migrants’, as illustrated in (tw17):
(tw17) A #Melilla, il y a des centaines de jeunes hommes qui veulent passer la frontière de force : nous savons faire la différence entre {les réfugiés de guerre ukrainiens – avec qui il est normal que la France soit solidaire – | et ces #migrants économiques}. #Elysée2022 (@MLP_officiel, 2022-03-03)
English translation: In #Melilla, there are hundreds of young men trying to cross the border by force: we know the difference between {Ukrainian war refugees – with whom it is normal for France to show solidarity – | and these economic migrants}.
By contrast, the left-wing subcorpus displays a more stable thematic profile, especially for migrant, which remains predominantly embedded in list environments linked to humanitarian/asylum and social vulnerability referents. The lexeme immigration also shows a left/right contrast, although in a more gradual and heterogeneous way: in FR-R, its list environments progressively shift toward security/judicial and economic/institutional domains, whereas in FR-L they remain more strongly associated with humanitarian/asylum framing. When security/judicial conjuncts appear with immigration in the left-wing subcorpus, they often occur in explicitly contrastive or metadiscursive formulations that refer to the ideological positioning of political opponents, as illustrated in (tw18):
(tw18) Pour @mlp_officiel, “il y a un lien entre {l’immigration | et le terrorisme}”. Peuple métissé, nous vous méprisons, madame #f_inter #NoPasaran; (L_@JLMelenchon, 2012-04-19)
English translation: According to @mlp_officiel, “there is a link between {immigration | and terrorism}.” Mixed-race people, we despise you, madam. #f_inter #NoPasaran.
Taken together, these results suggest that list constructions do not merely reflect co-occurrence patterns, but participate in the short-term diachronic restructuring of ideological framing by stabilizing different ad hoc categories across political orientations.
Systematic differences thus emerge between the left- and right-wing subcorpora in lexical choices, syntactic constructions, and semantic-role distributions. These findings confirm that a fine-grained, multi-layer annotation scheme can support the study of ideological variation in political discourse on CMC platforms such as Twitter/X by linking lexical forms to linguistic constructions, discourse-level categorization, and micro-diachronic variation.
(6) Implications/Applications
(6.1) Diachronic relevance
The diachronic structure of the corpus, spanning 2011–2022, makes it possible to investigate temporal dynamics on migration-related discourse over a time span that also captures the growing use of Twitter as a privileged tool of political communication, followed by its relative decline from 2019 onwards. The dataset supports the analysis of changes in the frequency, distribution, and combinatory patterns of migr-lexicon occurrences across time, including peaks associated with major political events. It therefore enables the study of gradual change, periods of acceleration, and stabilization phenomena in political language.
Beyond individual findings, the corpus provides a basis for modeling language change in digital environments, where discourse is persistent, searchable, and recirculated. This persistence differentiates social media corpora from traditional textual corpora and has implications for the temporality of recursivity of linguistic and discursive change.
(6.2) Reusability and scope
FR-MIGR-TWIT 2.0 is structured to support reuse in multiple research contexts, including:
– micro-diachronic lexical and syntactic studies,
– comparative political discourse analysis (including account-level comparisons),
– computational modeling of language change and discursive variation,
– studies of categorization, reference, and ideological framing in discourse
As detailed in section (2), FR-MIGR-TWIT 2.0 is entirely encoded in UTF-8 and available in CSV and XML formats. The full annotation layers and metadata were not distributed entirely in TEI/XML format. Nevertheless, the TEI/XML version of FR-MIGR-TWIT incorporates TEI-compatible structural principles, including unique textual identifiers (xml:id), user attribution (who), temporal metadata (when), and utterance-like segmentation (<u>, <seg>). In this respect, the corpus may be loaded into software environments such as TXM and IRaMuTeQ. The TEI/XML structure and metadata organization may facilitate interoperability with other French CMC corpora documenting political communication, such as Polititweet (CoMeRe repository) (Chanier et al., 2014).
More broadly, FR-MIGR-TWIT 2.0 is part of the multilingual MIGR-TWIT corpus family, which includes a UK module of migration-related political tweets, namely UK-R-MIGR-RA-2012-2022. The MIGR-TWIT corpus family was developed within the framework of the OLiNDiNUM Linguistic Observatory of Digital Discourse corpus infrastructure (Pietrandrea et al., 2026), alongside resources such as UK-EU-DEBATE-20-21, thereby creating opportunities for comparative analyses of the right-wing political discourse on migration and online public debate across languages and political contexts in Europe.
The combination of rich metadata, multi-layer annotation, and explicit corpus-construction criteria ensures transparency, reproducibility, and methodological transferability. As such, the corpus may be reused not only for qualitative and quantitative linguistic analysis, but also as a resource for annotation experiments.
Notes
[1] The MIGR-TWIT Corpus is composed of three subcorpora, including FR-R-MIGR-TWIT-2011-2022 and UK-R-MIGR-RA-TWIT-2012-2022, created by Battaglia, Blandino, Jeon and Pietrandrea (2022) (Zenodo repository of the French and UK right-wing migration tweet corpora), as well as FR-L-MIGR-TWIT-2011-2022 (Pietrandrea & Jeon, 2023) (Zenodo repository of the French left-wing migration tweet corpus). The French modules were manually annotated according to the multi-layer linguistic annotation model presented in this article. The UK module focuses on right-wing political tweets including migr-derived terms, as well as refugee(s) and asylum. Rather than being annotated using the same annotation scheme, it was investigated through a combination of topic modeling and corpus linguistic methods (Blandino, 2023).
[2] A further frequency analysis of migr-tweets was subsequently conducted for the accounts @lesRepublicains and @Renaissance. Due to methodological limitations related to access to the Twitter API v2 Academic Research track after June 2023, the analysis covered the entire activity period of the accounts, from their creation date until 27 May 2026. Migr-tweets were counted manually using Twitter/X Advanced Search and divided by the total number of posts displayed on each account’s main page on 28 May 2026.
AI Declaration
Generative AI tools were used in a limited capacity during manuscript preparation to assist with editorial revision and, where appropriate, language rephrasing. They were also used as a support tool in the development and debugging of Python and R scripts used in the research workflow, including for format conversion, data cleaning, statistical testing, data visualization, and the management of large datasets. All AI-generated suggestions were reviewed, verified, and adapted by the authors. The design of the study, data collection, corpus annotation, analysis, and interpretation of results were carried out by the authors.
Acknowledgements
The authors would like to thank Elena Battaglia for contributing to the extraction and completion of the 2011–2016 subset of the FR-R module during an earlier phase of corpus creation. The authors also thank Lelia Pasquetti, Clara Defrenne and Charlotte Malherbe for their participation in the annotation campaign.
Author Contributions
Sangwan Jeon: Conceptualization, Methodology, Data Curation (including corpus compilation and annotation), Validation, Visualization, Writing – Original Draft, Writing – Review & Editing.
Paola Pietrandrea: Conceptualization, Methodology, Project Administration, Funding Acquisition, Supervision, Validation, Writing – Original Draft, Writing – Review & Editing.
