False Positive Risk in AI Detection of L2 Academic Writing: A Case Study
Abstract
AI -text detectors are increasingly used in education and scholarly publishing. While these tools are meant to protect academic integrity, they may also create new inequities for multilingual writers. This exploratory case study investigates why legitimate second language (L2) academic prose can be misclassified as AI -generated. We start from the premise that L2 academic writing functions as a distinct register whose recurrent linguistic features create predictability, which is one of the main metrics to detect AI -generated text. Using a single -author corpus repeatedly flagged as machine--generated, we develop a detector -inspired simulation capturing five proxies of high predictability: low perplexity, low burstiness, high formulaic language, consistent formality, and low lexical diversity. We simulated two common detection models: a prevalence model (flagging frequent proxies) and a cluster--based model (flagging co -occurrence of features). Prevalence -based detection flagged 62% of the sample manuscripts, while cluster -based detection flagged 38% with three or more co -occurring proxies. Although dense clustering (four to five proxies) was uncommon (7.8%), our sample was repeatedly flagged by both logics. These findings do not establish equivalence with commercial detection systems, but they suggest that highly conventionalized L2 academic prose may face a heightened risk of false positive AI flagging. We term this phenomenon the “detection paradox”.
© 2026 Rania Za’rour, Abdel Rahman Mitib Altakhaineh, published by SAN University
This work is licensed under the Creative Commons Attribution-NonCommercial-NoDerivatives 4.0 License.