Skip to main content
Have a personal or library account? Click to login
False Positive Risk in AI Detection of L2 Academic Writing: A Case Study Cover

False Positive Risk in AI Detection of L2 Academic Writing: A Case Study

Open Access
|Jul 2026

References

  1. Abd‑Elaal, E.‑S., Gamage, S. J. E., & Mills, J. E. (2022). Assisting academics to identify computer -generated writing. European Journal of Engineering Education, 47(5), 725—745. DOI: 10.1080/03043797.2022.2046709.
  2. Ädel, A., & Erman, B. (2012). Recurrent word combinations in academic writing by native and non -native speakers of English: A lexical bundles approach. English for Specific Purposes, 31(2), 81—92. DOI: 10.1016/j.esp.2011.08.004.
  3. Ajideh, P., Zohrabi, M., & Ohbatalab, R. (2025). Stance markers in academic writing: Native vs. non -native (Iranian) authorship in hard and soft sciences research articles. Critical Literary Studies, 7(2). DOI: 10.22034/cls.2025.63772.
  4. Ahmed, J., & Ahmed, N. (2026). When Human Writing is Marked AI written: Investigating Turnitin’s Misclassification of Students’ Writing. Research Journal for Social Affairs, 4(1), 209—215. DOI: 10.71317/RJSA.004.01.0702.
  5. Aldeen, S. D., Abbas, T., & Abbas, A. R. (2025). Review of Detecting Text generated by ChatGPT Using Machine and Deep -Learning Models: A Tools and Methods Analysis. Diyala Journal of Engineering Sciences, 18(1), 34—54. DOI: 10.24237/djes.2025.18102.
  6. Alfehaid, A., & Alkhatib, N. (2024). (Non)-Conformity to Native English Norms in Postgraduate Students’ Writing in UK Universities: Perspectives of Native and Non -Native Students and Academic Staff. International Journal of Arabic‑ ‑English Studies (IJAES), 24(1), 1—20. DOI: 10.33806/ijaes.v24i1.562.
  7. Al Fattah, N. (2018). Cognitive load theory in the context of second language academic writing. Higher Education Pedagogies, 3(1), 385—402. DOI: 10.1080/23752696.2018.1513812.
  8. Albelihi, H. H. M., & Al ‑Ahdal, A. (2024). Overcoming error fossilization in academic writing: Strategies for Saudi EFL learners to move beyond the plateau. Asian ‑Pacific Journal of Second and Foreign Language Education, 9(75). DOI: 10.1186/s40862-024-00303-y.
  9. Altakhaineh, A. R. M., Younes, A. S., & Allawama, A. (2024). A corpus--driven study of gratitude in English acknowledgements by Arabic -speaking MA students: constructing L2 academic writer identity. Cogent Arts & Humanities, 11(1). DOI: 10.1080/23311983.2024.2346361.
  10. Bender, E. M., Gebru, T., McMillan ‑Major, A., & Shmitchell, S. (2021). On the dangers of stochastic parrots: Can language models be too big? In: Proceedings of the 2021 ACM Conference on Fairness, Accountability, and Transparency (pp. 610—623). Association for Computing Machinery. DOI: 10.1145/3442188.3445922.
  11. Bajwa, N. H., König, C. J., & Kunze, T. (2020). Evidence -based understanding of introductions of research articles. Scientometrics, 124, 195—217. DOI: 10.1007/ s11192-020-03475-9.
  12. Caines, A., Benedetto, L., Taslimipoor, S., Davis, C., Gao, Y., Andersen, O. et al. (2023). On the application of large language models for language teaching and assessment technology. arXiv:2307.08393. DOI: 10.48550/arXiv.2307.08393.
  13. Canli, Z., & Yağız, O. (2024). A contrastive investigation into the non -native speakers of English academicians’ academic writing cognitions and challenges in the first and second languages. Arab World English Journal, 15(1), 117—131. DOI: 10.24093/awej/vol15no1.8.
  14. Casal, J. E., & Yoon, J. (2023). Frame -based formulaic features in L2 writing pedagogy: Variants, functions, and student writer perceptions in academic writing. English for Specific Purposes, 71, 102—114. DOI: 10.1016/j. esp.2023.03.004.
  15. Connor, U. (1996). Contrastive rhetoric. Cambridge University Press.
  16. Doughman, J., Afzal, O. M., Toyin, H. O., Shehata, S., Nakov, P., & Talat, Z. (2025, January). Exploring the limitations of detecting machine-generated text. In: Proceedings of the 31st International Conference on Computational Linguistics (pp. 4274—4281). DOI: 10.48550/arXiv.2406.11073.
  17. Du, Z., & Hashimoto, K. (2023). TCNAEC: Advancing sentence -level revision evaluation through diverse non -native academic English insights. IEEE Access, 11, 144939—144952. DOI: 10.1109/ACCESS.2023.3342862.
  18. Eid, F. M. S., & Mutahar, M. S. A. (2025). Bridging rhetorical differences: Arabic textual metaphors in academic writing and translation. European Journal of Arts, Humanities and Social Sciences, 2(2), 183—201. DOI: 10.59324/ ejahss.2025.2(2).21.
  19. Ellis, N. C., Simpson–Vlach, R. I. T. A., & Maynard, C. (2008). Formulaic language in native and second language speakers: Psycholinguistics, corpus linguistics, and TESOL. TESOL Quarterly, 42(3), 375—396. DOI: 10.1002/j.1545-7249.2008. tb00137.x.
  20. Flitcroft, M. A., Sheriff, S. N., Wolfrath, N., Maddula, R., McConnell, L., Xing, Y., & Kothari, A. N. (2024). Performance of artificial intelligence content detectors using human and artificial intelligence -generated scientific writing. Annals of Surgical Oncology, 31(10), 6387—6393. DOI: 10.1245/s10434-024-15549-6.
  21. Flowerdew, J. (2001). Attitudes of journal editors to nonnative speaker contributions. TESOL Quarterly, 35(1), 121—150. DOI: 10.2307/3587862.
  22. Ganjavi, A., & Conner, M. B. A. (2024). Publishers’ and journals’ instructions to authors on use of generative artificial intelligence in academic and scientific publishing: A bibliometric analysis. BMJ, 386, e077192. DOI: 10.1136/bmj--2023-077192.
  23. Canagarajah, A. S. (2002). A geopolitics of academic writing (Vol. 163). University of Pittsburgh Press.
  24. Gao, C. A., & Howard, F. N. S. (2023). Comparing scientific abstracts generated by ChatGPT to real abstracts with detectors and blinded human reviewers. NPJ Digital Medicine, 6(1), 75. DOI: 10.1038/s41746-023-00819-6.
  25. Gehrmann, S., Strobelt, H., & Rush, A. M. (2019). GLTR: Statistical detection and visualization of generated text. In: Proceedings of the 57th Annual Meeting of the Association for Computational Linguistics: System Demonstrations (pp. 111—116). Association for Computational Linguistics. DOI: 10.18653/v1/P19--3019.
  26. Giray, L. (2024). The problem with false positives: AI detection unfairly accuses scholars of AI plagiarism. The Serials Librarian, 85(5—6), 181—189. DOI: 10.1080/0361526X.2024.2433256.
  27. Giray, L., Santos, K., & Manalang, F. (2025). Beyond policing: AI writing detection tools, trust, academic integrity, and their implications for college. Internet Reference Services Quarterly, 29(1), 83—116. DOI: 10.1080/10875301.2024.2437174.
  28. Gosselin, R. D. (2025). AI detectors are poor Western blot classifiers: A study of accuracy and predictive values. PeerJ, 13, e18988. DOI: 10.7717/peerj.18988.
  29. Granić, A., & Marangunić, N. (2019). Technology acceptance model in educational context: A systematic literature review. British Journal of Educational Technology, 50(5), 2572—2593.DOI: 10.1111/bjet.12864.
  30. Guo, B., Zhang, X., Wang, Z., Jiang, M., Nie, J., Ding, Y., Yue, J., & Wu, Y. (2023). How close is ChatGPT to human experts? Comparison corpus, evaluation, and detection. arXiv:2301.0759. DOI: 10.48550/arXiv.2301.07597.
  31. Guo, H., Cheng, S., Zhang, K., Shen, G., & Zhang, X. (2025). CodeMirage: A multi -lingual benchmark for detecting AI-generated and paraphrased source code from production -level LLMs. arXiv:2506.11059. DOI: 10.48550/ arXiv.2506.11059.
  32. Halliday, M. A. K., & Matthiessen, C. M. I. M. (2014). Halliday’s introduction to functional grammar (4th ed.). Routledge. DOI: 10.4324/9780203431269.
  33. Haq, Z. U., Naeem, H., Naeem, A., Iqbal, F., & Zaeem, D. (2023). Comparing human and artificial intelligence in writing for health journals: An exploratory study. medRxiv. DOI: 10.1101/2023.02.22.23286322.
  34. Hinkel, E. (2001). Matters of cohesion in L2 academic texts. Applied Language Learning, 12(2), 111—132.
  35. Hinkel, E. (2002). Second language writers’ text: Linguistic and rhetorical features. Routledge.
  36. Hinkel, E. (2003). Simplicity without elegance: Features of sentences in L1 and L2 academic texts. TESOL Quarterly, 37(2), 275—301. DOI: 10.2307/3588505.
  37. Hyland, K. (2016). Academic publishing and the myth of linguistic injustice. Journal of Second Language Writing, 31, 58—69. DOI: 10.1016/j.jslw.2016.01.005.
  38. Ippolito, D., Duckworth, D., Callison ‑Burch, C., & Eck, D. (2020). Automatic detection of generated text is easiest when humans are fooled. In: Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics (pp. 1808—1822). Association for Computational Linguistics. DOI: 10.18653/v1/2020. acl-main.164.
  39. Jarvis, S., & McNamara, D. S. (2012). Detecting the first language of second language writers using automated indices of cohesion, lexical sophistication, syntactic complexity, and conceptual knowledge. In: S. Jarvis & S. A. Crossley (Eds.), Approaching language transfer through text classification: Explorations in the detection ‑based approach (pp. 106—126). Channel View Publications. DOI: 10.21832/9781847696991-005.
  40. Jarvis, S., & Pavlenko, A. (2008). Crosslinguistic influence in language and cognition. Routledge.
  41. Junina, A. K., Strauss, P. L., Wood, J. K., & Grant, L. (2025). A mixed-method inquiry into Arabic -speaking students’ experiences with English academic writing at the undergraduate level. Asia Pacific Journal of Education, 45(1), 260—277. DOI: 10.1080/02188791.2022.2101986.
  42. Kar, S. K., Bansal, T., Modi, S., & Singh, A. (2025). How Sensitive Are the Free AI -detector Tools in Detecting AI -generated Texts? A Comparison of Popular AI--detector Tools. Indian Journal of Psychological Medicine, 47(3), 275—278. DOI: 10.1177/02537176241247934.
  43. Kashiha, H., & Chan, S. H. (2015). A little bit about: Differences in native and non -native speakers’ use of formulaic language. Australian Journal of Linguistics, 35(4), 297—310. DOI: 10.1080/07268602.2015.1067132.
  44. Kehkashan, T., Riaz, R. A., Al ‑Shamayleh, A. S., Akhunzada, A., Ali, N., Hamza, M., & Akbar, F. (2025). AI -generated text detection: A comprehensive review of methods, datasets, and applications. Computer Science Review, 58, 100793. DOI: 10.1016/j.cosrev.2025.100793.
  45. Lambrecht, K. (1994). Information structure and sentence form: Topic, focus, and the mental representations of discourse referents. Cambridge University Press. DOI: 10.1017/CBO9780511620607.
  46. Laufer, B., & Waldman, T. (2011). Verb—noun collocations in second language writing: A corpus analysis of learners’ English. Language Learning, 61(2), 647— 672. DOI: 10.1111/j.1467-9922.2010.00621.x.
  47. Liang, W., Yükselgonul, M., & Mao, Y. J. (2023). GPT detectors are biased against non -native English writers. Patterns, 4(7), 100779. DOI: 10.1016/j. patter.2023.100779.
  48. Lillis, T. M., & Curry, M. J. (2010). Academic writing in global context. Routledge.
  49. Liu, J. Q., & Hui, K. A. Y. (2024). The great detectives: Humans versus AI detectors in catching large language model -generated medical writing. International Journal for Educational Integrity, 20(1), 8. DOI: 10.1007/s40979-024-00155-6.
  50. Lu, X. (2011). A corpus -based evaluation of syntactic complexity measures as indices of college -level ESL writers’ language development. TESOL Quarterly, 45(1), 36—62. DOI: 10.5054/tq.2011.240859.
  51. Mavrou, I., & Chao, J. (2023). What does linguistic distance predict when it comes to L2 writing of adult immigrant learners of Spanish? Written Communication, 40(3), 943—975. DOI: 10.1177/07410883231169511.
  52. Miralles ‑González, P., Huertas ‑Tato, J., Martín, A., & Camacho, D. (2025). Not all tokens are created equal: Perplexity Attention Weighted Networks for AI generated text detection. arXiv:2501.03940. DOI: https://doi.org/10.48550/arXiv.2501.03940.
  53. Mitchell, E., Lee, Y., Khazatsky, A., Manning, C. D., & Finn, C. (2023). DetectGPT: Zero ‑Shot Machine ‑Generated Text Detection using Probability Curvature (Version 2). arXiv:2301.11305. DOI: 10.48550/ARXIV.2301.11305.
  54. Mu, C., & Zhang, L. J. (2018). Understanding Chinese multilingual scholars’ experiences of publishing research in English. Journal of Scholarly Publishing, 49(4), 397—418. DOI: 10.3138/jsp.49.4.02.
  55. Muñoz‑Ortiz, A., Gómez‑Rodríguez, C., & Vilares, D. (2024). Contrasting linguistic patterns in human and LLM -generated news text. Artificial Intelligence Review, 57(10), 265. DOI: 10.1007/s10462-024-10903-2.
  56. Murdoch, Y. D., Lim, H., & Cho, J. (2021). Writing center visitors: Influence of L1 writing skills on students’ exophonic writings. SAGE Open, 11(4). DOI: 10.1177/21582440211062234.
  57. Odri, G. A., & Yoon, D. J. Y. (2023). Detecting generative artificial intelligence in scientific articles: Evasion techniques and implications for scientific integrity. Orthopaedics & Traumatology: Surgery & Research, 109(8), 103706. DOI: 10.1016/j.otsr.2023.103706.
  58. Opara, C. (2025). Distinguishing AI -generated and human -written text through psycholinguistic analysis. In: International Conference on Artificial Intelligence in Education (pp. 212—219). Cham: Springer Nature Switzerland. DOI: 10.48550/ arXiv.2505.01800.
  59. Öztürk, Y., & Taşçı, S. (2023). A corpus -based analysis of lexical bundles in non -native post graduate academic writing and a potential L1 influence. REFLections, 30(2), 488—505. DOI: 10.61508/refl.v30i2.267463.
  60. Pan, F. (2018). A multidimensional analysis of L1—L2 differences across three advanced levels. Southern African Linguistics and Applied Language Studies, 36(2), 117—131. DOI: 10.2989/16073614.2018.1476162.
  61. Paquot, M., & Granger, S. (2012). Formulaic language in learner corpora. Annual Review of Applied Linguistics, 32, 130—149. DOI: 10.1017/S0267190512000098.
  62. Pedersen, A. M. (2011). Writing Across Languages, Disciplines, and Sources: Second Language Writers in Jordan. Across the Disciplines, 8(1), 109—119. DOI: 10.37514/ATD-J.2011.8.1.06.
  63. Picazo ‑Sanchez, P., & Ortiz ‑Martin, L. (2024). Analysing the impact of ChatGPT in research: P. Picazo -Sanchez and L. Ortiz -Martin. Applied Intelligence, 54(5), 4172—4188. DOI: 10.1007/s10489-024-05298-0.
  64. Popkov, A. A., & Barrett, T. S. (2025). AI vs academia: Experimental study on AI text detectors’ accuracy in behavioral health academic writing. Accountability in Research, 32(7), 1072—1088. DOI: 10.1080/08989621.2024.2331757.
  65. Shin, Y. K. (2019). Do native writers always have a head start over nonnative writers? The use of lexical bundles in college students’ essays. Journal of English for Academic Purposes, 40, 1—14. DOI: 10.1016/j.jeap.2019.04.004.
  66. Showemimo, P. (2025). The AI Paradox in Academia: When Writing Too Well Raises Red Flags. SSRN. DOI: 10.2139/ssrn.5255976.
  67. Shchemeleva, I. (2021). There’s no discrimination: The association between journals, title rephrasing of the given reviewer suggestions, prevalence of text -recycling, and publishers’ policies in English. Publications, 9(1), 8. DOI: 10.3390/publications9010008.
  68. Shehata, A. M. K., & Eldakar, M. A. M. (2018). Publishing research in the international context: An analysis of Egyptian social sciences scholars’ academic writing behaviour. The Electronic Library, 36(5), 910—924. DOI: 10.1108/EL-01-2017-0005.
  69. Subandowo, D., Sárdi, C., & Thresia, F. (2025). An investigation of English academic writing strategies employed by Indonesian graduate students in an English medium instruction (EMI) context. Asian ‑Pacific Journal of Second and Foreign Language Education, 10(1), 38. DOI: 10.1186/s40862-025-00345-w.
  70. Taşçı, S., & Öztürk, Y. (2021). Post -predicate that -clauses controlled by verbs in native and non -native academic writing: A corpus -based study. Australian Journal of Applied Linguistics, 4(1), 18—33. DOI: 10.29140/ajal.v4n1.486.
  71. Tufts, B., Zhao, X., & Li, L. (2025). A practical examination of AI -generated text detectors for large language models. In: Findings of the Association for Computational Linguistics: NAACL 2025 (pp. 4824—4841). DOI: 10.48550/ arXiv.2412.05139.
  72. Uchendu, A., Le, T., Shu, K., & Lee, D. (2020). Authorship attribution for neural text generation. In: Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing (EMNLP) (pp. 8384—8395). Association for Computational Linguistics. DOI: 10.18653/v1/2020.emnlp-main.673.
  73. Valipour, V., Assadi, N., & Asl, H. D. (2017). The generic structures and lexico -grammaticality in English academic research papers. Southern African Linguistics and Applied Language Studies, 35(2), 169—182. DOI: 10.2989/16073614.2017.1373365.
  74. Wang, X., Zhang, W., & Rajtmajer, S. (2024). Monolingual and multilingual misinformation detection for low -resource languages: A comprehensive survey. arXiv:2410.18390. DOI: 10.48550/arXiv.2410.18390.
  75. Weber‑Wulff, D., Anohina ‑Naumeca, A., & Bjelobaba, S. L. (2023). Testing of detection tools for AI -generated text. International Journal for Educational Integrity, 19(1), 1—39. DOI: 10.1007/s40979-023-00146-z.
  76. Won, D., Shin, Y. K., Kim, H. & Yoo, I. W. (2025). Advancing Language Assessment with GPT: Is It Nonnative -Language Friendly? Language Assessment Quarterly, 1, 1—18. DOI: 10.1080/15434303.2024.2444349.
  77. Zaini, A., & Ollerhead, S. (2019). Reverse contrastive rhetoric in expository writing: Transfer and power relations at work. Southern African Linguistics and Applied Language Studies, 37(1), 41—61. DOI: 10.2989/16073614.2019.1609364.
Language: English
Page range: 37 - 78
Submitted on: Mar 16, 2026
Accepted on: Apr 26, 2026
Published on: Jul 3, 2026
Published by: SAN University
In partnership with: Paradigm Publishing Services
Publication frequency: 2 issues per year

© 2026 Rania Za’rour, Abdel Rahman Mitib Altakhaineh, published by SAN University
This work is licensed under the Creative Commons Attribution-NonCommercial-NoDerivatives 4.0 License.