
Pedagogical Efficacy of LLM-Generated Synthetic Data Versus Real-World Clinical Records: A Randomized Controlled Non-Inferiority Trial
By: Zhongkun Zuo, Jianyu Fang, Muyao Ye and Zedong Li
References
- Elstein AS. Beyond multiple-choice questions and essays: the need for a new way to assess clinical competence. Academic medicine: journal of the Association of American Medical Colleges. 1993;68:244–249. DOI: 10.1097/00001888-199304000-00002
- Cooper N, Bartlett M, Gay S, Hammond A, Lillicrap M, Matthan J, et al. Consensus statement on the content of clinical reasoning curricula in undergraduate medical education. Medical teacher. 2021;43:152–159. DOI: 10.1080/0142159X.2020.1842343
- Donkin R, Yule H, Fyfe T. Online case-based learning in medical education: a scoping review. BMC medical education. 2023;23:564. DOI: 10.1186/s12909-023-04520-w
- Bakkum MJ, Hartjes MG, Piët JD, Donker EM, Likic R, Sanz E, et al. Using artificial intelligence to create diverse and inclusive medical case vignettes for education. Br J Clin Pharmacol. 2024;90:640–648. DOI: 10.1111/bcp.15977
- Baranowski MLH, Stoff BK. Should medical students follow up on skin biopsy results? When education conflicts with patient privacy. J Am Acad Dermatol. 2018;78:1229–1231. DOI: 10.1016/j.jaad.2017.09.028
- Bittner JG, Logghe HJ, Kane ED, Goldberg RF, Alseidi A, Aggarwal R, et al. A Society of Gastrointestinal and Endoscopic Surgeons (SAGES) statement on closed social media (Facebook®) groups for clinical education and consultation: issues of informed consent, patient privacy, and surgeon protection. Surg Endosc. 2019;33:1–7. DOI: 10.1007/s00464-018-6569-2
- Giuffrè M, Shung DL. Harnessing the power of synthetic data in healthcare: innovation, application, and privacy. NPJ digital medicine. 2023;6:186. DOI: 10.1038/s41746-023-00927-3
- Kıyak YS, Górski S, Tokarek T, Pers M, Kononowicz AA. Large language models for generating key-feature questions in medical education. Med Educ Online. 2025;30:2574647. DOI: 10.1080/10872981.2025.2574647
- Zhang Q, Huang Z, Huang Y, Wang G, Zhang R, Yang J, et al. Generative AI in medical education: feasibility and educational value of LLM-generated clinical cases with MCQs. BMC medical education. 2025;25:502. DOI: 10.1186/s12909-025-08085-8
- Stretton B, Kovoor J, Arnold M, Bacchi S. ChatGPT-Based Learning: Generative Artificial Intelligence in Medical Education. Medical science educator. 2024;34:215–217. DOI: 10.1007/s40670-023-01934-5
- Emekli E, Emekli E, Özel B. Artificial Intelligence-Assisted Generation of Case Scenarios and Multiple-Choice Questions in Psychiatry: A Pilot Study. Acad Psychiatry. 2025. DOI: 10.1007/s40596-025-02298-1
- Zhou N, Wu Q, Wu Z, Marino S, Dinov ID. DataSifterText: Partially Synthetic Text Generation for Sensitive Clinical Notes. Journal of medical systems. 2022;46:96. DOI: 10.1007/s10916-022-01880-6
- Abd-Alrazaq A, AlSaad R, Alhuwail D, Ahmed A, Healy PM, Latifi S, et al. Large Language Models in Medical Education: Opportunities, Challenges, and Future Directions. JMIR medical education. 2023;9:
e48291 . DOI: 10.2196/48291 - Dar SUH, Seyfarth M, Ayx I, Papavassiliu T, Schoenberg SO, Siepmann RM, et al. Unconditional latent diffusion models memorize patient imaging data. Nature biomedical engineering. 2026;10:458–472. DOI: 10.1038/s41551-025-01468-8
- Herrmann-Werner A, Festl-Wietek T, Holderried F, Herschbach L, Griewatz J, Masters K, et al. Assessing ChatGPT’s Mastery of Bloom’s Taxonomy Using Psychosomatic Medicine Exam Questions: Mixed-Methods Study. Journal of medical Internet research. 2024;26:
e52113 . DOI: 10.2196/52113 - Wang C, Sun N, Pan X, Liu P, Ma Z, Wang W, et al. The impact of integrating generative artificial intelligence into medical education on short-term learning outcomes: a systematic review and meta-analysis of randomized controlled trials. BMC medical education [Internet]. 2026. DOI: 10.1186/s12909-026-09320-6
- Furfaro D, Celi LA, Schwartzstein RM. Artificial Intelligence in Medical Education: A Long Way to Go. Chest. 2024;165:771–774. DOI: 10.1016/j.chest.2023.11.028
- Ranganathan P, Pramesh CS, Aggarwal R. Non-inferiority trials. Perspectives in clinical research. 2022;13:54–57. DOI: 10.4103/picr.picr_245_21
- Cox K. Stories as case knowledge: case knowledge as stories. Medical education. 2001;35:862–866. DOI: 10.1046/j.1365-2923.2001.01016.x
- Dudas RA, Bannister SL. Enhancing Diagnostic Reasoning in Medical Education Through Patient Stories and Illness Scripts. Pediatrics [Internet]. 2025;156. DOI: 10.1542/peds.2025-072981
- Leon M, Feng R, Quiroz Flores M, Pelletier G, Bethencourt D, Shibata M, et al. Blinded two-phase evaluation of large language models in complex cardiac surgery: task-specific performance and human-AI collaboration. Frontiers in digital health. 2026;8:769467. DOI: 10.3389/fdgth.2026.1769467
- Thampy H, Willert E, Ramani S. Assessing Clinical Reasoning: Targeting the Higher Levels of the Pyramid. Journal of general internal medicine. 2019;34:1631–1636. DOI: 10.1007/s11606-019-04953-4
- Mee J, Pandian R, Wolczynski J, Morales A, Paniagua M, Harik P, et al. An experimental comparison of multiple-choice and short-answer questions on a high-stakes test for medical students. Advances in health sciences education: theory and practice. 2024;29:783–801. DOI: 10.1007/s10459-023-10266-3
- Abdulnour R-EE, Gin B, Boscardin CK. Educational Strategies for Clinical Supervision of Artificial Intelligence Use. N Engl J Med. 2025;393:786–797. DOI: 10.1056/NEJMra2503232
- Goh E, Gallo R, Hom J, Strong E, Weng Y, Kerman H, et al. Large Language Model Influence on Diagnostic Reasoning: A Randomized Clinical Trial. JAMA Netw Open. 2024;7:
e2440969 . DOI: 10.1001/jamanetworkopen.2024.40969 - Nielsen JPS, Mikkelsen AK, Kuenzel J, Sebelik ME, Madani G, Yang T-L, et al. Evaluation of Multiple-Choice Tests in Head and Neck Ultrasound Created by Physicians and Large Language Models. Diagnostics (Basel, Switzerland) [Internet]. 2025;15. DOI: 10.3390/diagnostics15151848
- Kıyak YS, Kaya AB, Emekli E. Validity of AI-generated multiple-choice questions in medical education: a systematic review. Postgraduate Medical Journal [Internet]. 2026. DOI: 10.1093/postmj/qgag057
- Liu P, Zhang J, Chen S, Chen S. Human-AI teaming in healthcare: 1 + 1 > 2? npj Artificial Intelligence [Internet]. 2025. DOI: 10.1038/s44387-025-00052-4
- Reiter JP, Mitra R. Estimating Risks of Identification Disclosure in Partially Synthetic Data. Journal of Privacy and Confidentiality [Internet]. 2009. DOI: 10.29012/jpc.v1i1.567
- Gordon D, Rencic JJ, Lang VJ, Thomas A, Young M, Durning SJ. Advancing the assessment of clinical reasoning across the health professions: Definitional and methodologic recommendations. Perspectives on medical education. 2022;11:108–114. DOI: 10.1007/s40037-022-00701-3
- Katsanos AH, Lioutas V-A, Yperzeele L, Ullberg T, Li L, Ramage ER, et al. Perception and acquaintance of stroke specialists on non-inferiority trials: An international survey. Journal of stroke and cerebrovascular diseases: the official journal of National Stroke Association. 2025;34:08132. DOI: 10.1016/j.jstrokecerebrovasdis.2024.108132
- Chen L, Zaharia M, Zou J. How Is ChatGPT’s Behavior Changing Over Time? Harvard Data Science Review [Internet]. 2024;6. DOI: 10.1162/99608f92.5317da47
DOI: https://doi.org/10.5334/pme.2535 | Journal eISSN: 2212-277X
Language: English
Page range: 759 - 770
Submitted on: Mar 2, 2026
Accepted on: Aug 20, 2026
Published on: Sep 9, 2026
Published by: Ubiquity Press
In partnership with: Paradigm Publishing Services
© 2026 Zhongkun Zuo, Jianyu Fang, Muyao Ye, Zedong Li, published by Ubiquity Press
This work is licensed under the Creative Commons Attribution 4.0 License.