Skip to main content
Have a personal or library account? Click to login
Pedagogical Efficacy of LLM-Generated Synthetic Data Versus Real-World Clinical Records: A Randomized Controlled Non-Inferiority Trial Cover

Pedagogical Efficacy of LLM-Generated Synthetic Data Versus Real-World Clinical Records: A Randomized Controlled Non-Inferiority Trial

Open Access
|Sep 2026

References

  1. Elstein AS. Beyond multiple-choice questions and essays: the need for a new way to assess clinical competence. Academic medicine: journal of the Association of American Medical Colleges. 1993;68:244249. DOI: 10.1097/00001888-199304000-00002
  2. Cooper N, Bartlett M, Gay S, Hammond A, Lillicrap M, Matthan J, et al. Consensus statement on the content of clinical reasoning curricula in undergraduate medical education. Medical teacher. 2021;43:152159. DOI: 10.1080/0142159X.2020.1842343
  3. Donkin R, Yule H, Fyfe T. Online case-based learning in medical education: a scoping review. BMC medical education. 2023;23:564. DOI: 10.1186/s12909-023-04520-w
  4. Bakkum MJ, Hartjes MG, Piët JD, Donker EM, Likic R, Sanz E, et al. Using artificial intelligence to create diverse and inclusive medical case vignettes for education. Br J Clin Pharmacol. 2024;90:640648. DOI: 10.1111/bcp.15977
  5. Baranowski MLH, Stoff BK. Should medical students follow up on skin biopsy results? When education conflicts with patient privacy. J Am Acad Dermatol. 2018;78:12291231. DOI: 10.1016/j.jaad.2017.09.028
  6. Bittner JG, Logghe HJ, Kane ED, Goldberg RF, Alseidi A, Aggarwal R, et al. A Society of Gastrointestinal and Endoscopic Surgeons (SAGES) statement on closed social media (Facebook®) groups for clinical education and consultation: issues of informed consent, patient privacy, and surgeon protection. Surg Endosc. 2019;33:17. DOI: 10.1007/s00464-018-6569-2
  7. Giuffrè M, Shung DL. Harnessing the power of synthetic data in healthcare: innovation, application, and privacy. NPJ digital medicine. 2023;6:186. DOI: 10.1038/s41746-023-00927-3
  8. Kıyak YS, Górski S, Tokarek T, Pers M, Kononowicz AA. Large language models for generating key-feature questions in medical education. Med Educ Online. 2025;30:2574647. DOI: 10.1080/10872981.2025.2574647
  9. Zhang Q, Huang Z, Huang Y, Wang G, Zhang R, Yang J, et al. Generative AI in medical education: feasibility and educational value of LLM-generated clinical cases with MCQs. BMC medical education. 2025;25:502. DOI: 10.1186/s12909-025-08085-8
  10. Stretton B, Kovoor J, Arnold M, Bacchi S. ChatGPT-Based Learning: Generative Artificial Intelligence in Medical Education. Medical science educator. 2024;34:215217. DOI: 10.1007/s40670-023-01934-5
  11. Emekli E, Emekli E, Özel B. Artificial Intelligence-Assisted Generation of Case Scenarios and Multiple-Choice Questions in Psychiatry: A Pilot Study. Acad Psychiatry. 2025. DOI: 10.1007/s40596-025-02298-1
  12. Zhou N, Wu Q, Wu Z, Marino S, Dinov ID. DataSifterText: Partially Synthetic Text Generation for Sensitive Clinical Notes. Journal of medical systems. 2022;46:96. DOI: 10.1007/s10916-022-01880-6
  13. Abd-Alrazaq A, AlSaad R, Alhuwail D, Ahmed A, Healy PM, Latifi S, et al. Large Language Models in Medical Education: Opportunities, Challenges, and Future Directions. JMIR medical education. 2023;9:e48291. DOI: 10.2196/48291
  14. Dar SUH, Seyfarth M, Ayx I, Papavassiliu T, Schoenberg SO, Siepmann RM, et al. Unconditional latent diffusion models memorize patient imaging data. Nature biomedical engineering. 2026;10:458472. DOI: 10.1038/s41551-025-01468-8
  15. Herrmann-Werner A, Festl-Wietek T, Holderried F, Herschbach L, Griewatz J, Masters K, et al. Assessing ChatGPT’s Mastery of Bloom’s Taxonomy Using Psychosomatic Medicine Exam Questions: Mixed-Methods Study. Journal of medical Internet research. 2024;26:e52113. DOI: 10.2196/52113
  16. Wang C, Sun N, Pan X, Liu P, Ma Z, Wang W, et al. The impact of integrating generative artificial intelligence into medical education on short-term learning outcomes: a systematic review and meta-analysis of randomized controlled trials. BMC medical education [Internet]. 2026. DOI: 10.1186/s12909-026-09320-6
  17. Furfaro D, Celi LA, Schwartzstein RM. Artificial Intelligence in Medical Education: A Long Way to Go. Chest. 2024;165:771774. DOI: 10.1016/j.chest.2023.11.028
  18. Ranganathan P, Pramesh CS, Aggarwal R. Non-inferiority trials. Perspectives in clinical research. 2022;13:5457. DOI: 10.4103/picr.picr_245_21
  19. Cox K. Stories as case knowledge: case knowledge as stories. Medical education. 2001;35:862866. DOI: 10.1046/j.1365-2923.2001.01016.x
  20. Dudas RA, Bannister SL. Enhancing Diagnostic Reasoning in Medical Education Through Patient Stories and Illness Scripts. Pediatrics [Internet]. 2025;156. DOI: 10.1542/peds.2025-072981
  21. Leon M, Feng R, Quiroz Flores M, Pelletier G, Bethencourt D, Shibata M, et al. Blinded two-phase evaluation of large language models in complex cardiac surgery: task-specific performance and human-AI collaboration. Frontiers in digital health. 2026;8:769467. DOI: 10.3389/fdgth.2026.1769467
  22. Thampy H, Willert E, Ramani S. Assessing Clinical Reasoning: Targeting the Higher Levels of the Pyramid. Journal of general internal medicine. 2019;34:16311636. DOI: 10.1007/s11606-019-04953-4
  23. Mee J, Pandian R, Wolczynski J, Morales A, Paniagua M, Harik P, et al. An experimental comparison of multiple-choice and short-answer questions on a high-stakes test for medical students. Advances in health sciences education: theory and practice. 2024;29:783801. DOI: 10.1007/s10459-023-10266-3
  24. Abdulnour R-EE, Gin B, Boscardin CK. Educational Strategies for Clinical Supervision of Artificial Intelligence Use. N Engl J Med. 2025;393:786797. DOI: 10.1056/NEJMra2503232
  25. Goh E, Gallo R, Hom J, Strong E, Weng Y, Kerman H, et al. Large Language Model Influence on Diagnostic Reasoning: A Randomized Clinical Trial. JAMA Netw Open. 2024;7:e2440969. DOI: 10.1001/jamanetworkopen.2024.40969
  26. Nielsen JPS, Mikkelsen AK, Kuenzel J, Sebelik ME, Madani G, Yang T-L, et al. Evaluation of Multiple-Choice Tests in Head and Neck Ultrasound Created by Physicians and Large Language Models. Diagnostics (Basel, Switzerland) [Internet]. 2025;15. DOI: 10.3390/diagnostics15151848
  27. Kıyak YS, Kaya AB, Emekli E. Validity of AI-generated multiple-choice questions in medical education: a systematic review. Postgraduate Medical Journal [Internet]. 2026. DOI: 10.1093/postmj/qgag057
  28. Liu P, Zhang J, Chen S, Chen S. Human-AI teaming in healthcare: 1 + 1 > 2? npj Artificial Intelligence [Internet]. 2025. DOI: 10.1038/s44387-025-00052-4
  29. Reiter JP, Mitra R. Estimating Risks of Identification Disclosure in Partially Synthetic Data. Journal of Privacy and Confidentiality [Internet]. 2009. DOI: 10.29012/jpc.v1i1.567
  30. Gordon D, Rencic JJ, Lang VJ, Thomas A, Young M, Durning SJ. Advancing the assessment of clinical reasoning across the health professions: Definitional and methodologic recommendations. Perspectives on medical education. 2022;11:108114. DOI: 10.1007/s40037-022-00701-3
  31. Katsanos AH, Lioutas V-A, Yperzeele L, Ullberg T, Li L, Ramage ER, et al. Perception and acquaintance of stroke specialists on non-inferiority trials: An international survey. Journal of stroke and cerebrovascular diseases: the official journal of National Stroke Association. 2025;34:08132. DOI: 10.1016/j.jstrokecerebrovasdis.2024.108132
  32. Chen L, Zaharia M, Zou J. How Is ChatGPT’s Behavior Changing Over Time? Harvard Data Science Review [Internet]. 2024;6. DOI: 10.1162/99608f92.5317da47
DOI: https://doi.org/10.5334/pme.2535 | Journal eISSN: 2212-277X
Language: English
Page range: 759 - 770
Submitted on: Mar 2, 2026
Accepted on: Aug 20, 2026
Published on: Sep 9, 2026
Published by: Ubiquity Press
In partnership with: Paradigm Publishing Services

© 2026 Zhongkun Zuo, Jianyu Fang, Muyao Ye, Zedong Li, published by Ubiquity Press
This work is licensed under the Creative Commons Attribution 4.0 License.