
Pedagogical Efficacy of LLM-Generated Synthetic Data Versus Real-World Clinical Records: A Randomized Controlled Non-Inferiority Trial
Abstract
Background: Expert-reviewed clinical cases generated by large language models (LLMs) may supplement case resources in medical education, but their short-term educational performance relative to real-case-derived teaching materials remains uncertain. We compared immediate post-training test performance after teaching with the two types of case materials and assessed non-inferiority against a prespecified margin.
Methods: We conducted a prospective, parallel-group, randomized non-inferiority trial. Through the Wenjuanxing online platform, participants were randomized 1:1 to learn with either real-case-derived teaching cases compiled by clinicians and reviewed by experts or AI-generated clinical cases produced by Gemini 3.0 Pro from fully de-identified matched real cases and reviewed by three senior general surgery specialists with full-professor rank. The primary outcome was the total score on an independent 10-item immediate post-training test (0–10 points), with a prespecified non-inferiority margin of –0.5 points. Secondary outcomes included the training-phase performance score, learning efficiency index, single-item mental effort rating, case realism, and case-source judgment.
Results: A total of 403 participants were randomized, of whom 386 were included in the modified intention-to-treat analysis: 192 in the real-case group and 194 in the AI-generated case group. The mean post-training test score was 4.95 (SD, 3.35) in the real-case group and 4.61 (SD, 3.35) in the AI-generated case group. The mean difference (AI-generated minus real-case group) was –0.335 points (95% CI, –1.006 to 0.337). Because the lower bound of the confidence interval was below the prespecified non-inferiority margin of –0.5 points, non-inferiority was not demonstrated (one-sided P = 0.314). No significant between-group differences were observed in the training-phase performance score, learning efficiency index, or single-item mental effort rating. AI-generated cases received lower realism ratings for Level 3 cases. The proportion of participants with at least one high-confidence completely incorrect response was 1.6% in the real-case group and 2.1% in the AI-generated case group.
Conclusions: In this short-term, text-based online case-learning setting, no statistically significant between-group difference was observed in immediate post-training test performance; however, non-inferiority of AI-generated clinical cases relative to real-case-derived teaching materials was not demonstrated.
© 2026 Zhongkun Zuo, Jianyu Fang, Muyao Ye, Zedong Li, published by Ubiquity Press
This work is licensed under the Creative Commons Attribution 4.0 License.