Skip to main content
Have a personal or library account? Click to login
Pedagogical Efficacy of LLM-Generated Synthetic Data Versus Real-World Clinical Records: A Randomized Controlled Non-Inferiority Trial Cover

Pedagogical Efficacy of LLM-Generated Synthetic Data Versus Real-World Clinical Records: A Randomized Controlled Non-Inferiority Trial

Open Access
|Sep 2026

Figures & Tables

Figure 1

Preparation of study materials, participant flow, and the shared online educational process. Phase A shows how five fully de-identified, matched real cases served as the common source for both groups’ materials: experienced clinicians compiled them into structured real-case-derived teaching cases, while matched AI-generated teaching cases were produced using Gemini 3.0 Pro through Google AI Studio, with model outputs not manually edited. Each case pair entered the study only after unanimous review by three senior general surgery specialists with full-professor rank. Phase B shows automatic 1:1 randomization on the Wenjuanxing platform and the modified intention-to-treat set. After randomization, 9 participants in the real-case group and 8 in the AI-generated case group were excluded because the total response time was shorter than 5 minutes or because responses showed a completely uniform pattern; there were no other withdrawals, group changes, or exclusions. Phase C outlines the shared online educational process. Participants were not informed of their group or of the case source; this statement does not imply that case-source masking was necessarily successful.

Table 1

Baseline characteristics.

CHARACTERISTICREAL CASES N = 192AI-GENERATED CASES N = 194P VALUE
Training stage0.150
    Year 348 (25.0%)67 (34.5%)
    Year 467 (34.9%)53 (27.3%)
    Intern42 (21.9%)36 (18.6%)
    Resident35 (18.2%)38 (19.6%)
Male gender96 (50.0%)99 (51.0%)0.800
Prior surgical rotation99 (51.6%)85 (43.8%)0.130
Baseline confidence0.300
    13 (1.6%)1 (0.5%)
    233 (17.2%)31 (16.0%)
    393 (48.4%)82 (42.3%)
    454 (28.1%)62 (32.0%)
    59 (4.7%)18 (9.3%)
Learning preference0.900
    Preference 157 (29.7%)58 (29.9%)
    Preference 270 (36.5%)75 (38.7%)
    Preference 365 (33.9%)61 (31.4%)

[i] Values are n (%). P values are from Pearson chi-square or Fisher exact tests, as appropriate.

Table 2

Item analysis of the 15-item training-phase assessment.

ITEMDIFFICULTY INDEXCORRECTED ITEM-TOTAL CORRELATIONUPPER-LOWER DISCRIMINATIONALPHA IF ITEM DELETED
Case 1: diagnosis0.6760.5950.7400.908
Case 1: reasoning0.6300.6080.7720.908
Case 1: decision-making0.6090.5780.7720.909
Case 2: diagnosis0.5490.6570.8530.906
Case 2: reasoning0.4950.6100.7780.907
Case 2: decision-making0.3940.5560.7180.909
Case 3: diagnosis0.5570.6840.9020.905
Case 3: reasoning0.5230.6500.8270.906
Case 3: decision-making0.4380.6900.8750.905
Case 4: diagnosis0.4270.6890.9090.905
Case 4: reasoning0.3730.5790.7260.908
Case 4: decision-making0.2980.4710.5440.912
Case 5: diagnosis0.4380.7030.9010.904
Case 5: reasoning0.3680.6300.7840.907
Case 5: decision-making0.2980.4580.5270.912

[i] Overall Cronbach alpha = 0.913 (bootstrap 95% CI 0.901 to 0.924). Difficulty index is the proportion correct; higher values indicate easier items.

Table 3

Learning outcomes.

OUTCOMEREAL CASES N = 192AI-GENERATED CASES N = 194EFFECT ESTIMATE (AI – REAL)P VALUE
Training-phase score (0–15)7.27 (4.97)6.88 (4.85)–0.390.419
Post-training score (0–10)4.95 (3.35)4.61 (3.35)–0.335 (–1.006 to 0.337)0.314
Pass rate (score >= 6)82/192 (42.7%)77/194 (39.7%)–3.0 percentage points0.618

[i] Values are mean (SD) unless otherwise stated. For the primary post-training outcome, the prespecified non-inferiority margin was –0.5 points; the confidence interval crossed this margin, so non-inferiority was not demonstrated. P = 0.314 is the one-sided non-inferiority P value. Training-phase score used the Wilcoxon rank-sum test; pass rate used the Yates-corrected chi-square test.

Figure 2

Learning process and case perceptions. (A) Learning efficiency index, defined as the training-phase performance score divided by cumulative reading time for the five cases (in minutes). (B) Distribution of case-source judgments; each participant made five judgments, and percentages were calculated by randomized group. (C) Mean case realism ratings summarized by prespecified case-complexity level. Case 1 was Level 1, Cases 2 and 3 were Level 2, and Cases 4 and 5 were Level 3. Error bars indicate descriptive 95% confidence intervals.

Table 4

Learner experience and attitudes.

OUTCOMEREAL CASES N = 192AI-GENERATED CASES N = 194EFFECT OR COMPARISONP VALUE
Perceived mental effort (1–9)4.79 (1.41)4.74 (1.44)–0.05 (–0.34 to 0.23)0.707
Overall satisfaction (1–5)4.22 (0.72)4.30 (0.73)Categorical distribution0.490
AI attitude: support113 (58.9%)115 (59.3%)
AI attitude: neutral58 (30.2%)64 (33.0%)Three-category distribution0.521
AI attitude: oppose21 (10.9%)15 (7.7%)

[i] Mental effort was measured with a single-item 1–9 scale, not the NASA-TLX. The mental-effort comparison used Welch’s t test; satisfaction and AI attitude used categorical chi-square tests.

Figure 3

Exploratory and safety-related analyses. (A) Participant-level mean proportion correct for diagnosis-, reasoning-, and decision-making-labeled items among the 15 training-phase items. (B) Mean post-training test scores summarized by training stage; error bars indicate descriptive 95% confidence intervals. (C) Proportion of participants with at least one high-confidence completely incorrect response, defined as a per-case score of 0/3 with a confidence rating of 4/4. Error bars indicate exact 95% confidence intervals; Fisher’s exact test P = 1.000.

DOI: https://doi.org/10.5334/pme.2535 | Journal eISSN: 2212-277X
Language: English
Page range: 759 - 770
Submitted on: Mar 2, 2026
Accepted on: Aug 20, 2026
Published on: Sep 9, 2026
Published by: Ubiquity Press
In partnership with: Paradigm Publishing Services

© 2026 Zhongkun Zuo, Jianyu Fang, Muyao Ye, Zedong Li, published by Ubiquity Press
This work is licensed under the Creative Commons Attribution 4.0 License.