
Figure 1
Preparation of study materials, participant flow, and the shared online educational process. Phase A shows how five fully de-identified, matched real cases served as the common source for both groups’ materials: experienced clinicians compiled them into structured real-case-derived teaching cases, while matched AI-generated teaching cases were produced using Gemini 3.0 Pro through Google AI Studio, with model outputs not manually edited. Each case pair entered the study only after unanimous review by three senior general surgery specialists with full-professor rank. Phase B shows automatic 1:1 randomization on the Wenjuanxing platform and the modified intention-to-treat set. After randomization, 9 participants in the real-case group and 8 in the AI-generated case group were excluded because the total response time was shorter than 5 minutes or because responses showed a completely uniform pattern; there were no other withdrawals, group changes, or exclusions. Phase C outlines the shared online educational process. Participants were not informed of their group or of the case source; this statement does not imply that case-source masking was necessarily successful.
Table 1
Baseline characteristics.
| CHARACTERISTIC | REAL CASES N = 192 | AI-GENERATED CASES N = 194 | P VALUE |
|---|---|---|---|
| Training stage | 0.150 | ||
| Year 3 | 48 (25.0%) | 67 (34.5%) | |
| Year 4 | 67 (34.9%) | 53 (27.3%) | |
| Intern | 42 (21.9%) | 36 (18.6%) | |
| Resident | 35 (18.2%) | 38 (19.6%) | |
| Male gender | 96 (50.0%) | 99 (51.0%) | 0.800 |
| Prior surgical rotation | 99 (51.6%) | 85 (43.8%) | 0.130 |
| Baseline confidence | 0.300 | ||
| 1 | 3 (1.6%) | 1 (0.5%) | |
| 2 | 33 (17.2%) | 31 (16.0%) | |
| 3 | 93 (48.4%) | 82 (42.3%) | |
| 4 | 54 (28.1%) | 62 (32.0%) | |
| 5 | 9 (4.7%) | 18 (9.3%) | |
| Learning preference | 0.900 | ||
| Preference 1 | 57 (29.7%) | 58 (29.9%) | |
| Preference 2 | 70 (36.5%) | 75 (38.7%) | |
| Preference 3 | 65 (33.9%) | 61 (31.4%) |
[i] Values are n (%). P values are from Pearson chi-square or Fisher exact tests, as appropriate.
Table 2
Item analysis of the 15-item training-phase assessment.
| ITEM | DIFFICULTY INDEX | CORRECTED ITEM-TOTAL CORRELATION | UPPER-LOWER DISCRIMINATION | ALPHA IF ITEM DELETED |
|---|---|---|---|---|
| Case 1: diagnosis | 0.676 | 0.595 | 0.740 | 0.908 |
| Case 1: reasoning | 0.630 | 0.608 | 0.772 | 0.908 |
| Case 1: decision-making | 0.609 | 0.578 | 0.772 | 0.909 |
| Case 2: diagnosis | 0.549 | 0.657 | 0.853 | 0.906 |
| Case 2: reasoning | 0.495 | 0.610 | 0.778 | 0.907 |
| Case 2: decision-making | 0.394 | 0.556 | 0.718 | 0.909 |
| Case 3: diagnosis | 0.557 | 0.684 | 0.902 | 0.905 |
| Case 3: reasoning | 0.523 | 0.650 | 0.827 | 0.906 |
| Case 3: decision-making | 0.438 | 0.690 | 0.875 | 0.905 |
| Case 4: diagnosis | 0.427 | 0.689 | 0.909 | 0.905 |
| Case 4: reasoning | 0.373 | 0.579 | 0.726 | 0.908 |
| Case 4: decision-making | 0.298 | 0.471 | 0.544 | 0.912 |
| Case 5: diagnosis | 0.438 | 0.703 | 0.901 | 0.904 |
| Case 5: reasoning | 0.368 | 0.630 | 0.784 | 0.907 |
| Case 5: decision-making | 0.298 | 0.458 | 0.527 | 0.912 |
[i] Overall Cronbach alpha = 0.913 (bootstrap 95% CI 0.901 to 0.924). Difficulty index is the proportion correct; higher values indicate easier items.
Table 3
Learning outcomes.
| OUTCOME | REAL CASES N = 192 | AI-GENERATED CASES N = 194 | EFFECT ESTIMATE (AI – REAL) | P VALUE |
|---|---|---|---|---|
| Training-phase score (0–15) | 7.27 (4.97) | 6.88 (4.85) | –0.39 | 0.419 |
| Post-training score (0–10) | 4.95 (3.35) | 4.61 (3.35) | –0.335 (–1.006 to 0.337) | 0.314 |
| Pass rate (score >= 6) | 82/192 (42.7%) | 77/194 (39.7%) | –3.0 percentage points | 0.618 |
[i] Values are mean (SD) unless otherwise stated. For the primary post-training outcome, the prespecified non-inferiority margin was –0.5 points; the confidence interval crossed this margin, so non-inferiority was not demonstrated. P = 0.314 is the one-sided non-inferiority P value. Training-phase score used the Wilcoxon rank-sum test; pass rate used the Yates-corrected chi-square test.

Figure 2
Learning process and case perceptions. (A) Learning efficiency index, defined as the training-phase performance score divided by cumulative reading time for the five cases (in minutes). (B) Distribution of case-source judgments; each participant made five judgments, and percentages were calculated by randomized group. (C) Mean case realism ratings summarized by prespecified case-complexity level. Case 1 was Level 1, Cases 2 and 3 were Level 2, and Cases 4 and 5 were Level 3. Error bars indicate descriptive 95% confidence intervals.
Table 4
Learner experience and attitudes.
| OUTCOME | REAL CASES N = 192 | AI-GENERATED CASES N = 194 | EFFECT OR COMPARISON | P VALUE |
|---|---|---|---|---|
| Perceived mental effort (1–9) | 4.79 (1.41) | 4.74 (1.44) | –0.05 (–0.34 to 0.23) | 0.707 |
| Overall satisfaction (1–5) | 4.22 (0.72) | 4.30 (0.73) | Categorical distribution | 0.490 |
| AI attitude: support | 113 (58.9%) | 115 (59.3%) | ||
| AI attitude: neutral | 58 (30.2%) | 64 (33.0%) | Three-category distribution | 0.521 |
| AI attitude: oppose | 21 (10.9%) | 15 (7.7%) |
[i] Mental effort was measured with a single-item 1–9 scale, not the NASA-TLX. The mental-effort comparison used Welch’s t test; satisfaction and AI attitude used categorical chi-square tests.

Figure 3
Exploratory and safety-related analyses. (A) Participant-level mean proportion correct for diagnosis-, reasoning-, and decision-making-labeled items among the 15 training-phase items. (B) Mean post-training test scores summarized by training stage; error bars indicate descriptive 95% confidence intervals. (C) Proportion of participants with at least one high-confidence completely incorrect response, defined as a per-case score of 0/3 with a confidence rating of 4/4. Error bars indicate exact 95% confidence intervals; Fisher’s exact test P = 1.000.
