Skip to main content
Have a personal or library account? Click to login
Performance of ChatGPT and GPT-4 on Polish National Specialty Exam (NSE) in Ophthalmology Cover

Performance of ChatGPT and GPT-4 on Polish National Specialty Exam (NSE) in Ophthalmology

Open Access
|Sep 2024

Figures & Tables

Table 1.

Overall proportion of correct and incorrect answers (McNemar’s test: chi-squared = 13.4; df=1; p <0.001; OR (95% CI) = 3.78 (1.78–8.96))

LLMGPT-4
GPT-3.5Correct answerYesNoRow total
Yes28 (28.6%)9 (9.2%)37 (37.8%)
No34 (34.7%)27 (27.6%)61 (62.2%)
Column total62 (63.3%)36 (36.7%)N = 98

[i] PA= (a+b)/N= 37/98= 0.3776 (0.2879, 0.4764);

[ii] PB= (a+c)/N= 62/98= 0.6327 (0.5339, 0.7214);

[iii] PB – PA= 0.2551 (0.1197, 0.3905);

[iv] p1= 34/61= 0.5574 (0.4332, 0.6750);

[v] p2= 28/37= 0.7568 (0.5990, 0.8664);

[vi] Cohen’s k: 0.1758 (0.009, 0.343).

Table 2.

Distribution of correct and incorrect answers for Physiology & Diagnostics question category (McNemar’s test: chi-squared = 2.5; df=1; p=0.1138; OR (95% CI) = 4 (0.798–38.666))

LLMGPT-4
GPT-3.5Correct answerYesNoRow total
Yes6 (28.6%)2 (9.5%)8 (38.1%)
No8 (38.1%)5 (23.8%)13 (61.9%)
Column total14 (66.7%)7 (33.3%)N = 21

[i] PA= (a+b)/N= 8/21= 0.381 (0.2075, 0.5912);

[ii] PB= (a+c)/N= 14/21= 0.6667 (0.4537, 0.8281);

[iii] PB - PA= 0.2857 (−0.0038, 0.5752);

[iv] p1= 8/13= 0.6154 (0.3552, 0.8229);

[v] p2= 6/8= 0.75 (0.4093, 0.9285);

[vi] Cohen’s k: 0.118 (−0.235, 0.471).

Table 3.

Distribution of correct and incorrect answers for Clinical & Case Questions question category (McNemar’s test: chi-squared = 4.083; df=1; p=0.0433; OR (95% CI)= 5 (1.066–46.933))

LLMGPT-4
GPT-3.5Correct answerYesNoRow total
Yes17 (39.5%)2 (4.7%)19 (44.2%)
No10 (23.3%)14 (32.6%)24 (55.8%)
Column total27 (62.8%)16 (37.2%)N = 43

[i] PA= (a+b)/N= 19/43= 0.4419 (0.3043, 0.5889);

[ii] PB= (a+c)/N= 27/43= 0.6279 (0.4786, 0.7562);

[iii] PB – PA= 0.186 (−0.0211, 0.3931);

[iv] p1= 10/24= 0.4167 (0.2447, 0.6117);

[v] p2= 17/19= 0.8947 (0.6860, 0.9706);

[vi] Cohen’s k: 0.458 (0.214, 0.702).

Table 4.

Distribution of correct and incorrect answers for Treatment & Pharmacology question category (McNemar’s test: chi-squared = 1.5; df=1; p=0.2207; OR (95% CI) = 5 (0.559–236.488))

LLMGPT-4
GPT-3.5Correct answerYesNoRow total
Yes2 (20%)1 (10%)3 (30%)
No5 (50%)2 (20%)7 (70%)
Column total7 (70%)3 (30%)N = 10

[i] PA= (a+b)/N= 3/10 = 0.3 (0.1078, 0.6032);

[ii] PB= (a+c)/N= 7/10 = 0.7 (0.3968, 0.8922);

[iii] PB – PA= 0.4 (−0.0017; 0.8017);

[iv] p1= 5/7= 0.7143 (0.3589, 0.9178);

[v] p2= 2/3= 0.6667 (0.2077, 0.9385);

[vi] Cohen’s k: −0.034 (−0.492, 0.423).

Table 5.

Distribution of correct and incorrect answers for Surgery question category (McNemar’s test: chi-squared = 1.125; df=1; p=0.2888; OR (95% CI) = 3 (0.536–30.393))

LLMGPT-4
GPT-3.5Correct answerYesNoRow total
Yes1 (7.7%)2 (15.4%)3 (23.1%)
No6 (46.2%)4 (30.8%)10 (76.9%)
Column total7 (53.9%)6 (46.2%)N = 13

[i] PA= (a+b)/N= 3/13= 0.2308 (0.0818, 0.5026);

[ii] PB= (a+c)/N= 7/13= 0.5385 (0.2914, 0.7679);

[iii] PB – PA= 0.3077 (−0.0471, 0.6625);

[iv] p1= 6/10= 0.6 (0.3127, 0.8318);

[v] p2= 1/3= 0.333 (0.0615, 0.7923);

[vi] Cohen’s k: −0.182 (−0.626, 0.262).

Table 6.

Distribution of correct and incorrect answers for Pediatrics question category (McNemar’s test: chi-squared = 0.571; df=1; p=0.4497; OR (95% CI) = 2.5 (0.409–26.253))

LLMGPT-4
GPT-3.5Correct answerYesNoRow total
Yes2 (18.2%)2 (18.2%)4 (36.4%)
No5 (45.5%)2 (18.2%)7 (63.6%)
Column total7 (63.6%)4 (36.4%)N = 11

[i] PA= (a+b)/N= 4/11= 0.3636 (0.1517, 0.6462);

[ii] PB= (a+c)/N= 7/11= 0.6364 (0.3538, 0.8483);

[iii] PB – PA= 0.2727 (−0.1293, 0.6748);

[iv] p1= 5/7= 0.7143 (0.3589, 0.9178);

[v] p2= 2/4= 0.5 (0.15, 0.85);

[vi] Cohen’s k: −0.185 (−0.708, 0.338)

Table 7.

Distribution of correct/false answers allocated for level of confidence for GPT3.5

Level of confidenceGPT-3.5
CorrectIncorrect
Definitely sure1823
Very sure1938
Almost sure--
Not very sure--
Definitely not sure--

[i] p3=(18/41)= 0.4390 (0.2989, 0.5896)

[ii] p4=(19/57)= 0.3333 (0.2249, 0.4628)

Table 8.

Distribution of correct/false answers allocated for level of confidence for GPT4

Level of confidenceGPT-4
CorrectIncorrect
Definitely sure81
Very sure4019
Almost sure1416
Not very sure--
Definitely not sure--

[i] p3=(8/9)= 0.8889 (0.5650, 0.9801)

[ii] p4=(54/89)= 0.6067 (0.5029, 0.7018)

Table 9.

Distribution of certainty levels between LLMs

LLMGPT-4
GPT-3.5Level of confidenceDefinitely sureVery sureAlmost sureNot very sureDefinitely not sureTotal
Definitely sure2 (2.04 %)18 (18.37%)21 (21.43%)--41 (41.84%)
Very sure7 (7.14%)41 (41.84%)9 (9.18%)--57 (58.16%)
Almost sure------
Not very sure------
Definitely not sure------
Total9 (9.18%)59 (60.20%)30 (30.61%)--N = 98
Language: English
Page range: 111 - 116
Submitted on: Jan 11, 2024
Accepted on: Jun 19, 2024
Published on: Sep 23, 2024
Published by: Hirszfeld Institute of Immunology and Experimental Therapy
In partnership with: Paradigm Publishing Services
Publication frequency: 1 issue per year

© 2024 Marcin Ciekalski, Maciej Laskowski, Agnieszka Koperczak, Maria Śmierciak, Sebastian Sirek, published by Hirszfeld Institute of Immunology and Experimental Therapy
This work is licensed under the Creative Commons Attribution-NonCommercial-NoDerivatives 4.0 License.