1. Background
1.1. A General Overview of Large Language Models in Radiology
1.2. The Specific Role of Magnetic Resonance Imaging Acquisition Protocols and Turkish Society of Radiology 2018 Magnetic Resonance Imaging and Computed Tomography Acquisition Standards Guideline
1.3. The Research Gap and Study Objectives
2. Objectives
3. Materials and Methods
3.1. Study Design
3.2. Participant Radiologists
| Variables | Radiologists (name initials) | Radiology experience (y) | MRI experience (y) | Certification |
|---|---|---|---|---|
| JRR | JRR1 (S.E.E.); JRR2 (H.K.) | 2 | 1 | — |
| SRR | SRR1 (Y.Ö.); SRR2 (M.K.) | 4 | 3 | — |
| JR | JR1 (Y.C.G.); JR2 (T.C.); JR3 (E.Ç.) | 7 | 6 | Board-certified (EDiR) |
| SR | SR1 (S.D.); SR2 (R.S.Ö.); SR3 (A.Ö.); SR4 (H.G.H.Ç.) | 23 | 20 | — |
Abbreviations: MRI, magnetic resonance imaging; JRR, junior radiology resident; SRR, senior radiology resident; JR, junior radiologist; EDiR, European Diploma in Radiology; SR, senior radiologist.
3.3. Question Development and Validation
| Question type | Question number per section (n) | Sections | Total number of questions (n) | Purpose of question type |
|---|---|---|---|---|
| OEQs | 15 | 7 | 105 | Assess factual knowledge of MRI acquisition standards |
| CBQs | 15 | 7 | 105 | Identify single, key MRI sequence for a given clinical scenario |
Abbreviations: OEQs, open-ended questions; MRI, magnetic resonance imaging; CBQs, case-based questions.
3.4. Prompting and Model Input Procedures
3.5. Performance Evaluation
3.6. Statistical Analysis
4. Results
4.1. Open-ended Question Performance
| Variables | Mean Likert score (point) | 95% CI lower (point) | 95% CI upper (point) |
|---|---|---|---|
| Claude 3.5 Sonnet | 3.51 | 3.39 | 3.63 |
| ChatGPT-4o with canvas | 3.45 | 3.28 | 3.62 |
| ChatGPT-4o | 3.39 | 3.22 | 3.56 |
| ChatGPT-o1 | 3.25 | 3.08 | 3.42 |
| Claude 3 Opus | 3.24 | 3.08 | 3.4 |
| Mistral Large 2 | 3.25 | 3.12 | 3.38 |
| Llama 3.1 405B | 3.17 | 3.04 | 3.3 |
| Gemini 1.5 Pro | 3.13 | 2.99 | 3.27 |
| JRR1 | 1.95 | 1.73 | 2.17 |
| JRR2 | 1.91 | 1.65 | 2.17 |
| SRR1 | 2.36 | 2.13 | 2.59 |
| SRR2 | 2.36 | 2.13 | 2.59 |
| JR1 | 3.06 | 2.92 | 3.2 |
| JR2 | 3.07 | 2.93 | 3.21 |
| SR1 | 3.22 | 3.09 | 3.35 |
| SR2 | 3.23 | 3.1 | 3.36 |
Abbreviations: CI, confidence interval; JRR, junior radiology resident; SRR, senior radiology resident; JR, junior radiologist; SR, senior radiologist.
| Variables | Claude 3 Opus | Claude 3.5 Sonnet | ChatGPT-4o | Mistral Large 2 | ChatGPT-4o with canvas | Gemini 1.5 Pro | ChatGPT-o1 | Llama 3.1 405B | JRR-1 | JRR-2 | SRR-1 | SRR-2 | JR-1 | JR-2 | SR-1 | SR-2 |
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| Claude 3 Opus | - | 0.0020 | 0.0640 | 0.9070 | 0.0100 | 0.2470 | 0.9110 | 0.3400 | 0.0002 | 0.0002 | 0.0002 | 0.0002 | 0.0420 | 0.0400 | 0.8390 | 0.7590 |
| Claude 3.5 Sonnet | 0.0020 | - | 0.1580 | 0.0002 | 0.4550 | 0.0001 | 0.0040 | 0.0001 | 0.0001 | 0.0001 | 0.0001 | 0.0001 | 0.0001 | 0.0002 | 0.0002 | 0.0002 |
| ChatGPT-4o | 0.0640 | 0.1580 | - | 0.0910 | 0.3170 | 0.0050 | 0.0800 | 0.3200 | 0.0001 | 0.0001 | 0.0003 | 0.0003 | 0.0010 | 0.0010 | 0.4900 | 0.5200 |
| Mistral Large 2 | 0.9070 | 0.0002 | 0.0910 | - | 0.0230 | 0.1350 | 0.9590 | 0.2580 | 0.0001 | 0.0001 | 0.0002 | 0.0003 | 0.0340 | 0.0380 | 0.7230 | 0.7020 |
| ChatGPT-4o with canvas | 0.0100 | 0.4550 | 0.3170 | 0.0230 | - | 0.0010 | 0.0050 | 0.0010 | 0.0001 | 0.0001 | 0.0002 | 0.0002 | 0.0010 | 0.0010 | 0.0120 | 0.0110 |
| Gemini 1.5 Pro | 0.2470 | 0.0001 | 0.0050 | 0.1350 | 0.0010 | - | 0.3390 | 0.6150 | 0.0003 | 0.0002 | 0.0002 | 0.0001 | 0.3800 | 0.3110 | 0.2700 | 0.2340 |
| ChatGPT-o1 | 0.9110 | 0.0040 | 0.0800 | 0.9590 | 0.0050 | 0.3390 | - | 0.3790 | 0.0001 | 0.0001 | 0.0002 | 0.0003 | 0.0350 | 0.0310 | 0.7920 | 0.7000 |
| Llama 3.1 405B | 0.3400 | 0.0001 | 0.3200 | 0.2580 | 0.0010 | 0.6150 | 0.3790 | - | 0.0001 | 0.0001 | 0.0001 | 0.0001 | 0.1700 | 0.1230 | 0.5400 | 0.5080 |
| JRR-1 | 0.0002 | 0.0001 | 0.0001 | 0.0001 | 0.0001 | 0.0003 | 0.0001 | 0.0001 | - | 0.7510 | 0.0002 | 0.0003 | 0.0002 | 0.0002 | 0.0002 | 0.0002 |
| JRR-2 | 0.0002 | 0.0001 | 0.0001 | 0.0001 | 0.0001 | 0.0002 | 0.0001 | 0.0001 | 0.7510 | - | 0.0002 | 0.0002 | 0.0001 | 0.0002 | 0.0001 | 0.0001 |
| SRR-1 | 0.0002 | 0.0001 | 0.0003 | 0.0002 | 0.0002 | 0.0002 | 0.0003 | 0.0001 | 0.0003 | 0.0002 | - | 0.8210 | 0.0001 | 0.0002 | 0.0002 | 0.0002 |
| SRR-2 | 0.0002 | 0.0001 | 0.0003 | 0.0003 | 0.0002 | 0.0001 | 0.0003 | 0.0001 | 0.0003 | 0.0002 | 0.8210 | - | 0.0002 | 0.0002 | 0.0001 | 0.0001 |
| JR-1 | 0.0420 | 0.0002 | 0.0010 | 0.0340 | 0.0010 | 0.3800 | 0.0350 | 0.1700 | 0.0002 | 0.0001 | 0.0001 | 0.0002 | - | 0.8600 | 0.0002 | 0.0002 |
| JR-2 | 0.0400 | 0.0002 | 0.0010 | 0.0380 | 0.0010 | 0.3110 | 0.0310 | 0.1230 | 0.0002 | 0.0002 | 0.0002 | 0.0002 | 0.8600 | - | 0.0003 | 0.0002 |
| SR-1 | 0.8390 | 0.0002 | 0.4900 | 0.7230 | 0.0120 | 0.2700 | 0.7920 | 0.5400 | 0.0002 | 0.0001 | 0.0002 | 0.0001 | 0.0002 | 0.0003 | - | 0.6400 |
| SR-2 | 0.7590 | 0.0002 | 0.5200 | 0.7020 | 0.0110 | 0.2340 | 0.7000 | 0.5080 | 0.0002 | 0.0001 | 0.0002 | 0.0001 | 0.0002 | 0.0002 | 0.6400 | - |
Abbreviations: JRR, junior radiology resident; SRR, senior radiology resident; JR, junior radiologist; SR, senior radiologist.
a P-values are obtained from Wilcoxon test.
4.2. Case-based Question Accuracy
| Variables | CBQ accuracy (%) | 95% CI lower (%) | 95% CI upper (%) |
|---|---|---|---|
| Claude 3.5 Sonnet | 84 | 78 | 90 |
| SR2 | 90 | 84 | 96 |
| SR1 | 88 | 82 | 94 |
| ChatGPT-o1 | 69 | 61 | 77 |
| Llama 3.1 405B | 68 | 60 | 76 |
| JR1 | 69 | 61 | 77 |
| ChatGPT-4o | 65 | 57 | 73 |
| JR2 | 65 | 57 | 73 |
| Mistral Large 2 | 64 | 56 | 72 |
| ChatGPT-4o with canvas | 62 | 54 | 70 |
| SRR2 | 60 | 52 | 68 |
| SRR1 | 59 | 51 | 67 |
| Claude 3 Opus | 58 | 50 | 66 |
| JRR1 | 57 | 49 | 65 |
| Gemini 1.5 Pro | 56 | 48 | 64 |
| JRR2 | 50 | 42 | 58 |
Abbreviations: CBQ, case-based question; CI, confidence interval; SR, senior radiologist; JR, junior radiologist; SRR, senior radiology resident; JRR, junior radiology resident.
| Variables | Abdomen and pelvis | Brain | Cardiothoracic | Spinal | Head and neck | Musculoskeletal | Breast | P |
|---|---|---|---|---|---|---|---|---|
| Claude 3 Opus (CBQ) | 0.239 X2 | |||||||
| False | 5 (33.3) | 5 (33.3) | 5 (33.3) | 5 (33.3) | 5 (33.3) | 7 (46.7) | 5 (33.3) | |
| True | 10 (66.7) | 10 (66.7) | 10 (66.7) | 10 (66.7) | 10 (66.7) | 8 (53.3) | 10 (66.7) | |
| Claude 3 Opus Likert Score (OEQ) | 0.592 K | |||||||
| Mean ± SD | 3.07± 0.594 | 3.47± 0.516 | 3.07 ± 0.799 | 3.33 ± 0.617 | 3.40 ± 0.632 | 3.13 ± 0.990 | 3.20 ± 0.676 | |
| Median | 3 | 3 | 3 | 3 | 3 | 3 | 3 | |
| Claude 3.5 Sonnet (CBQ) | 0.710 X2 | |||||||
| False | 1 (6.7) | 1 (6.7) | 1 (6.7) | 3 (20.0) | 4 (26.7) | 2 (13.3) | 3 (20.0) | |
| True | 14 (93.3) | 14 (93.3) | 14 (93.3) | 12 (80.0) | 11 (73.3) | 13 (86.7) | 12 (80.0) | |
| Claude-3.5 Sonnet Likert Score (OEQ) | 0.670 K | |||||||
| Mean ± SD | 3.60 ± 0.507 | 3.60± 0.507 | 3.60 ± 0.507 | 3.67 ± 0.488 | 3.67 ± 0.488 | 3.47 ± 0.516 | 3.47 ± 0.516 | |
| Median | 4 | 4 | 4 | 4 | 4 | 3 | 3 | |
| ChatGPT-4o (CBQ) | 0.050 X2 | |||||||
| False | 2 (13.3) | 2 (13.3) | 2 (13.3) | 4 (26.7) | 9 (60.0) | 6 (40.0) | 7 (46.7) | |
| True | 13 (86.7) | 13 (86.7) | 13 (86.7) | 11 (73.3) | 6 (40.0) | 9 (60.0) | 8 (53.3) | |
| ChatGPT-4o Likert Score (OEQ) | 0.717 K | |||||||
| Mean ± SD | 3.47 ± 0.516 | 3.47± 0.516 | 3.53 ± 0.640 | 3.33 ± 0.724 | 3.47 ± 0.640 | 3.20 ± 1.014 | 3.13 ± 0.834 | |
| Median | 3 | 3 | 4 | 4 | 4 | 3 | 3 | |
| Mistral Large 2 (CBQ) | 0.210 X2 | |||||||
| False | 7 (46.7) | 2 (13.3) | 2 (13.3) | 6 (40.0) | 8 (53.3) | 6 (40.0) | 3 (20.0) | |
| True | 8 (53.3) | 13 (86.7) | 13 (86.7) | 9 (60.0) | 7 (46.7) | 9 (60.0) | 12 (80.0) | |
| Mistral Large 2 Likert Score (OEQ) | 0.423 K | |||||||
| Mean ± SD | 3.20 ± 0.561 | 3.20± 0.561 | 3.33 ± 0.617 | 3.13 ± 0.516 | 3.33 ± 0.617 | 3.47 ± 0.640 | 3.07 ± 0.458 | |
| Median | 3 | 3 | 3 | 3 | 3 | 4 | 3 | |
| ChatGPT-4o with canvas (CBQ) | 0.709 X2 | |||||||
| False | 4 (26.7) | 4 (26.7) | 4 (26.7) | 5 (33.3) | 6 (40.0) | 6 (40.0) | 7 (46.7) | |
| True | 11 (73.3) | 11 (73.3) | 11 (73.3) | 10 (66.7) | 9 (60.0) | 9 (60.0) | 8 (53.3) | |
| ChatGPT-4o with canvas Likert Score (OEQ) | 0.128 K | |||||||
| Mean ± SD | 3.20 ± 0.775 | 3.53± 0.640 | 3.80 ± 0.414 | 3.60 ± 0.737 | 3.53 ± 0.640 | 3.27 ± 1.033 | 3.20 ± 0.676 | |
| Median | 3 | 3 | 4 | 3 | 4 | 4 | 3 | |
| Gemini 1.5 Pro (CBQ) | 0.648 X2 | |||||||
| False | 5 (33.3) | 6 (40.0) | 6 (40.0) | 5 (33.3) | 6 (40.0) | 7 (46.7) | 6 (40.0) | |
| True | 10 (66.7) | 9 (60.0) | 9 (60.0) | 10 (66.7) | 9 (60.0) | 8 (53.3) | 9 (60.0) | |
| Gemini 1.5 Pro Likert Score (OEQ) | 0.761 K | |||||||
| Mean ± SD | 3.20 ± 0.561 | 2.93± 0.799 | 3.07 ± 0.594 | 3.27 ± 0.594 | 3.33 ± 0.617 | 3.07 ± 0.799 | 3.07 ± 0.704 | |
| Median | 3 | 3 | 3 | 3 | 3 | 3 | 3 | |
| ChatGPT-o1 (CBQ) | 0.948 X2 | |||||||
| False | 4 (26.7) | 5 (33.3) | 5 (33.3) | 5 (33.3) | 5 (33.3) | 3 (20.0) | 6 (40.0) | |
| True | 11 (73.3) | 10 (66.7) | 10 (66.7) | 10 (66.7) | 10 (66.7) | 12 (80.0) | 9 (60.0) | |
| ChatGPT-o1 Likert Score (OEQ) | 0.912 K | |||||||
| Mean ± SD | 3.27 ± 0.799 | 3.47± 0.516 | 3.20 ± 0.775 | 3.27 ± 0.594 | 3.27 ± 0.704 | 3.20 ± 1. 014 | 3.07 ± 0.799 | |
| Median | 3 | 3 | 3 | 3 | 3 | 3 | 3 | |
| Llama 3.1 405B (CBQ) | 0.459 X2 | |||||||
| False | 6 (40.0) | 5 (33.3) | 5 (33.3) | 3 (20.0) | 4 (26.7) | 5 (33.3) | 3 (20.0) | |
| True | 9 (60.0) | 10 (66.7) | 10 (66.7) | 12 (80.0) | 11 (73.3) | 10 (66.7) | 12 (80.0) | |
| Llama 3.1 405B Likert Score (OEQ) | 0.752 K | |||||||
| Mean ± SD | 3.20 ± 0.561 | 3.27± 0.458 | 3.13 ± 0.516 | 3.07 ± 0.594 | 3.00 ± 0.535 | 3.27 ± 0.594 | 3.27 ± 0.594 | |
| Median | 3 | 3 | 3 | 3 | 3 | 3 | 3 |
Abbreviations: X2,chi-square test; CBQ, case based questions; OEQ, open-ended questions; K, Kruskal-wallis test; SD, standart deviation.
a Values are expressed as No. (%), Mean ± SD, or Median.
| Variables | Claude 3 Opus | Claude 3.5 Sonnet | ChatGPT-4o | Mistral Large 2 | ChatGPT-4o with canvas | Gemini 1.5 Pro | ChatGPT-o1 | Llama 3.1405B | JRR-1 | JRR-2 | SRR-1 | SRR-2 | JR-1 | JR-2 | SR-1 | SR-2 |
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| Claude 3 Opus | - | 0.0001 | 0.0907 | 0.1650 | 0.1700 | 0.7280 | 0.0300 | 0.0270 | 0.5510 | 0.6710 | 0.2810 | 0.4100 | 0.0300 | 0.2810 | 0.0002 | 0.0002 |
| Claude 3.5 Sonnet | 0.0001 | - | 0.0020 | 0.0003 | 0.0001 | 0.0002 | 0.0150 | 0.0120 | 0.0001 | 0.0001 | 0.0001 | 0.0001 | 0.0001 | 0.0001 | 0.0002 | 0.0002 |
| ChatGPT-4o | 0.0970 | 0.0020 | - | 1 | 0.5130 | 0.1760 | 0.5960 | 0.6350 | 0.2650 | 0.0500 | 0.7600 | 0.5960 | 0.6430 | 0.7750 | 0.0001 | 0.0001 |
| Mistral Large 2 | 0.1650 | 0.0003 | 1 | - | 0.8620 | 0.2300 | 0.4860 | 0.4050 | 0.4610 | 0.0610 | 1 | 0.6880 | 0.4880 | 0.8880 | 0.0001 | 0.0001 |
| ChatGPT-4o with canvas | 0.1700 | 0.0001 | 0.5130 | 0.8620 | - | 0.3550 | 0.1940 | 0.2620 | 0.1510 | 0.0700 | 1 | 0.7850 | 0.2500 | 1 | 0.0002 | 0.0002 |
| Gemini 1.5 Pro | 0.7280 | 0.0002 | 0.1760 | 0.2300 | 0.3550 | - | 0.0310 | 0.0430 | 0.8770 | 0.3710 | 0.4960 | 0.6880 | 0.0990 | 0.4700 | 0.0001 | 0.0001 |
| ChatGPT-o1 | 0.0300 | 0.0150 | 0.5960 | 0.4860 | 0.1940 | 0.0310 | - | 1 | 0.1090 | 0.0070 | 0.3600 | 0.2220 | 1 | 0.3600 | 0.0030 | 0.0010 |
| Llama 3.1 405B | 0.0270 | 0.0120 | 0.6350 | 0.4050 | 0.2620 | 0.0430 | 1 | - | 0.1530 | 0.0030 | 0.3710 | 0.2720 | 1 | 0.4010 | 0.0010 | 0.0001 |
| JRR-1 | 0.5510 | 0.0001 | 0.2650 | 0.4610 | 0.1510 | 0.8770 | 0.1090 | 0.1530 | - | 0.3130 | 0.6260 | 0.8900 | 0.1090 | 0.6580 | 0.0001 | 0.0002 |
| JRR-2 | 0.6710 | 0.0001 | 0.0500 | 0.0610 | 0.0700 | 0.3710 | 0.0070 | 0.0030 | 0.3130 | - | 0.0800 | 0.1360 | 0.0140 | 0.1060 | 0.0002 | 0.0002 |
| SRR-1 | 0.2810 | 0.0001 | 0.7600 | 1 | 1 | 0.4960 | 0.3600 | 0.3710 | 0.6260 | 0.0800 | - | 0.8900 | 0.3370 | 1 | 0.0002 | 0.0002 |
| SRR-2 | 0.4100 | 0.0001 | 0.5960 | 0.6880 | 0.7850 | 0.6880 | 0.2220 | 0.2720 | 0.8900 | 0.1360 | 0.8900 | - | 0.2720 | 0.8740 | 0.0002 | 0.0002 |
| JR-1 | 0.0300 | 0.0001 | 0.6430 | 0.4880 | 0.2500 | 0.0990 | 1 | 1 | 0.1090 | 0.0140 | 0.3370 | 0.2720 | - | 0.4420 | 0.0002 | 0.0010 |
| JR-2 | 0.2810 | 0.0001 | 0.7750 | 0.8880 | 1 | 0.4700 | 0.3600 | 0.4010 | 0.6580 | 0.1060 | 1 | 0.8740 | 0.4420 | - | 0.0800 | 0.1360 |
| SR-1 | 0.0002 | 0.5410 | 0.0001 | 0.0001 | 0.0002 | 0.0001 | 0.0030 | 0.0010 | 0.0001 | 0.0002 | 0.0002 | 0.0002 | 0.0002 | 0.0800 | - | 0.8900 |
| SR-2 | 0.0002 | 0.0002 | 0.0001 | 0.0001 | 0.0002 | 0.0001 | 0.0010 | 0.0001 | 0.0002 | 0.0002 | 0.0002 | 0.0002 | 0.0010 | 0.1360 | 0.8900 | - |
Abbreviations: JRR, junior radiology resident; SRR, senior radiology resident; JR, junior radiologist; SR, senior radiologist.
a P-values are obtained from McNemar test.





