Two reviewers (M.T. and S.H.Y.) independently assessed the methodological quality and risk of bias of all 27 included studies; disagreements were resolved by consensus, with a third reviewer available for arbitration. Two complementary instruments were applied. Risk of bias and applicability were evaluated using QUADAS-2 (
8) across its four domains—patient selection, index test (the deep learning classifier), reference standard, and flow and timing—with each domain rated as low, high, or unclear. Reporting completeness and adherence to good practice for AI diagnostic studies were assessed using the Checklist for Artificial Intelligence in Medical Imaging (CLAIM) (
9). The appraisal followed the CLAIM 2024 update (
10), which is now the reference standard for AI medical-imaging studies, and the core items scored here are consistent with it. For each study, we recorded adherence to core CLAIM items covering data source and eligibility, definition and adequacy of the reference standard, data partitioning, model and training description, performance metrics and their reported uncertainty, internal and external validation, and comparison with clinicians, and expressed adherence as the percentage of applicable items met. Because most included studies were diagnostic prediction-model studies rather than classical diagnostic-accuracy studies, the QUADAS-2 signalling questions were interpreted with attention to machine-learning-specific sources of bias, particularly data leakage arising from image- or slice-level rather than patient-level partitioning and the presence or absence of a genuinely held-out or external test set. For the same reason, risk of bias was additionally considered against the domains of PROBAST (
11), the instrument designed for prediction-model studies, which reinforced the patient-selection and analysis concerns identified by QUADAS-2. Domain-level judgements for every study are reported in Tables A2 in the Supplementary File (QUADAS-2) and A3 (CLAIM).