The CAD system is used to assist radiologists in breast mass discrimination. There have been many studies on the CAD system investigating the benefit of its output for radiologist diagnosis. Some studies have shown that CAD enhances the diagnostic performance of radiologists (
4,
5,
9-
13).
S-detect™ is a recently developed breast cancer CAD system using deep-learning algorithm with big data which providing assistance in morphological analysis by the ACR BI-RADS (
4,
5,
9-
12). Several studies have reported that S-detect™ could enhance the diagnostic performance of radiologists (
5,
6,
9-
12,
14). In addition, several studies have published that S-detect™ is a useful diagnostic tool that could be clinically used to enhance the specificity, PPV, and accuracy of breast US, regardless of the degree or radiologist’s experience (
9,
12). However, it is known that combining CAD with breast US is more useful than CAD alone (
9,
10). Our study aimed to determine the merits of the CAD system in breast US reading using breast phantom. In one previous study, five readers including residents retrospectively reviewed the CAD images and gave their assessments; inter-rater agreement was measured with Cohen’s kappa value (
11). The conclusion of that study was that S-detect™ is a feasible tool for characterization of breast lesions; it has a potential as a teaching tool for the less experienced operators. However, in that study, the residents only reviewed and assessed images taken by a radiologist with 32 years of experience in breast imaging (
11). In our study, we used breast phantom, not real patients. This is the greatest strength of our research. Junior residents who are trained for 2 hours can not perform ultrasound on real patients. We would like to demonstrate that breast phantom enables breast ultrasound education and training for starters.
First, we studied the reliability of CAD system for breast US. We used breast phantom, which enables multi-reader analysis for the same lesion. There have been several papers discussing breast phantom as a tool for breast US training (
15-
18), but this is the first study to directly apply breast phantom for reliability studies for breast CAD system. In our conclusion, there was better agreement of lexicons and final assessment in US than in CAD. The kappa values of the final decision on senior and junior groups were more variable on CAD than US, especially, in the junior group, there was greater inconsistency of CAD than US. Similar to the breast US, the inter- and intra-readers variability exists in CAD. In one previous study, moderate agreement (κ = 0.58) was seen in the final assessment between the CAD and dedicated breast radiologist (
5). The kappa value (κ = 0.44) between residents’ CAD result and dedicated breast radiologists in our study was lower than that (κ = 0.58) of the previous study. In order for CAD to be used properly as a dedicated breast radiologist, radiologists must get a proper US shot and then apply CAD. In our study, the CAD results for the junior group (beginner or starter) varied and were inconsistent. The kappa value of CAD was lower than that of the US. The statistical value was limited because of the small number of lesions in the breast phantom. However, the variability and inconsistency of the junior group were difficult to ignore. Therefore, we suggest that minimum training and experience for breast US is indispensable for better use of breast CAD.
In the subjective combined conclusion, the kappa value was improved in the junior group. When analyzing each lexicon, the kappa value of shape, orientation, and margin on US were significantly higher than those on CAD. The kappa value of echogenicity on CAD was higher only than that of US. So, we found that combining with breast US could improve the reliability in this study.
We also evaluated the diagnostic performance of the CAD system with breast US, by the junior and senior readers. The AUC was higher in CAD than US, while conjunctive combination result was the best. In addition, the diagnostic performance of CAD in the senior group was better than that of the junior group similar to US. We also found that combining CAD system with breast US could improve the diagnostic performance in this study. In one previous study, AUC improved for both the experienced and inexperienced readers (0.84 to 0.86 and 0.73 to 0.80) after the addition of CAD (
9). In our study, AUC improved for both senior and junior resident groups (0.779 to 0.917 and 0.756 to 0.846) after conjunctive combination. In another study, CAD was a useful additional diagnostic tool for breast US in all radiologists, with benefits differing depending on the radiologist’s level of experience. Compared with the experienced radiologists, the less experienced radiologists had significantly improved NPV (0.867 to 0.94 and 0.533 to 0.762) and AUC (0.823 to 0.839 and 0.623 to 0.759) with CAD assistance. In contrast, experienced radiologists had significantly improved specificity (0.525 to 0.542 and 0.661 to 0.661) and PPV (0.556 to 0.585 and 0.649 to 0.649) with CAD assistance. Interobserver variability of US features and final assessment by categories were significantly improved and moderate agreement was seen in the final assessment after CAD combination regardless of the radiologist’s experience (
10). In our study, combination of US with CAD improved the reliability and diagnostic performance, especially in the junior group.
There are limitations in our study. First, the data used in this study were derived from too few lesions (n = 14). The sample volume is very low which could decrease the accuracy of the study. This is probably the major cause of why data from our study did not yield statistically significant results. Specifically, the sample size is too small to show the significant difference of diagnostic performances between senior and junior groups. Since our study was based on breast phantom, the small number of lesions was inevitable. However, using breast phantom has several advantages. It can result in more reproducible results, it is objective, and studies can be repeated many times. In the future, various studies, especially the reliability test could be applied using various phantoms. Second, in our study, the CAD system did not analyze calcification because the number of lesions including calcification in our phantom was insufficient for reliability analysis. For the same reason, associated features such as duct change were not analyzed in detail. Third, since our study was a study using breast phantom, pathologic confirmation was not possible and the gold standard was reaffirmed by dedicated breast radiologists. Therefore, there is a limit in deriving the diagnostic performance from this. Finally, when the reader selected the representative image, which could be differed in CAD depends on the readers. When the reader identified the lesion and touched the center of the lesion in the monitor, a ROI was drawn along the border of the mass either automatically by the CAD program. Several drawn borders were presented on the screen, and the reader selected the most appropriate one. BI-RADS lexicon and final assessment classifications were automatically analyzed and displayed by the CAD system. However, the readers selected the representative US image, touched the center of the lesion, and selected the most appropriate CAD image. Any change in any of these steps could make a different effect on CAD result based on the readers.
In conclusion, the combination of US with CAD improved the reliability and diagnostic performance, especially in the junior group. As mentioned earlier, several studies have shown that CAD is useful for inexperienced radiologists. However, in our study, the junior group actually meant beginners, and the CAD results of the junior group (beginner or starter) were variable and inconsistent. Therefore, minimum training and experience for breast US is indispensable for the better use of breast CAD, and combination of US with CAD is useful for all readers.