This systematic review aimed to assess the role of LUS in predicting surfactant therapy in premature neonates with RDS. The included studies were based on cohort studies of clinical trials where the bedside LUS was more typically used for evaluation, moving from an optional to a routine examination (
33). Most of the included articles chose the European guidelines to evaluate the use of surfactant after the occurrence of RDS and then used lung ultrasound as the investigative metric for evaluation. However, some essential parameters, such as blood gas analysis, CXR, and clinical signs, cannot be overlooked. As experienced clinicians, they combined all the data to assess the severity of RDS. We summarized these studies and found that the LUS was good at evaluating pulmonary changes in the illness and accurately predicted the timing of surfactant use. Certainly, this does not mean that the LUS is free of controversy, so we would like to discuss the advantages and limitations of the LUS in detail.
As shown in
Tables 1 and
2, they placed the LUS study in a tertiary NICU, but the ultrasound operation was not necessarily accomplished by the ultrasound physicians. Indeed, five studies (
11,
21,
26,
27,
29) chose to train their NICU physicians as ultrasound operators, and only Raimondi et al. (
8) chose ultrasound physicians. Neonatologists receive at least 1 - 3 months of theoretical and practical training before the study, undergo testing, and perform ultrasound examinations under the supervision of senior physicians. To obtain clear ultrasound images of the lungs, the operators choose high-frequency hockey-stick linear probes for study in neonates. Theoretically, all neonates < 36 weeks are eligible for the study, except those suffering from congenital disease. In practice, neonates < 34 weeks are more likely to have RDS, while RDS in neonates > 34 weeks is not exclusively caused by RDS, as it may be caused by wet lung or pneumonia. Moreover, neonates > 34 weeks have relatively mature alveolar development and do not necessarily require surfactant treatment (
34). Considering
Tables 1 and
2, it can be seen that the cut-offs for LUS are not consistent and may produce heterogeneity.
We further evaluated the recruited studies against the four domains (patient selection, index test, reference standard, flow, and timing) and analyzed the results using the QUADAS-2 tool (
Figure 2). Most patients received the same reference standard, allowing for the correct classification of the target disease. A single-blinded or double-blinded method was used to ensure the reliability of the test. A pre-January cohort study was used to ensure high article quality for the study type of articles. Overall, the seven articles selected were of high quality and data reliability. Only a few articles had minor problems with the selection of time intervals. For example, appropriate intervals between the gold standard and new methodological measurements were not reported, which may result in a high risk of bias. Our results show that prediction with LUS is best performed within 1 - 2 hours after birth or with CPAP support and not later than 24 hours because the LUS pattern in neonates can mutate, even so-called "black slips" (deterioration of the LUS grade) occur in the first hours of life. Complete clearance of airway fluid cannot be achieved in the first four hours. Premature neonates are more affected by this "black slip" and even seem to have fluid refill in the airways without end-expiratory pressure (
35). Also, the changes are most pronounced 1 - 3 hours compared to 5 - 10 minutes after birth, probably due to ventilator fatigue (
30). These findings help understand why most studies have chosen to perform the LUS two hours after birth around the LUS time to predict surfactant needs better.
Although the meta-analysis found good pooled sensitivity and specificity of 86% and 79%, respectively, with a diagnostic odds ratio of 45.99, reflecting the high diagnostic correctness of LUS, some parameters showed a high degree of heterogeneity, such as I
2 (specificity) of 93.6% and I
2 (positive likelihood ratio) of 95.5% (
Figure 3). These heterogeneities may be due to cut-off or non-cut-off effects. We analyzed the non-cut-off effect using meta-regression and Deeks’ funnel plots. We categorized the data into subgroups according to "year of publication," "cases," "study design," and "gestational age" and compared within-group differences to identify sources of heterogeneity (
Table 3). We found no significant differences in the meta-regressions for the subgroup analyses, except for a slight deviation in the number of cases (P = 0.04). Deeks’ funnel plot (
Figure 4) shows the publication bias of the articles, which may be due to the lack of negative reviews, as the selected articles all had positive feedback on LUS (
36,
37). The sources of heterogeneity in the non-cut-off effect are complex and include study populations, indicator tests, reference standards (
38,
39), and so on. As shown in
Figure 7, the articles by Brat11 on the LUS prediction of 100% sensitivity to surfactant may introduce heterogeneity.
To explore the accuracy of the LUS and sources of heterogeneity, we plotted summary receiver characteristic curves (
Figure 5) with an AUC of 94% and no "shoulder arm" feature (
40). However, it is clear from the articles that the cut-offs are inconsistent, and a possible cut-off effect may exist. First, the effect of gestational age on LUS was considered, and although previous meta-regression showed no significant difference in LUS between < 34 and > 34 weeks, it cannot be ruled out that there is no difference in cut-offs of 28 - 34, 26 - 28, or 34 - 37 weeks. De Martino et al. (
26) argued that the different cut-offs of < 28 weeks vs. > 28 weeks would impact the indications for substance use. Secondly, the disease severity may impact the cut-offs, as Vardar et al. (
27) found that a score of 4 - 6 was suitable for predicting surfactant use in mild patients, with a score of 10 or higher indicating severe lung changes and emphasized that the LUS was suitable for predicting the early detection of RDS, but not for predicting late complications of chronic lung disease. Apparently, two ranges of cut-offs can be seen, with three papers (
11,
21,
26) choosing 4 - 6 as the predictive value and four papers (
8,
27-
29) choosing 8 - 12 as the predictive value, respectively. Nevertheless, a score of 8 - 12 has not been shown to be superior to a score of 4 - 6, although some studies have argued that choosing cut-offs greater than 6 reaches the optimal range in terms of specificity and sensitivity (
31). In any case, their study was statistically based on their own data, and the optimal cut-offs can be reliably obtained from the summary receiver operating characteristic curve with no standardized range of cut-offs. In practice, though, mechanically ventilated patients have a higher mean airway pressure and may require a higher cut-off. In addition, the lungs of premature neonates are inherently immature, and the cut-offs for very premature neonates may be different from those for late premature neonates, making it inappropriate to use the same cut-offs for prediction. We compared two different cut-offs using forest plots, which showed a difference between scores of 4 - 6 and 8 - 12 (
Figure 6), and this inconsistency in cut-offs led to a cut-off effect. Uncertainty in the range of cut-offs can greatly affect the objectivity of LUS, and therefore, reasonable cut-offs should be established for different cases. For example, different cut-offs are used to predict the likelihood of RDS according to different gestational ages and lesion degrees. Considering all data, it is concluded that cut-offs of 4 - 6 scores apply to premature neonates at 34 weeks, whereas cut-offs of 8 - 12 scores apply to premature neonates around 28 weeks or those with disease progression, suggesting the need for mechanical ventilation.
Since the LUS is a quantitative indicator of the diagnostic description of lung ultrasound, many authors have questioned its applicability and accuracy (
41). Nevertheless, the sensitivity and specificity of the LUS for predicting the onset and progression of RDS remain high when analyzed from the study data. The LUS can help neonatologists predict disease early and guide the use of surfactants and mechanical ventilation. The LUS is currently used in RDS, transient tachypnea of the newborn, pneumonia and chronic lung disease, etc. However, the disadvantage of the LUS is that it only provides a numerical quantitative assessment of the severity of lung lesions. It is difficult to accurately assess patients with severe and complex lesions, such as pulmonary hemorrhage, pneumothorax, and hernias; therefore, the original images should be kept for analysis and combined with CXR or even computed tomography (CT) if necessary (
42,
43). Additionally, evaluating lung diseases by ultrasound is closely related to the ultrasound equipment and the operator’s experience. For example, when using ultrasound to probe lung regions, the B-line may show hypertrophy inhomogeneity at different frequencies or always prefer open harmonics, affecting the imaging of both the B-line and the A-line (
44,
45). Improper selection of ultrasound probes can also affect the scoring accuracy seriously, and it is important to understand how to select convex, micro-convex, and sector probes.