The findings in this study showed that, as we believed, the incidence of T2DM is much lower in the study population than the prevalence of T2DM in the general population. Military personnel is chosen according to pre-employment medical tests. In addition, the military lifestyle demands particular conditions, including regular physical activity, more mobility, a healthier dietary program, and periodic medical examination.
Previous studies have demonstrated increased T2DM risk associated with physical inactivity (
28-
30). As the results showed, the mean age and BMI were almost similar between officers and staff members. Therefore, the higher T2DM incidence in staff may confirm physical inactivity and sedentary behaviors in this group. Some studies have reported a stressful lifestyle as a risk factor for T2DM (
7,
31). We used rank as a marker for socioeconomic status. However, the lower prevalence of T2DM in conscripts is probably more related to their age and BMI circumstances than to their life status.
The risk of T2DM was categorized into modifiable and non-modifiable factors (
32). Among the variables in this study, only age was non-modifiable. Obesity is a well-established risk factor for T2DM (
33). The incidence of overweight or obesity in the study population was around 56.8%, and about 2.3% of T2DM individuals were overweight or obese. Kuwahara et al. (
11) suggested that preventing weight gain plays an important role in the reduction of T2DM risk. Diabetic dyslipidemia is an abnormal change in lipid profile as a consequence of T2DM (
34). Previous studies have illustrated how insulin resistance in T2DM patients causes high TG levels and decreased HDL cholesterol levels (
34-
36). T2DM individuals have elevated LDL cholesterol levels, but they may not have higher LDL levels (
37). Our findings of lipid profiles in T2DM patients are consistent with the reports of the aforementioned studies.
In this study, we evaluated three classification methods, and all of them suffered from the class-imbalanced problem. The use of SMOTE sampling technique was useful to resolve the imbalanced data problem.
Among classification methods, logistic regression has been extensively used in scientific research to measure the association between dependent and independent variables. Logistic regression is a parametric model and works based on a pre-determined set of variables. Therefore, its classification performance depends on the given model. Due to the intricate relationship among underlying features, this method may not have enough power to accurately classify subjects (
38).
By contrast, the decision tree is a non-parametric method mainly developed to classify the population rather to test the significance of variables on outcome (
39). However, the major drawback of the decision tree is moderate-to-high variance, which is an important cause of decision tree weak performance (
40).
RF is also a model-free classification technique, working based on an ensemble of decision trees. The salient feature of RF is low variance due to randomness grown of many trees (
41). Consequently, RF is less prone to overfitting and is better in generalization. In addition, RF provides a measure of variable importance, which is more informative than choosing a group of variables that their combination is predictive. Khalilia et al. (
13) showed that RF has superior performance compared to SVM, bagging, and boosting methods in disease prediction. Casanova et al. (
14) pointed out that the accuracy of RF in classification of diabetic retinopathy participants was much higher than the accuracy of logistic regression.
Typically, the performance of statistical models is assessed using predictive accuracy. However, study of diseases requires a relatively high rate of correct classification of patients. Our results confirm that RF is more powerful in finding complex relations among risk factors. Specifically, with regard to sensitivity and specificity, the RF more correctly classified cases.