In the first step of the present study, gas chromatography (GC) analyses were applied on thirty eight different samples of heroin from three locations in Serbia. The data of gas chromatography (GC) analyses are summarized in
Table 1. as the peak area ratio of four secondary components, namely: acetyl thebaol (TEB), 6-monoacetyl morphine (MAM), papaverine (PAP), noscapine (NOS), and the main psychoactive component of diacetyl morphine (DAM). All these components were identified according to their retention times. In the second step, we focused our efforts on developing the MLR models that can determine the geographical origin of heroin samples. A set of twenty eight collected data (samples 1-28) was used for MLR modeling. TEB/DAM was used as a dependent variable in the regression analysis, and MAM/DAM, PAP/DAM and NOS/DAM were used as independent variables.
MLR procedure was used to model the relationships between the data of GC analyses. The stepwise regression (SWR) method was used to derive the most significant models as a calibration models for prediction of TEB/DAM peak ratio of seized heroin samples. The specifications for the best selected MLR models are shown in
Table 2.
The statistical quality of the generated models was checked by statistical parameters: correlation coefficient (
r), standard error of estimation or standard deviation (
s), and
F-test (Fisher›s value) for statistical significance (
16–
18). Correlation coefficient
r (or coefficient of multiple determination) is a relative measure of the fit by the regression equation. Correspondingly, it represents the part of variation in the observed data that is explained by the regression. Standard deviation is measured by the error mean square, which expresses the variation of the residuals or the variation from the regression line. Thus, standard deviation is an absolute measure of the quality of the fit and should have a low value for the regression to be significant. The
F-test reflects the ratio of the variance explained by the model and the variance due to the error in regression. High value of the
F-test indicates that the model is statistically significant.
It is well known that there are three important components in any chemometric-regression analysis: the development of the models, validation of the models and the utilization of developed models. Validation is a crucial aspect of any regression analysis (
19). For testing the validity of the predictive power of selected models leave one out (LOO) technique was used. The developed models were validated by the calculation of the following statistical parameters: PRESS, SSY, S
PRESS, r2CV, and
r2adj (
Table 3.). These parameters were calculated from the following equations:
where, Yobs, Ycalc and Ymean are observed, calculated and mean values; n is number of the samples and p is number of independent parameters.
PRESS is an acronym for prediction sum of squares. It is used to validate a regression model regarding to its predictability. To calculate PRESS, each observation is individually omitted. The remaining n-1 observations are used to calculate a regression and estimate the value of the omitted observation. This is done n times, once for each observation. The difference between the actual Y value, Yobs, and the predicted Ycalc, is so-called the prediction error. The sum of the squared prediction errors is the PRESS value. The smaller PRESS is, the better predictability of the model is achieved. SSY are the sums of squares associated with the corresponding sources of variation. These values are in terms of the dependent variable, Y.
The above PRESS value can be used to compute an r2CV statistic, called r2 cross validated parameter, which reflects the prediction ability of the model. This is a good way to validate the prediction of a regression model without selecting another sample or splitting the data. It is very possible to have a high r2 and a very low r2CV. When this occurs, it implies that the fitted model is data dependent. This parameter ranges from below zero to above one. When outside the range of 0-1, it is truncated to stay within this range. Adjusted r-squared (r2adj) is an adjusted version of r2. The adjustment seeks to remove the distortion due to a small sample size.
In many cases
r2CV and
r2adj are taken as a proof of the high predictive ability of MLR models. A high value of these statistical characteristics (>0.5) is considered as a proof of the high predictive ability of the model. However, some recent reports have proved the opposite (
20). Although, the low value of
r2CV for the training set can indeed serve as an indicator of a low predictive ability of a model, the opposite is not necessarily true. Thus, the high value of LOO
r2CV is the necessary condition for a model to have a high predictive power, but it is not a sufficient one.
Although models showed good internal consistency, they may not be applicable for the analogs which were never used in the generation of the correlation. It is proven that the only way to estimate the true predictive power of a model is to test it on a sufficiently large collection of the samples from an external test set. The test set must include no less than five samples, whose properties and structures must cover the range of properties and structures of the samples from the training set. This application is necessary for obtaining trustful statistics for comparison between the observed and predicted values for these compounds. Therefore, the external extrapolation power of the model was further authenticated by a test set of ten heroin samples.
The values of TEB/DAM peak ratio of an external set of heroin samples (samples 29-38) were calculated by the models. These data are compared with experimentally obtained values of TEB/DAM ratio (
Table 4.
Figure 1.). From the data presented in
Table 4. it is shown that high agreement between experimental and predicted TEB/DAM ratio was obtained (the residual values are small, indicating the good predictability of the established models). According to the reference (
16) without the validation of the MLR models by using the external test set, we could not come to a right conclusion about high predictive ability of derived models.
To investigate the existence of a systemic error in developing the MLR models, the residuals of predicted TEB/DAM peak ratio values were plotted against the experimental values in
Figure 2. The propagation of the residuals on both sides of zero indicates that no systematic error exists in the development of regression models as suggested by Jalali-Heravi
et al. (
21). It indicates that these models can be successfully applied to predict the geographic origin of seized heroin samples using the GC results. Therefore, the randomness of the residuals and their low values indicate that the obtained mathematical models can predict the dependent variable with acceptable error. According to the Variance Inflation Factor (VIF), which was lower than 10 for all the obtained models, it can be concluded that there is no multi collinearity present in the established models.
HCA was performed on the TEB/DAM, MAM/DAM, PAP/DAM and NOS/DAM peak ratios of the analysed heroin samples in order to reveal the similarities among them. Clustering was based on the Euclidean distance and single linkage algorithm. The obtained dendrogram is presented in
Figure 3. As it can be seen from the presented dendrogram, on the basis of TEB/DAM, MAM/DAM, PAP/DAM and NOS/DAM peak ratios, the most similar heroin samples come from border crossing (
2) and Novi Sad municipality (
3).These entities are placed into the separate cluster, while the samples that belong to border crossing (
1) are significantly different than the others.
WWR test was applied firstly on the TEB/DAM, MAM/DAM, PAP/DAM and NOS/DAM peak ratios together. The established null hypothesis “two samples come from populations having the same distribution” was confirmed for the samples that are seized at border crossing (
2) and Novi Sad municipality (
Table 5). This result confirms the finding obtained with HCA method. According to WWR test, the samples from Novi Sad municipality and border crossing (
1) do not belong to the same population, as well as the samples from border crossings (
1) and (
2).
Testing the H
0 hypothesis for the examined samples on the basis of TEB/DAM, MAM/DAM, PAP/DAM and NOS/DAM peak ratios separately, resulted in the same way as in previous WWR analysis, except in the case which included TEB/DAM peak ratio (
Table 6). In this case, all the three types of samples do not belong to the same population. It can indicate that exactly TEB/DAM peak ratio can be used as discriminating factor for the analysed samples. As it is shown in the MLR analysis, this ratio actually is dependent variable which is predicted based on the other determined peak ratios.