Principal component analysis of the data set
Principal Component Analysis is a variable reduction procedure. Principal Components (PCs) are able to detect internal relations between characteristics of a set of objects, thus enabling a drastic reduction of the dimensionality of the original raw data. This reduction is achieved by transforming the original matrix to a new one, whose set of variables-termed PCs-appear to be orthogonal to each other (uncorrelated) and ordered so that the first few, with descending importance, retains most of the variance content from the total set of original variables (
32).
| Parameter* | Value |
|---|
| Population size | 30 chromosomes |
| Response | cross-validated% explained variance |
| Maximum number of variables selected in the same chromosome | 30 |
| Probability of mutation | 1% |
| Number of runs | 100 |
| Window size for smoothing | 3 |
| Number of compounds (Table 1) | Observation activity | PCR
| PLS
| GA-PLS
|
|---|
| Predicted | Error (%) | Predicted | Error (%) | Predicted | Error (%) |
|---|
| 11 | 5.00 | 5.23 | 4.60 | 5.19 | 3.80 | 5.06 | 1.20 |
| 18 | 5.10 | 5.36 | 5.09 | 5.29 | 3.72 | 5.12 | 0.39 |
| 21 | 5.60 | 5.03 | -10.17 | 5.11 | -8.75 | 5.54 | -1.07 |
| 23 | 5.00 | 4.23 | -15.40 | 4.39 | -12.20 | 4.96 | -0.80 |
| 26 | 8.30 | 8.68 | 4.57 | 8.51 | 2.53 | 8.33 | 0.36 |
| 33 | 7.85 | 8.06 | 2.67 | 7.94 | 1.14 | 7.81 | 0.51 |
| 40 | 4.37 | 4.01 | -8.23 | 4.16 | -4.80 | 4.32 | -1.14 |
| 53 | 8.24 | 8.75 | 6.19 | 8.34 | 1.21 | 8.26 | 0.24 |
| 63 | 5.68 | 5.93 | 4.40 | 5.88 | 3.52 | 5.72 | 0.70 |
| 68 | 6.66 | 6.86 | 3.00 | 6.84 | 2.70 | 6.81 | 2.25 |
| 71 | 7.11 | 6.56 | -7.73 | 6.72 | -5.48 | 7.14 | 0.42 |
| 74 | 8.13 | 8.69 | 6.88 | 8.57 | 5.41 | 8.09 | -0.49 |
| Methods | Data set | R2 | Q2* |
|---|
| PCR | Training | 0.7929 | 0.7812 |
| Test | 0.7822 | 0.7346 |
| PLS | Training | 0.8427 | 0.8109 |
| Test | 0.8126 | 0.8033 |
| GA-PLS | Training | 0.9412 | 0.9371 |
| Test | 0.9208 | 0.9124 |
| Model | R2 | Q2 | NF* | Reference |
|---|
| MLR | 0.900 | 0.745 | 9 | |
| | | | (34) |
| PLS | 0.889 | 0.860 | 9 | |
| MLR | 0.815 | 0.783 | 5 | |
| | | | (35) |
| MLR | 0.811 | 0.778 | 6 | |
| NN | 0.919 | 0.779 | 6 | |
| | | | (36) |
| MLR | 0.856 | 0.814 | 4 | |
| NN | 0.850 | 0.878 | 4 | |
| | | | (37) |
| SVM | 0.874 | 0.867 | 4 | |
| PCR | 0.793 | 0.781 | 7 | |
| PLS | 0.842 | 0.812 | 5 | This work |
| GA-PLS | 0.941 | 0.937 | 3 | |
2D images and unfolding step of the 107 chemical structures to give the X-matrix. The arrow in structure indicates the coordinate of a pixel in common among the whole series of compounds, used in the 2D alignment step
Principal components analysis of the 2D image descriptors for the data set, (a) PC2 versus PC1, (b) PC3 versus PC1 and (c) PC3 versus PC2
The RMSECV versus number of latent variables
Selected regions by genetic algorithms
Accordingly, PC1 is defined in the direction of maximum variation of the whole dataset. PC2 is the direction that describes the maximum variance in the orthogonal subspace to PC1. The PCA was performed with the calculated structure descriptors for the whole dataset to detect the homogeneities in the dataset, and also to show the spatial location of the samples to assist the separation of the data in the training and test sets. The PCA results showed that three principal components (PC1, PC2, and PC3) described 94.53% of the overall variables, as follows: PC1 = 46.24%, PC2 = 31.59% and PC3 = 16.76%. Most of the variance is accounted for in the 3 first PCs. Their score plot is a reliable presentation of the spatial distribution of the points in the dataset. As can be seen in
Figure 2, there is no clear clustering between compounds. The data separation is very important in the development of reliable and robust QSAR models. The quality of the prediction depends on the dataset used to develop the model. For regression analysis, the dataset was separated into two groups, a training set (91 data) and a prediction set (16 data), according to the Kennard-Stones algorithm. As shown in
Figure 2, the distribution of the compounds in each subset seems to be relatively well-balanced over the space of the principal components.
PCR and PLS modeling
The general purpose of the linear regression method is to quantify the relationship between several independent or predictor variables and a dependent variable. PLS is a linear modelling technique where the information in the descriptor matrix X is projected onto a small number of underlying (‘latent’) variables called PLS components or latent variables. The matrix Y is simultaneously used for the estimation of the ‘latent’ variables in X, which will be the most relevant for the Y variables prediction. Independent or predictor variables could cause pixel changes in descriptors of image of molecules, their principal components or latent variables. In multivariate calibration, such as PCR and PLS models, a predictive model can be obtained by selecting the optimum number of components using a cross-validation technique. In the cross-validation technique, one or more samples in the dataset are omitted, and the rederived PLS model is used to predict the biological activity of the omitted samples. This process is repeated until the biological activity of all samples in the dataset has been predicted once. The number of principal factors (nLV) of PLS is an important parameter in the modelling. The parameter is determined on the basis of assessing root mean square error of calibration (RMSEC) and root mean square error of cross validation (RMSECV). The number of PLS factors included in the model was chosen in accordance with the lowest RMSECV. As shown in
Figure 3, the RMSECV is minimized when the value of LVs is 7 and 5, and thus, the optimum LVs for the training set of PCR and PLS methods were respectively chosen to be 7 and 5. Prior to the PCR and PLS analysis, the dataset was mean-centred.
GA-PLS modeling
In the multivariate imaging analysis, the number of dependent variables is very large, so data reduction is very necessary. To find the more convenient set of descriptors in PLS modelling, genetic algorithms were used. The hybrid method that integrates GA as a powerful optimization tool and PLS as a robust statistical tool are applied to variable selection and modelling. After the running of GAs for pixel variables, the selected pixel descriptors were used for the running of PLS. When GA-PLS was used, the number of latent variables reduced to 3 (
Figure 3). In each feature selection method, the variables remaining after the exclusion of non-significant parameters were cross-correlated in order to select the most relevant parameters concerning the following criteria: 1)
p <0.05; 2) having the highest correlation with experimental data; and 3) having the lowest correlation with each other (33). The range of selected pixel descriptors is shown in
Figure 4. According to the descriptors selected by genetic algorithms, it was found that the maximum structural effects are in a, b, c, and d regions (
Figure 4). It seems that regions b, c, d—due to having the different functional groups—have a greater impact on the anti-HIV activity. This is because substituting O or S instead of X in the region a does not have a large impact on response. Selected areas in all the molecules are not identical in structure.
Model validation and prediction of anti-HIV activity
In
Table 3, the predicted values of activity obtained by the PCR, PLS and GA-PLS methods and the per cent relative errors of prediction are presented. The data observed and predicted activity for GA-PLS are distributed about a straight line with the corresponding slope and intercept equal to 0.9987 and 0.0085 respectively, which are nearly close to the perfect values: one and zero, correspondingly. The relative errors of prediction are between -1.14% and 2.25%. This was obtained by using the GA-PLS method, which shows the high-quality predictive capability of the developing QSAR model. The data presented in
Table 3 indicate that the GA-PLS model has good statistical quality with low prediction errors, while the GA-PLS model uses fewer latent variables.
Table 3 also shows RMSEP and RSEP to predict the activity of anti-HIV activity. Other statistical parameters have been to evaluate the suitability of the models developed for predicting the activity of the studied compounds, and this includes cross validation coefficient (Q
2 and R
2). An inspection of the results of the table reveals higher R
2 and Q
2 values and lower RMSCEV and RMSEP for the GA-PLS method compared with their counterparts. These results showed GA-PLS is significantly better than that of the other models. These parameters are listed in
Table 4, and show good statistical qualities.
The results were summarized and compared to the other models obtained by some works on the same set of HEPT derivatives in
Table 5. These results suggest the MIA-QSAR method is a useful tool, as promising as the most refined widely applied 2D methodologies, to correlate real pIC50 with pIC
50 provided by descriptors from modelled structures for this series of anti-HIV compounds. Also, this comparative table makes it clear that MIA is at least as predictive as these 2D refined methodologies, being, therefore, a much less expensive alternative to propose new HETP derivatives, since MIA-QSAR showed a Q
2 superior to all models available in the literature for this series of compounds.
Molecular design
As an application of the proposed method, we investigated GA-PLS model to predict the anti-HIV activity of five new HETP compounds on which biological tests were not performed yet.
Table 6 shows the chemical structure of five new HETP compounds and their activity calculated by this proposed method. According to GA-PLS model, we have found the new HEPT 5 molecules (
Table 6).