Sample collection and preparation
A total of 31 samples of lime fruit (
Citrus latifolia) originated from Jahrom city, IR. Iran were directly obtained from the local market of Tehran, IR Iran between April and December 2018. Lime fruit samples were gently squeezed by a manual citrus juicer (MCP 3500, Bosch, Germany), and homogenized using the Ultra-Turrax homogenizer (T8; IKA, Staufen, Germany). Twenty-five adulterated lime juice samples were kindly donated by the Iranian Food and Drug Administration. These samples were detected as adulterated samples based on citric acid to iso-citric acid ratio. The samples with a citric acid to iso-citric acid ratio over 300 were considered as non-genuine samples (
23). In these samples, adulteration was performed by the addition of water and subsequently citric acid as an acidifying agent.
Spectral collection (Portable NIRS)
A miniaturized research model NIRS device (Tellspec
®, Tellspec Inc., Toronto, Canada) connected to a smartphone was used in this study. Tellspec is equipped with two integrated halogen tungsten lamps and a single 1mm InGaAs detector on the same side which makes it able to operate as a diffuse reflectance NIRS. Exposure time, wavelength resolution, and accuracy were 0.635 ms, 12 nm and 2 nm, respectively (
15,
24). Three diffuse reflectance spectra at three random spots were acquired for each sample in the spectral range of 900–1700 nm (11,111-5,882 cm
-1) which included 256 points with 3 nm spectral steps. The averaged spectra of three acquired scans from each sample were subsequently used for fingerprinting and data elaboration.
Statistical analysis
Data preprocessing
Different preprocessing techniques including multiplicative scatter correction (MSC), standard normal variate (SNV), and 2nd-order derivative (2nd-Dv) were conducted on the whole spectra of genuine and adulterated juices before performing unsupervised and supervised algorithms. These preprocessing techniques were applied as they are the most widely used algorithms in NIRS in both reflectance and transmittance mode.
Principal component analysis
In order to visualize a description of the dataset, a multivariate statistical analysis was performed on the dataset, and different preprocessing techniques were conducted on the whole spectra of genuine and adulterated juices to find out which preprocessing technique could discriminate adulterated samples from genuine ones. PCA as a dimension-reduction tool was used to reduce the number of variables. PCA transforms the correlated variables into the uncorrelated variables called principal components (
25). To find out the variables which were more important in sample clustering, PC score plot was generated.
Partial Least Squares Discriminant Analysis
For sample clustering and making predictive models based on the state of adulteration, PLS-DA classifier was used to distinguish the different groups. PLS-DA builds regression models to correlate the information in the X block (
i.e., raw data) to binary Y variables (
i.e., groups, class membership,
etc.) by using the PLS algorithm (
26). This approach was utilized to maximize the covariance between the independent variables X and the corresponding dependent variable Y (
20). During model optimization, different preprocessing techniques were applied. The optimal number of factors also known as latent variables was selected based on the root mean square error of cross-validation (RMSECV) during cross-validation. In this case, RMSECV was plotted against the number of factors and the optimum number of factors that minimized the cross-validation error was selected.
k-nearest neighbors algorithm
K-NN as a pattern recognition technique was used for the classification of the samples. This algorithm attempts to categorize a new sample by computing the distance of that sample to all of the samples in the data matrix related to the training set (
27). The predicted class of an unknown sample depends on the class of its
k nearest neighbors. This model was applied following different data transforms and preprocessing methods mentioned before. During running the
k-NN model, the Euclidean distance that separates each pair of samples in the training set was calculated in the pirouette software. Following running the process, the optimal k value with the lowest validation error was selected.
Model validation
To evaluate the performance of generated models, internal and external validations were performed on two different data sets. For this purpose, the initial dataset was divided into two subsets of 70% and 30% by performing Kennard-stone algorithm. Forty uniformly distributed samples (22 genuine and 18 adulterated juices) were placed in the training set and 16 samples (9 genuine and 7 adulterated juices) were in the test set. By performing the data partitioning, the knowledge of training dataset did not affect the test dataset and the predictive power of the created model increased subsequently (
28). Leave-one-out cross-validation was applied on the training set for internal validation and the test set was used to externally validate the generated models. Data analysis was performed using Pirouette 4.5 software (Infometrix, Seattle, USA). A detailed workflow of data analysis is illustrated schematically in
Figure 1.
erated models several parameters including sensitivity, specificity, accuracy, and precision (Equations 1 to 4) were calculated. Matthew’s correlation coefficient (MCC) and kappa value were also compared across PLS-DA and k-NN models using the following equations (Equations 5 and 6). In equations 1 to 6, TP, TN, FP, FN, P
0, and P
e refer to true positive, true negative, false positive, false negative, the relative observed agreement among raters, and the hypothetical probability of chance agreement, respectively (
29,
30).
Sensitivity =
(Equation 1)
Specificity =
(Equation 2)
Accuracy =
(Equation 3)
Precision =
(Equation 4)
MCC =
(Equation 5)
Kappa =
(Equation 6)