Open-access A fair and efficient two-step procedure for sugarcane properties prediction based on near-infrared spectra

ABSTRACT

The growing demand for biofuels and sugar has prompted the selection of sugarcane cultivars with higher fiber content and apparent sucrose levels. In the initial phase of a breeding program, researchers might focus on classifying individuals as selected or not based on whether they display a property value above or below a specific threshold. Integrating classification methods with near-infrared (NIR) spectroscopy is essential for effective selection. This work outlines a two-step procedure designed to facilitate a fair comparison of the classifications under evaluation. We applied our approach to assess three classification techniques – Partial Least Squares Discriminant Analysis (PLS-DA), Support Vector Machines (SVM), and Random Forests (RF) – about their performance in predicting the classes of two sugarcane properties derived from NIR data. Our study utilized a dataset of 460 samples, categorized into high and low levels of fiber percentage (FIB) and apparent sucrose content (ASC), both of which are closely related to sucrose content. In the first step, the data was split into training and testing sets via the Kernard-Stone algorithm to determine the appropriate pretreatments for the NIR spectra. The selected pretreatment for each method was then applied in the second step, where the 460 samples were randomly divided again into training and test sets over ten repetitions. A key finding of this study is that the optimal set of pretreatments for a specific dataset varies depending on the classification method used. This two-step procedure streamlines the process for researchers to compare different classification techniques.

Keywords:
PLS-DA; machine learning; classification models; plant breeding

Introduction

Classification methods have been utilized with near-infrared (NIR) data to address classification challenges across various fields of study. In the area of sugarcane specifically, breeders aim to adopt statistical learning approaches to improve the efficacy of the breeding process (Moreira et al., 2021; Peternelli et al., 2017, 2018). This includes the application of NIR-base classification techniques (Peternelli et al., 2020; Porto et al., 2019; Sexton et al., 2020).

Regardless of the field of study and the application of NIR techniques in classification models, the information contained in NIR spectra is complex, self-correlated, and typically not available for analytical purposes without modification. Therefore, applying pretreatment to the spectra is essential to enhance the signal-to-noise ratio. The methods for these spectrum pretreatments vary across different studies (Ferreira, 2015; Pasquini, 2018).

In NIR studies that employ classification models, Partial Least Squares Discriminant Analysis (PLS-DA) is generally regarded as the preferred technique for analyzing NIR data (Pasquini, 2018). However, several other classification methods exist, each with unique characteristics. For instance, Support Vector Machines (SVM) demonstrate robustness when handling high-dimensional data and exhibit strong generalization capabilities (Moreira et al., 2021; Torres and Reverón, 2014). Another notable approach is the Random Forest (RF) method, which is versatile and adept at managing data with non-linear features and high dimensionality (Chemura et al., 2017; Moreira et al., 2021). The literature is rich with applications of NIR data analyses for classification purposes, detailing various aspects of data splitting for cross-validation, methods for comparing results, and approaches for drawing inferences about the comparisons of classification methods. Yet, based solely on these references, it remains unclear how researchers should effectively execute the process of comparing the classification methods in question, optimally split the available data, and apply statistical tests to make probabilistic inferences regarding the classification methods under comparison.

This study presents a two-step approach for conducting comparative analyses of the predictive performance of classification models that is both fair and easy to execute. We demonstrated our method by comparing the performance of three classification techniques – PLS-DA, SVM, and RF – using NIR data from a sugarcane experiment to classify clones based on two industrial quality characteristics: apparent sucrose and fiber content.

Materials and Methods

Reference analysis and NIR data

The study comprised 460 samples of crushed stalks from a population of clones. The apparent sucrose content (ASC) in the cane and the fiber percentage (FIB) in sugarcane were determined according to the recommendations outlined in the CONSECANA manual (CONSECANA, 2006).

Ten stalks from ten different clumps within a sugarcane family were randomly selected from each double-row plot. The stalks were cut at the soil level with a machete. Subsequently, the green tops, adherent leaves, and leaf sheaths were removed before the stalks were bunched and weighed using a dynamometer. The ten stalks were then crushed using a stationary forage chopper (model EN-6500, Nogueira & Brait company). A subsample of 500 g of the crushed stalks was obtained and subjected to pressure in a hydraulic press (PL011 Model, Dedini, Inc.) at 250 kg cm–2 (24.5 MPa) for 1 minute.

After pressing, the juice and the remaining fiber cake were collected and taken to the laboratory for analysis. The juice's percentage of sucrose (POL) in the sugarcane juice was measured using polarimetry with a saccharimeter (Model SDA2500), following the bleaching of the solution with lead acetate Pb(C2H3O2)2. The remaining fiber cake was then weighed (WC) to estimate its fiber content (CONSECANA, 2006) using the formula FIB=0.08×WC+0.876. Additionally, the apparent sucrose content was calculated with the equation ASC=POL×(1–0.01×FIB)× C, where C is the coefficient that converts juice sucrose into sugarcane sucrose. It is calculated using the formula C=1.0313–0.00575×FIB. The final values are expressed based on the total fresh biomass (500 g of crushed stalks).

For the NIR analyses, subsamples of 100 g of crushed stalks were collected and immediately dried in a forced air oven at 50 °C for 24 h or until a constant mass was reached. After drying, the subsamples were ground using a mill equipped with a 0.4 mm bottom sieve and then packed in a plastic bag for storage. NIR spectra were measured under controlled laboratory conditions, maintained at a room temperature of 21 °C. The instrument utilized was a Fourier Transform Near-Infrared (FT-NIR) spectrometer (model Antaris™ II, Thermo Scientific Inc.). The operating conditions of the instrument were set to a resolution of 4 cm–1 over a wavenumber range of 10000 to 4000 cm–1, using the diffuse reflectance mode with log (1 / R), where R represents the reflectance measurement. Each scan moved the accessory to cover various sample positions, totaling six positions. For each sample, 192 scans were conducted and averaged to produce the final spectrum.

Statistical analysis and comparison of methods

Initially, a descriptive analysis was conducted to categorize the data set into two classes based on the averages of FIB and ASC. Values above the mean of the characteristic of interest were designated as class 1, while those below the mean were classified as class 2. This classification method was selected for its simplicity. However, researchers may choose to compare classification methods without defining a threshold, as demonstrated by Peternelli and Andrade (2023), or select an alternative threshold that is more suitable for their specific application.

To enhance the analysis, graphs were constructed to visualize the spectra and assist in selecting the appropriate pretreatments. The pretreatments evaluated in this study included First Derivative (1D) and Second Derivative (2D), both of which employed the Savitzky-Golay procedure with a window size of 15 and a polynomial degree of 2. Additional methods considered were Multiplicative Scatter Correction (MSC) and Standard Normal Variate (SNV), along with their combinations, followed by Mean Centering (MC) (Ferreira, 2015). It is worth noting that other pretreatment options are available in the literature (Alsberg et al., 1997; Engel et al., 2013; Fearn, 2000; Ferreira, 2015), as mentioned in other references (Chong and O’Shea, 2013; Corrêdo et al., 2021; Phetpan et al., 2018; Phuphaphud et al., 2020; Sexton et al., 2020).

For practical purposes and to simplify the application of the procedure, the classification methods examined in this study include PLS-DA, SVM, and RF. Detailed information about these methods can be found in the specialized literature (Hastie et al., 2009; James et al., 2013) and will not be elaborated upon in this manuscript. The proposed procedure consists of 2 steps: Step 1 involves identifying the most effective pretreatment for each classification method, while Step 2 focuses on comparing these classification methods.

Procedure – Step 1

The data set was previously stratified using the Kennard-Stone algorithm (Kennard and Stone, 1969), resulting in two subsets: a training set containing 368 samples (80 %) and a testing set comprising 92 samples (20 %). This separation was conducted individually for ASC and FIB, ensuring that each characteristic maintained the same class proportion as the overall dataset. The primary objective of this algorithm is to select a subset of samples that best represents the maximum variability present in the entire dataset (Kennard and Stone, 1969), making it a popular choice in NIR research (Ferreira et al., 2022; Li et al., 2018). The Kennard-Stone algorithm ensures consistency by selecting the same training and testing samples for a given dataset whenever the same configuration is applied. In this context, it was utilized to identify the optimal pretreatment or their combination in the first step of the procedure. A cross-validation analysis was conducted on the training set to determine the model parameters for each set of pretreatments. The pretreatment or combination that yielded the lowest classification error on the testing set was considered the most suitable for the specific classification method to be employed in the second step.

Procedure – Step 2

Following the selection of pretreatments for each classification method (Step 1), the dataset, comprising 460 samples, was randomly partitioned into training (368 samples, representing 80 % of the dataset) and testing (92 samples, accounting for 20 % of the dataset) sets to compare the classification methods. Each instance of random sampling is referred to as a "repetition" in this study. In Step 2, a model was constructed for each repetition using the pretreatments defined in Step 1; thus, no evaluation of pretreatments was conducted in Step 2. A total of ten repetitions were performed, resulting in the creation of ten classification models. This approach enabled the establishment of confidence intervals and the application of statistical tests (Ferreira et al., 2022).

In all classification methods considered in this work (Steps 1 and 2), it is crucial to define the optimal values for their parameters. Specifically, In PLS-DA, the number of latent variables (nLV) is essential (Barker and Rayens, 2003); for SVM, we utilized the penalty parameter (C) and the Radial Basis kernel function; and for RF, the number of predictors (m) was determined (James et al., 2013). Other model parameters were set to their default values in the R package used for analysis. To aid in selecting parameter values, a k-fold cross-validation (k = 10) was employed using the training set samples. This approach divides the training set into ten equal parts, using one part for testing and the remaining nine parts for parameter estimation and model validation (Hastie et al., 2009). The parameters selected yielded the lowest estimated classification error during cross-validation.

An independent cross-validation was conducted for each testing set associated with the constructed models, following the methodology outlined by Ferreira et al. (2022). This process enabled the creation of a confusion matrix (Table 1), which displays the performance of each model in classifying the independent samples. Consequently, we calculated the average classification errors, sensitivity, and specificity. To compare the different methods, an analysis of variance (ANOVA) approach was employed.

Table 1
Representation of a confusion matrix between the actual classification and the classification based on the predicted values by the models.

The metrics used to evaluate the performance of each classification method are defined as follows (Altman and Bland, 1994a, b):

E r r o r = F P + F N T P + T N + F P + F N S e n s i t i v i t y = T P T P + F N S p e c i f i c i t y = T N T N + F P

The error metric refers to the classification error, indicating the percentage of samples that are misclassified. This metric is also known as the Apparent Error Rate (Chakraborty et al., 2021; Porto et al., 2019). The sensitivity metric measures the rate of true positives; that is, it assesses how often observations that belong to class h (h = 1 or 2) are correctly classified as class h. In contrast, the specificity metric evaluates the accurate negative rate, representing the proportion of observations that do not belong to class h and are correctly identified as such (Moreira et al., 2021).

Our application generated ten models for each classification method, corresponding to the ten repetitions conducted. We calculated the mean and standard deviation of the sensitivity, specificity, and error rate results derived from each classification method to compare their performances and identify the most effective method for classifying samples based on NIR data. As a result, we used statistical tests to compare the means and evaluate the classification methods.

Given that we conducted ten repetitions for each classification method, an ANOVA under a randomized complete block design (RCBD) was utilized to evaluate their performance. The blocking factor corresponds to a specific random selection that divides the data into training and testing sets, as established in Step 2. We selected the RCBD because we aimed to compare the classification methods while disregarding the influence of any random split of the spectra. An example of this approach can be found in the work by Ferreira et al. (2022), where the response variables included classification errors, sensitivity, and specificity produced by each classification method, with the blocks representing the ten independent repetitions performed (Ferreira et al., 2022). Following a significant result (p < α) from the analysis of variance, we can carry out multiple comparisons of the means of the classification methods when comparing more than two methods. For these multiple comparisons, the Student's t-test was employed at a significance level of 5 % (Steel et al., 1997), contingent upon obtaining a significant ANOVA result.

Analyses were conducted using MATLAB software (Matlab R2016a, 9.0, The MathWorks Inc.), PLS-Toolbox 8.2 (Eigenvector Research Inc.), and R (R Core Team, version 4.3.1). The function ‘kenSton’ from the R package ‘prospectr’ was employed for Kennard-Stone data splitting. For pretreatment analysis, we utilized a set of functions developed in our laboratory (Laboratory of Analysis and Research in Applied Statistics – LAPEA, www.lapea.ufv.br), which are accessible from the authors. R-base functions were applied for the remaining analysis.

The entire procedure is illustrated in a diagram (Figure 1).

Figure 1
A diagram referring to the two-step procedure performed. KS = Kennard-Stone algorithm; MC = Mean Centering; MSC = Multiplicative Scatter Correction; 1D = First Derivative; 2D = Second Derivative; SNV = Standard Normal Variate; PLS-DA = Partial Least Squares Discriminant Analysis; SVM = Support Vector Machines; RF = Random Forest; CV = cross-validation; nLV = number of latent variables; C = penalty parameter; m = number of predictors; TP = number of true positive cases; TN = number of actual negative cases; FP = number of false positive cases; FN = number of false negative cases; Error = Classification error; Sens = Sensitivity; Spc = Specificity.

Results

Descriptive data analysis

The ASC property exhibited a mean of 15.20 and a standard deviation of 1.83, while the FIB showed a mean of 13.66 with a standard deviation of 1.20. The fiber content values ranged from 10.53 % to 17.76 %, whereas apparent sucrose ranged from 9.31 % to 20.69 %. According to the established criteria for class definitions (threshold = property average), after separation using the Kenard-Stone algorithm, a total of 219 samples were classified into class 1 and 241 samples into class 2 for FIB. For ASC, 232 samples fell into class 1, and 228 samples were placed in class 2.

Spectra analysis

The sugarcane NIR spectra dataset used in this study comprises 460 samples and 3112 variables corresponding to the wavelengths obtained (Figure 2). To systematically select the optimal pretreatment, various NIR data pretreatments were initially evaluated using a training set generated through the Kennard-Stone algorithm to improve classification outcomes. The pretreatments tested included MSC, SNV, and both the First and Second Derivatives (Savitzky-Golay, window = 15 and polynomial degree = 2). The results for each pretreatment, along with the respective cross-validation errors (CV Error) for the PLS-DA, SVM, and RF methods in classifying FIB and ASC, are detailed in Table 2. A significant finding is that the most effective set of pretreatments applied to the spectra varies depending on the classification model used for the specific analysis. For ASC, the best PLS-DA model was achieved using a combination of MC, MSC, and Second Derivative pretreatments. Regarding the SVM method, the most suitable pretreatments were MC and MSC. The RF method demonstrated the lowest classification error when employing the centered on the mean and SNV pretreatments. For FIB, the optimal PLS-DA model was obtained with the combination of MC, MSC, and First Derivative pretreatments. Notably, the SVM method yielded a lower error without any pretreatments. The RF method showed improved results with the Center on Mean and Second Derivative pretreatments (Table 2).

Figure 2
Near-infrared spectra obtained from ground sugarcane bagasse to predict the contents of sucrose and fiber.
Table 2
Cross-validation error (CV Error) rates for different pretreatments in sugarcane near-infrared (NIR) data for classification of apparent sucrose percentage (ASC) and fiber percentage (FIB) using the Partial Least Squares Discriminant Analysis (PLS-DA), Support Vector Machines (SVM), and Random Forest (RF) models.

Upon analyzing the raw dataset, the error rates between PLS-DA and SVM are comparable, while RF exhibits an average increase in error rates of 26 %. For the ASC property, the optimal pretreatment for the SVM method results in a 9 % improvement in CV Error when no pretreatment is applied. Conversely, with the application of pretreatments, PLS-DA enhances the prediction accuracy of the ASC and FIB properties by 12 % and 13 %, respectively. RF also benefits from pretreatments, leading to a 15 % reduction in CV Errors for the ASC property and a 12 % reduction for the FIB property.

Adjustment of classification models

After selecting the appropriate pretreatments for each method, the original dataset of 460 samples was divided into training (368 samples) and testing (92 samples) sets. This process was repeated ten times, resulting in the random generation of ten classification models for each method. The individual results, including parameter values and corresponding CV Error rates for each method, are presented in Table 3.

Table 3
Parameter values and cross-validation error (CV Error) rates for the Partial Least Squares Discriminant Analysis (PLS-DA), Support Vector Machines (SVM), and Random Forest (RF) methods in the fitted models.

In each partition, the adjusted models were validated for classification by analyzing ASC and FIB, utilizing a testing set composed of 92 samples. This approach enabled the determination of each model's classification error, sensitivity, and specificity, allowing for a fair comparison of methods. Given that these models identify individuals with higher fiber and apparent sucrose contents, the classification parameters will be presented in relation to class 1. Since there are only two classes, the classification error values are equivalent, and the sensitivity values for class 2 are complementary to the specificity values for class 1 and vice versa.

There was variability among repetitions in the values of the classification parameters for each dataset used to build the models (Figure 3A-C).

Figure 3
A) Classification error (Error), B) Sensitivity, and C) Specificity for the Partial Least Squares Discriminant Analysis (PLS-DA), Support Vector Machines (SVM), and Random Forest (RF) models, referring to class 1 of apparent sucrose percentage (ASC) and fiber percentage (FIB) properties in the ten repetitions of the analysis. Each analysis corresponds to the original random training/testing sampling of the dataset. Results are shown based on the testing set.

On average, the RF method exhibits the highest error rates and lower sensitivity values for both variables. In contrast, the PLS-DA and the SVM demonstrate similar error, sensitivity, and specificity values across most repetitions (Table 4). To facilitate an exploratory analysis of these values, boxplots were constructed (Figure 4A-C), illustrating the dispersion of the evaluated measurements for each method.

Table 4
Mean and standard deviation (SD) values of the classification error (Error), sensitivity (Sens), and specificity (Spc) parameters obtained in each method for the classification of apparent sucrose percentage (ASC) and fiber percentage (FIB).
Figure 4
A) Boxplots of classification error, B) sensitivity, and C) specificity referring to class 1 for each classification method concerning the apparent sucrose percentage (ASC) and fiber percentage (FIB) properties. PLS-DA = Partial Least Squares Discriminant Analysis; SVM = Support Vector Machines; RF = Random Forest.

It is essential to note that the RF method frequently yields different outcomes compared to other methods, often resulting in lower classification metrics for both specificity and sensitivity. In contrast, the PLS-DA and SVM models do not demonstrate any significant difference in classification based on the ASC as opposed to the FIB.

Inferential comparison of classification methods

The ANOVA results comparing the three models (PSL-DA, SVM, and RF) reveal significant differences (p < 0.05) in mean classification error, sensitivity, and specificity regarding ASC. For FIB, the results also show significant differences (p < 0.05) for the mean classification error end sensitivity. For the pairwise comparisons, there is no significant difference (p > 0.05) in mean classification error and mean sensitivity between the PLS-DA and SVM methods; both differ from the RF (Figure 5), except for specificity concerning FIB. Additionally, PLS-DA and SVM demonstrate higher mean values for sensitivity and specificity metrics, along with lower classification error rates compared to RF.

Figure 5
Means and 95 % confidence intervals of each treatment and the Student's t-test for the evaluation metrics: classification error, sensitivity, and specificity for the apparent sucrose percentage (ASC) and fiber percentage (FIB). PLS-DA = Partial Least Squares Discriminant Analysis; SVM = Support Vector Machines; RF = Random Forest.

Discussion

Initial findings suggest that no specific pretreatment will universally suit all classification methods for both variables considered in this study; therefore, researchers should approach any analyses of this type with caution. In other words, the optimal choice of pretreatments that enhance classification metrics is contingent upon the data and the specific classification method employed (Sexton et al., 2020). The PLS-DA and RF methods demonstrated substantial improvements when additional pretreatment steps were applied. Conversely, the SVM method utilized only two pretreatments for ASC, while FIB data did not require any pretreatment. However, it is worth noting that PLS-DA exhibited a lower overall CV Error compared to RF. Another study evaluated the efficacy of pretreatments for SVM models in relation to other classification methods (Devos et al., 2014). The authors concluded that the precision improvement of the SVM-adjusted model, compared to PLS-DA, primarily stems from the non-linear adjustment characteristics of the SVM method, resulting in only a slight enhancement, concerning the use of pretreatments.

In the analysis comparing the three models, the assumptions of normality and homoscedasticity of the residuals were confirmed for all datasets (p > 0.05) (Steel et al., 1997; O’Neill and Mathews, 2000). However, in the context of sugarcane breeding, it is essential to place greater emphasis on classification error and sensitivity (Peternelli et al., 2017, 2018). These researchers argue that false positives are not a significant concern in plant selection because improperly selected plants can be discarded in the later stages of a breeding program.

A distinctive aspect of this work is the procedure we employed to ensure a fair comparison of the classifiers. We structured this process in two steps. In the first step, we partitioned the original dataset into training (368 samples) and testing (92 samples) sets using the Kennard-Stone algorithm (Kennard and Stone, 1969), which helps determine the most appropriate pretreatments for each classification method. This algorithm is considered the optimal approach for splitting data to enhance prediction accuracy (Ferreira et al., 2022). In the second step, the original dataset was randomly partitioned again into a training set (368 samples) and a testing set (92 samples), with ten repetitions. The pretreatments established in the first step were then applied. This two-step procedure enables a fair comparison of the classification methods in the second step, with each method optimized for its identified optimal pretreatments.

The classification results for ASC and FIB show that PLS-DA and SVM perform similarly (p > 0.05) and outperform RF in classifying the samples from the NIR data used in this study. This finding aligns with the conclusions of other researchers (Riccioli et al., 2018; Shao et al., 2015), who employed NIR and Electronic Nose data, respectively. In their studies, they found that the RF model was less effective than both PLS-DA and SVM.

Although the SVM and PLS-DA classifiers did not exhibit significant differences (p > 0.05) in the dataset analyzed in this study, numerous researchers have reported findings indicating a distinction between these two methods. A comparative analysis involving PLS-DA, SVM, and RF using NIR data revealed that the SVM exhibited higher sensitivity values compared to the other classifiers (Xu et al., 2017). Satisfactory results for identifying various types of rice flour using NIR spectroscopy with both the PLS-DA and SVM methods were reported by Sampaio et al. (2020). However, the authors emphasize that SVM demonstrates superior robustness over PLS-DA. Furthermore, a study comparing the PLS-DA and SVM methods, also utilizing NIR data, found that SVM had lower classification error rates than PLS-DA (Lu et al., 2014). In contrast, when PLS-DA, SVM, and RF methods were evaluated for their effectiveness in distinguishing spectrally similar tree species, the authors reported that PLS-DA outperformed both RF and SVM, achieving greater accuracy across all datasets (Richter et al., 2016).

Each classification method has unique characteristics that set it apart from the others. Support Vector Machines (SVM) demonstrate a strong capacity for generalization and adaptability in non-linear scenarios, exhibiting robustness even when dealing with large datasets. This robustness allows them to handle both minor and intentional variations in evaluation metrics (Wang et al., 2017; Zidi et al., 2018). Conversely, the PLS-DA approach effectively mitigates the impact of multicollinearity among variables (Liu et al., 2019) and is highly applicable to modeling high-dimensional data (Lee et al., 2018). The selection of a classification method ultimately hinges on the user's objectives and the nature of the data being analyzed. Both PLS-DA and the SVM have demonstrated high-performance applications in classifying NIR data samples across various fields, as evidenced in the studies by Gupta et al. (2018), Jianqiang et al. (2019), Porto et al. (2019), and Santana et al. (2020).

Furthermore, the selection of pretreatments applied to NIR spectra depends on the classification method employed. When comparing two or more classification methods, the proposed two-step procedure offers a sequential, straightforward, and fair approach. This method facilitates an inferential comparison that extends beyond merely assessing location measures. In this study, the PLS-DA and SVM methods demonstrated comparable and superior results relative to the RF method for classifying the percentage of apparent sucrose and fiber content. Nevertheless, all methods yielded satisfactory performance in terms of the classification metrics evaluated, establishing them as viable alternatives for refining NIR data classification models related to the apparent sucrose content and fiber percentage in sugarcane.

In summary, we determined that when conducting comparative studies of methods, the selection of NIR pretreatments should be tailored to each classification method individually. Furthermore, the proposed two-step procedure facilitates a fair comparison among the evaluated methods. In the present study, the PLS-DA and SVM methods demonstrated superior performance compared to the RF method in classifying the ASC and FIB traits using NIR spectra.

Declaration of use of AI Technologies

The authors declare that they did not use AI in analyzing or writing the manuscript.

Data availability statement

The scripts and a set of raw data needed to perform the main part of the analysis are available at https://doi.org/10.6084/m9.figshare.28886705.

Acknowledgments

The authors are thankful to Coordenação de Aperfeiçoamento de Pessoal de Nível Superior (CAPES) - finance code 001, Conselho Nacional de Desenvolvimento Científico e Tecnológico (CNPq) - processes 309662/2019-2 and 312316/2023-2, Financiadora de Estudos e Projetos (FINEP), and Fundação de Amparo à Pesquisa de Minas Gerais (FAPEMIG) for the financial support for research projects. The authors also thank the Rede Interuniversitária para o Desenvolvimento do setor Sucroenergético (RIDESA) for providing the data and constant financial support for the development of the breeding program.

References

  • Alsberg BK, Woodward AM, Kell DB. 1997. An introduction to wavelet transforms for chemometricians; a time-frequency approach. Chemometrics and Intelligent Laboratory Systems 37: 215-239. https://doi.org/10.1016/S0169-7439(97)00029-4
    » https://doi.org/10.1016/S0169-7439(97)00029-4
  • Altman DG, Bland JM. 1994a. Diagnostic tests 1: sensitivity and specificity. British Medical Journal 308: 1552. https://doi.org/10.1136/bmj.308.6943.1552
    » https://doi.org/10.1136/bmj.308.6943.1552
  • Altman DG, Bland JM. 1994b. Diagnostic tests 2: predictive values. British Medical Journal 309: 102. https://doi.org/10.1136/bmj.309.6947.102
    » https://doi.org/10.1136/bmj.309.6947.102
  • Barker M, Rayens W. 2003. Partial least squares for discrimination. Journal of Chemometrics 17: 166-173. https://doi.org/10.1002/cem.785
    » https://doi.org/10.1002/cem.785
  • Chakraborty SK, Mahanti NK, Mansuri SM, Tripathi MK, Kotwaliwale N, Jayas DS. 2021. Non-destructive classification and prediction of aflatoxin-B1 concentration in maize kernels using Vis–NIR (400-1000 nm) hyperspectral imaging. Journal of Food Science and Technology 58: 437-450. https://doi.org/10.1007/s13197-020-04552-w
    » https://doi.org/10.1007/s13197-020-04552-w
  • Chemura A, Mutanga O, Dube T. 2017. Separability of coffee leaf rust infection levels with machine learning methods at Sentinel-2 MSI spectral resolutions. Precision Agriculture 18: 859-881. https://doi.org/10.1007/s11119-016-9495-0
    » https://doi.org/10.1007/s11119-016-9495-0
  • Chong BF, O’Shea MG. 2013. Advancing energy cane cell wall digestibility screening by near-infrared spectroscopy. Applied Spectroscopy 67: 1160-1164. https://doi.org/10.1366/13-07003
    » https://doi.org/10.1366/13-07003
  • Conselho dos Produtores de Cana de Açúcar, Açúcar e Etanol do Estado de São Paulo [CONSECANA]. 2006. Manual de Instruções. 5ed. CONSECANA, Piracicaba, SP, Brazil (in Portuguese).
  • Corrêdo LP, Maldaner LF, Bazame HC, Molin JP. 2021. Evaluation of minimum preparation sampling strategies for sugarcane quality prediction by vis-NIR Spectroscopy. Sensors 21: 2195. https://doi.org/10.3390/s21062195
    » https://doi.org/10.3390/s21062195
  • Devos O, Downey G, Duponchel L. 2014. Simultaneous data pre-processing and SVM classification model selection based on a parallel genetic algorithm applied to spectroscopic data of olive oils. Food Chemistry 148: 124-130. https://doi.org/10.1016/j.foodchem.2013.10.020
    » https://doi.org/10.1016/j.foodchem.2013.10.020
  • Engel J, Gerretzen J, Szymańska E, Jansen JJ, Downey G, Blanchet L, et al. 2013. Breaking with trends in pre-processing? TrAC Trends in Analytical Chemistry 50: 96-106. https://doi.org/10.1016/j.trac.2013.04.015
    » https://doi.org/10.1016/j.trac.2013.04.015
  • Fearn T. 2000. On orthogonal signal correction. Chemometrics and Intelligent Laboratory Systems 50: 47-52. https://doi.org/10.1016/S0169-7439(99)00045-3
    » https://doi.org/10.1016/S0169-7439(99)00045-3
  • Ferreira MMC. 2015. Quimiometria: Conceitos, Métodos e Aplicações. Editora da UNICAMP, Campinas, SP, Brazil (in Portuguese).
  • Ferreira RA, Teixeira G, Peternelli, LA. 2022. Kennard-Stone method outperforms the Random Sampling in the selection of calibration samples in SNPs and NIR data. Ciência Rural 52: e20201072. https://doi.org/10.1590/0103-8478cr20201072
    » https://doi.org/10.1590/0103-8478cr20201072
  • Gupta O, Das AJ, Hellerstein J, Raskar R. 2018. Machine learning approaches for large scale classification of produce. Scientific Reports 8: 5226. https://doi.org/10.1038/s41598-018-23394-3
    » https://doi.org/10.1038/s41598-018-23394-3
  • Hastie T, Tibshirani R, Friedman J. 2009. The Elements of Statistical Learning: Data Mining, Inference, and Prediction. 2ed. Springer. New York, NY, USA.
  • James G, Witten D, Hastie T, Tibshirani R. 2013. An Introduction to Statistical Learning: With Applications in R. Springer, New York, NY, USA.
  • Jianqiang Z, Panpan Y, Weijuan L, Yanmei Y, Tianjun Y, Ying H, et al. 2019. Rapid and automatic classification of tobacco leaves using a hand-held DLP-based NIR spectroscopy device. Journal of the Brazilian Chemical Society 30: 1927-1932. https://doi.org/10.21577/0103-5053.20190105
    » https://doi.org/10.21577/0103-5053.20190105
  • Kennard RW, Stone LA. 1969. Computer Aided Design of Experiments. Technometrics 11: 137-148. https://doi.org/10.2307/1266770
    » https://doi.org/10.2307/1266770
  • Lee LC, Liong CY, Jemain AA. 2018. Partial least squares-discriminant analysis (PLS-DA) for classification of high-dimensional (HD) data: a review of contemporary practice strategies and knowledge gaps. The Analyst 143: 3526-3539. https://doi.org/10.1039/C8AN00599K
    » https://doi.org/10.1039/C8AN00599K
  • Li XY, Liu Y, Lv MR, Zou Y, Fan PP. 2018. Calibration transfer of soil total carbon and total nitrogen between two different types of soils based on visible-near-infrared reflectance spectroscopy. Journal of Spectroscopy 2018: 8513215. https://doi.org/10.1155/2018/8513215
    » https://doi.org/10.1155/2018/8513215
  • Liu YD, Xiao H, Xu H, Rao Y, Jiang X, Sun X. 2019. Visual discrimination of citrus HLB based on image features. Vibrational Spectroscopy 102: 103-111. https://doi.org/10.1016/j.vibspec.2019.04.001
    » https://doi.org/10.1016/j.vibspec.2019.04.001
  • Lu Y, Du C, Yu C, Zhou J. 2014. Classifying rapeseed varieties using Fourier transform infrared photoacoustic spectroscopy (FTIR-PAS). Computers and Electronics in Agriculture 107: 58-63. https://doi.org/10.1016/j.compag.2014.06.005
    » https://doi.org/10.1016/j.compag.2014.06.005
  • Moreira ÉFA, Barbosa MHP, Peternelli LA. 2021. Can statistical learning models make early selection among sugarcane families easier and still efficient? Crop Science 61: 456-465. https://doi.org/10.1002/csc2.20334
    » https://doi.org/10.1002/csc2.20334
  • O’Neill ME, Mathews K. 2000. A weighted least squares approach to Levene's test of homogeneity of variance. Australian and New Zealand Journal of Statistics 42: 81-100. https://doi.org/10.1111/1467-842X.00109
    » https://doi.org/10.1111/1467-842X.00109
  • Pasquini C. 2018. Near infrared spectroscopy: a mature analytical technique with new perspectives; a review. Analytica Chimica Acta 1026: 8-36. https://doi.org/10.1016/j.aca.2018.04.004
    » https://doi.org/10.1016/j.aca.2018.04.004
  • Peternelli LA, Moreira ÉFA, Nascimento M, Cruz CD. 2017. Artificial neural networks and linear discriminant analysis in early selection among sugarcane families. Crop Breeding and Applied Biotechnology 17: 299-305. http://dx.doi.org/10.1590/1984-70332017v17n4a46
    » http://dx.doi.org/10.1590/1984-70332017v17n4a46
  • Peternelli LA, Bernardes DP, Brasileiro BP, Barbosa MHP, Silva RHT. 2018. Decision trees as a tool to select sugarcane families. American Journal of Plant Sciences 9: 216-230. https://doi.org/10.4236/ajps.2018.92018
    » https://doi.org/10.4236/ajps.2018.92018
  • Peternelli LA, Gonçalves MTV, Fernandes JG, Brasileiro BP, Teófilo RF. 2020. Selection of sugarcane clones via multivariate models using near-infrared (NIR) spectroscopy data. Australian Journal of Crop Science 14: 889-896. https://doi.org/10.21475/ajcs.20.14.06.p2099
    » https://doi.org/10.21475/ajcs.20.14.06.p2099
  • Peternelli LA, Andrade ACB. 2023. Insights and protocols for discrimination of sugarcane clones by dissimilarity measures on RGB and NIR data. PLOS One 18: e0288508. https://doi.org/10.1371/journal.pone.0288508
    » https://doi.org/10.1371/journal.pone.0288508
  • Phetpan K, Udompetaikul V, Sirisomboon P. 2018. An online visible and near-infrared spectroscopic technique for the real-time evaluation of the soluble solids content of sugarcane billets on an elevator conveyor. Computers and Electronics in Agriculture 154: 460-466. https://doi.org/10.1016/j.compag.2018.09.033
    » https://doi.org/10.1016/j.compag.2018.09.033
  • Phuphaphud A, Saengprachatanarug K, Posom J, Maraphum K, Taira E. 2020. Non-destructive and rapid measurement of sugar content in growing cane stalks for breeding programmes using visible-near infrared spectroscopy. Biosystems Engineering 197: 76-90. https://doi.org/10.1016/j.biosystemseng.2020.06.012
    » https://doi.org/10.1016/j.biosystemseng.2020.06.012
  • Porto NA, Roque JV, Wartha CA, Cardoso W, Peternelli LA, Barbosa MHP, et al. 2019. Early prediction of sugarcane genotypes susceptible and resistant to Diatraea saccharalis using spectroscopies and classification techniques. Spectrochimica Acta - Part A: Molecular and Biomolecular Spectroscopy 218: 69-75. https://doi.org/10.1016/j.saa.2019.03.114
    » https://doi.org/10.1016/j.saa.2019.03.114
  • Riccioli C, Pérez-Marín D, Garrido-Varo A. 2018. Identifying animal species in NIR hyperspectral images of processed animal proteins (PAPs): comparison of multivariate techniques. Chemometrics and Intelligent Laboratory Systems 172: 139-149. https://doi.org/10.1016/j.chemolab.2017.12.003
    » https://doi.org/10.1016/j.chemolab.2017.12.003
  • Richter R, Reu B, Wirth C, Doktor D, Vohland M. 2016. The use of airborne hyperspectral data for tree species classification in a species-rich Central European forest area. International Journal of Applied Earth Observation and Geoinformation 52: 464-474. https://doi.org/10.1016/j.jag.2016.07.018
    » https://doi.org/10.1016/j.jag.2016.07.018
  • Sampaio PS, Castanho A, Almeida AS, Oliveira J, Brites C. 2020. Identification of rice flour types with near-infrared spectroscopy associated with PLS-DA and SVM methods. European Food Research and Technology 246: 527-537. https://doi.org/10.1007/s00217-019-03419-5
    » https://doi.org/10.1007/s00217-019-03419-5
  • Santana FB, Souza AM, Almeida MR, Breitkreitz MC, Filgueiras PR, Sena MM, et al. 2020. Didactic experiment of chemometrics for the classification of edible vegetable oils by fourier transform infrared spectroscopy and partial least squares discriminant analysis: a tutorial. Química Nova 43: 371-381 (in Portuguese, with abstract in English). https://doi.org/10.21577/0100-4042.20170480
    » https://doi.org/10.21577/0100-4042.20170480
  • Sexton J, Everingham Y, Donald D, Staunton S, White R. 2020. Investigating the identification of atypical sugarcane using NIR analysis of online mill data. Computers and Electronics in Agriculture 168: 105111. https://doi.org/10.1016/j.compag.2019.105111
    » https://doi.org/10.1016/j.compag.2019.105111
  • Shao X, Li H, Wang N, Zhang Q. 2015. Comparison of different classification methods for analyzing electronic nose data to characterize sesame oils and blends. Sensors 15: 26726-26742. https://doi.org/10.3390/s151026726
    » https://doi.org/10.3390/s151026726
  • Steel RGD, Torrie JH, Dickey DA. 1997. Principles and Procedures of Statistics: A Biometrical Approach. 3ed. McGraw-Hill, New York, NY, USA.
  • Torres A, Reverón J. 2014. Integration of rock physics, seismic inversion, and support vector machines for reservoir characterization in the Orinoco Oil Belt, Venezuela. The Leading Edge 33: 774-782. https://doi.org/10.1190/tle33070774.1
    » https://doi.org/10.1190/tle33070774.1
  • Wang H, Gu J, Wang S. 2017. An effective intrusion detection framework based on SVM with feature augmentation. Knowledge-Based Systems 136: 130-139. https://doi.org/10.1016/j.knosys.2017.09.014
    » https://doi.org/10.1016/j.knosys.2017.09.014
  • Xu JL, Riccioli C, Sun DW. 2017. Comparison of hyperspectral imaging and computer vision for automatic differentiation of organically and conventionally farmed salmon. Journal of Food Engineering 196: 170-182. https://doi.org/10.1016/j.jfoodeng.2016.10.021
    » https://doi.org/10.1016/j.jfoodeng.2016.10.021
  • Zidi S, Moulahi T, Alaya B. 2018. Fault detection in wireless sensor networks through SVM classifier. IEEE Sensors Journal 18: 340-347. https://doi.org/10.1109/JSEN.2017.2771226
    » https://doi.org/10.1109/JSEN.2017.2771226

*Corresponding author

<peternelli@ufv.br>

Conflict of interest

The authors have declared that there are no conflicts of interest.

Edited by:

Thomas Kumke

Publication Dates

  • Publication in this collection
    11 Aug 2025
  • Date of issue
    2025

History

  • Received
    13 Mar 2024
  • Accepted
    18 Feb 2025
location_on
Escola Superior de Agricultura "Luiz de Queiroz" USP/ESALQ - Scientia Agricola, Av. Pádua Dias, 11, 13418-900 Piracicaba SP Brazil, Phone: +55 19 3429-4401 / 3429-4486 - Piracicaba - SP - Brazil
E-mail: scientia@usp.br
rss_feed Acompanhe os números deste periódico no seu leitor de RSS
Ir para o topo Reportar erro