Open-access Implications of removing parameters from a mathematical model on relationships among traits in assays with growth-promoting bacteria in soybean crop

Abstract

The objetive assess the implications of removing parameter effects from a mathematical model on linear relationships in trials with growth-promoting bacteria in soybean crops. The design used consisted of randomized blocks, with four replications. First factor: i) Bradyrhizobium spp. + Azospirillum spp.; ii) Bradyrhizobium spp. + Pseudomonas fluorescens; ii) Bradyrhizobium spp. + Bacillus subtilis; iv) Bradyrhizobium spp. + Bacillus subtilis + Bacillus megaterium; v) Bradyrhizobium spp. + Azospirillum spp. + Pseudomonas fluorescens; vi) Bradyrhizobium spp. + Azospirillum spp. + Bacillus subtilis; vii) Bradyrhizobium spp. + Azospirillum spp. + Bacillus subtilis + Bacillus megaterium; viii) Bradyrhizobium spp. + Azospirillum spp. + Pseudomonas fluorescens + Bacillus subtilis + Bacillus megaterium; ix) no bacteria; and doses of phosphorus: 0, 50, 100 and 150 Kg ha-1. Parameters were removed and compared to the general model in the path and canonical correlation analyses. In the path analysis, there was a change in magnitude from 7.69% to 61.64%, and from 23.08% to 69.23%, concerning the direction of the direct effects. As for indirect effects, magnitude changed from 30.13% to 51.28%, and 33.97% to 55.77%, concerning the direction of the coefficients. In the canonical correlations, the explanatory power increased by up to 338.89% and 111.90%.

Key words
Glycine max L; Path analysis; Multicollinearity; Canonical correlations; Grain yield; Phosphorus

INTRODUCTION

Soybean (Glycine max L.), one of the main oilseeds grown in the world, is of great importance for industry and export, as well as widely used in human food and animal nutrition. Worldwide, 138.52 million hectares are used for soybean farming, with a production of 395.91 million tons of grains (Fao 2024). In Brazil, 147.35 million tons of soybeans are produced and 88.6 million tons of grains are exported, from a sown area of 45.98 million hectares, with an average production of 3,205 kg ha-1 (Conab 2024).

For high production ceilings, essential elements are necessary, such as water, temperature, solar radiation and adequate nutritional demand. In this context, phosphorus (P) stands out as one of the essential nutrients for the growth and development of crops, with much of it not being available in the soil for use by crops. Therefore, it is important to seek alternatives to increase the available phosphorus, such as using nutrient fixing and solubilizing microorganisms (Pavinato et al. 2020).

Phosphate solubilizing bacteria (PSB) play a vital role in converting insoluble phosphorus into plant-accessible forms, thereby improving nutrient uptake. They use phosphate solubilization mechanisms, such as production of organic acid, which is produced by Bacillus and Pseudomonas species and decreases soil pH, facilitating the solubilization of fixed phosphorus compounds (Iqbal et al. 2024). Solubilizing bacteria increase soil exoenzymatic activities, such as those involving phosphatases, which further aid in the mineralization of organic phosphorus (Cheng et al. 2023). Likewise, microbial interactions occur, which influence the structures of the microbial community, promoting synergistic relations that increase the availability of phosphorus (Pan & Cai 2023).

Bradyrhizobium japonicum is a prominent nitrogen fixing bacterium that forms symbiotic relationships with soybean roots, significantly increasing nitrogen availability and meeting the crop’s need for the biological nitrogen fixation (BNF) process (Faria et al. 2022). Inoculation with this bacterium, together with phosphorus solubilizing bacteria, increases the yield and quality of soybeans, with studies reporting rises in protein and oil content in seeds (Shome et al. 2022). The Pseudomonas fluorescens and Bacillus subtilis bacteria also have the ability to solubilize phosphate, increasing nutrient absorption and crop yield. However, the effectiveness of phosphate solubilizing bacteria can vary depending on soil conditions, with some strains showing limited impact on soybean growth in soils with low P content (Silvestrini et al. 2023).

In order for interactions between nutrient fixing and solubilizing bacteria and the crop to be ideal, it is necessary to understand the relationships between agronomic traits and their impact on grain yield after the use of these bioproducts. Therefore, it is essential to adopt statistical techniques that allow for identifying characteristics that can be directed towards genetic improvement and contribute to the selection of materials that present positive and facilitated interactions with these bacteria. Among the statistical techniques, path analysis and canonical correlation analysis can be highlighted.

Canonical correlation analysis (CCA) is a multivariate statistical technique, widely used to verify the association between groups of variables, in order to identify linear combinations that maximize the correlations between the two groups, in which the linear combinations are the canonical variables, and the existing associations are called canonical correlations (Hotelling 1936). In the literature, reports of the use of CCA are found for dual-purpose wheat, aiming to identify canonical correlations between morphological traits and components of genotype production, under different cutting managements in which the correlation depends on the management; the diameter of the main stem, the diameter of the tillers, the total number of tillers, as well as of the fertile ones, should be prioritized when it comes to grain yield and hectoliter weight in the selection of superior genotypes (Carvalho et al. 2015).

With respect to habanero pepper, to analyze the relationship among plant characteristics, such as root length and aerial part mass, the importance of canonical cross-loadings for reliable interpretations is emphasized (Diel et al. 2020). As for wheat, canonical correlations identified relationships between high and low heritability characteristics, suggesting that certain traits can be used for indirect selection of superior genotypes, which is crucial for improving yield (Silva et al. 2022). Concerning elephant grass and millet hybrids, it has been shown that morphoagronomic characteristics correlate with chemical ones, providing information for breeding programs that aim to improve forage quality (Daher et al. 2018).

Path analysis stands out as a consolidated method used in several areas for indirect selection and in breeding programs, which aim to assess the logical consequences of a hypothesis of cause-and-effect relationship as to associations between traits, decomposing these relationships into direct and indirect effects on a variable considered to be the main one (Wright 1923). Parameter estimates may be biased due to the complexity of the data, due to the main variable and a large number of explanatory variables, with a high degree of association between them, making it difficult to estimate the relationships of each explanatory variable, resulting in the so-called multicollinearity (Blalock 1963, Graham 2003). Research on corn crops has discovered that multicollinearity impacts the estimation of direct effects in path analyses, followed by multivariate non-normality (Toebe & Filho 2013).

Through path analysis, in a study with nitrogen application associated with inoculation with Bradyrhizobium japonicum, leaf area, thousand-grain weight and total dry weight were the characteristics with a direct effect on grain yield and protein content in the grain (Zuffo et al. 2020). In soybean cultivation (second harvest) in a subtropical climate, plant height was found to have a positive correlation with mean grain yield and is considered a characteristic to be studied in indirect selection, in a search for highly productive cultivars (Follmann et al. 2017). Through path analysis in soybeans, it was observed that an increase in respiratory activity and in seed percentage without mechanical damage indirectly contribute to raising the speed and emergence of soybean plants in the field (Wendt et al. 2017).

In multivariate analyses, mathematical model parameters relating to the experimental design are not considered, with average observations being used in general, without stratification of the effects of the parameters relating to the experimental design, factors under study and possible interactions; thus, maintaining the parameters of the mathematical model may lead to redundant conclusions (Sgarbossa et al. 2024b). By removing parameters from the mathematical model, stratifying the effects, one prevents possible results that may not evidence the true relationship among the variables studied, thus removing the influences of the treatments and of the experimental design used under the observations of the experimental unit.

Considering the scarcity of information on the subject and the importance of multivariate analyses such as path analysis and canonical correlation analysis, the objective of this study was to assess the implications of removing parameter effects from the mathematical model on linear relationships between variables, in path analysis and canonical correlation analysis, during a trial with growth-promoting bacteria in soybean crops.

MATERIALS AND METHODS

Study site and experimental design

The study used data from a field experiment conducted in the 2019/2020 harvest, in the municipality of Barra do Quaraí, RS, Brazil, under geographic coordinates 30°09’50.26” S, 57°30’21.40’’ W, 53 m above sea level. The region’s climate is type Cfa, according to the Köppen classification, characterized by an average air temperature of 20.1 °C to 21.0 °C, ranging from 0 to 38 °C, and an average annual precipitation of 1,240 mm (Tapiador et al. 2019). The soil of the experimental area is classified as Solodic Eutrophic Hydromorphic Planosol (Santos et al. 2018), with 30% clay; pH in H2O 4.6; Phosphorus (P) 1.5 mg L-1; Potassium (K) 102 mg L-1; Organic matter 1.3%; Calcium (Ca) 5.3 cmol L-1; Magnesium (Mg) 1.5 cmol L-1; Potential acidity (H + Al) 10.2 cmol L-1; Sulfur (S) 13.6 mg L-1; Cation exchange capacity 14.0 cmol L-1 and 50.6% sum of bases. The soil was fertilized in accordance with a soil chemical analysis and technical recommendations for the crop (CQFS 2016).

The experimental design used was of the randomized complete block type – a 9×4 bifactorial design with four replications, characterized by nine co-inoculation treatments with growth-promoting and phosphorus solubilizing bacteria: i) Bradyrhizobium spp. + Azospirillum spp.; ii) Bradyrhizobium spp. + Pseudomonas fluorescens; ii) Bradyrhizobium spp. + Bacillus subtilis; iv) Bradyrhizobium spp. + Bacillus subtilis + Bacillus megaterium; v) Bradyrhizobium spp. + Azospirillum spp. + Pseudomonas fluorescens; vi) Bradyrhizobium spp. + Azospirillum spp. + Bacillus subtilis; vii) Bradyrhizobium spp. + Azospirillum spp. + Bacillus subtilis + Bacillus megaterium; viii) Bradyrhizobium spp. + Azospirillum spp. + Pseudomonas fluorescens + Bacillus subtilis + Bacillus megaterium; and ix) Control (no use of bacteria); and four doses of phosphorus (P2O5): 0, 50, 100 and 150 kg ha-1, using the triple superphosphate (TSP) mineral fertilizer, with a concentration of 46% P2O5.

Sowing was carried out on November 11, 2019, using cultivar DM 66i68 IPRO, recommended for lowland environments (areas prone to flooding), with the aid of a mechanized seeder, with spacing between planting rows of 0.57 m and a population of 300,000 plants per hectare. The plots were delimited to be 5.7 m wide and 5 m long, totaling a useful area of 28.5 m². Inoculation was performed in the furrow at the time of sowing, using a volume of 50 L ha-1, with the liquid inoculant at the recommended doses of bacteria and in accordance with the defined treatments. The other management practices were carried out following the recommendations and technical specifications on soybean farming (Caraffa et al. 2019).

The variables measured were: number of nodules (NN), obtained by counting 12 plants per plot, at the R1 phenological stage (Fehr & Caviness 1977). To this end, the plants were collected randomly, preserving the root system, with the aid of a cutting shovel and in soil with field capacity. After counting, the samples were dried in an oven with forced air circulation at 60 °C until they reached constant mass. Subsequently, the dry mass of nodules (DMN) was quantified by weighing on a precision scale. Then, phosphorus content in the leaf tissue (PL) was quantified through analysis of 35 trifoliate leaves, in four replications per treatment. Drying was carried out in a forced air circulation oven at a temperature of 65 °C until constant mass was obtained. Afterwards, they were ground and subjected to nitric perchloric digestion and spectrometric determination with vanadate yellow (Carmo et al. 2000).

The mass of one thousand grains (MTG) was determined on an analytical scale through weighing of 1000 soybean grains per plot. The grain yield (GY) of the central 6 m2 of the plot was quantified, with the grain mass corrected to 13% moisture and converted to kg ha-1. Protein (PCG), oil (OCG), fiber (FCG) and ash (ACG) content, as well as fatty acids such as palmitic acid (PACG), stearic acid (SACG), oleic acid (OACG), linoleic acid (LACG) and linolenic acid (LAG), were also determined in the soybean grain samples in an analysis conducted using a Near Infrared Spectrophotometer (Pasquini 2018).

Statistical analysis

Initially, the assumptions of multivariate normality were tested using the Shapiro-Wilk normality test generalized by (Royston 1982). When they were not met, data transformation was performed, in accordance with the Box-Cox family (Box & Cox 1964). However, due to the occurrence of non-normal multivariate distribution, data transformation (λ value) was performed, which proved to be efficient in overcoming problems related to violation of the statistical assumption.

Subsequently, the statistical assumption of multicollinearity was tested among the explanatory traits (for path analysis) and among the traits that make up each group of variables (for canonical correlation analysis). Multicollinearity diagnosis was performed considering the variance inflation factor (VIF) and condition number (CN) statistics.

To this end, the correlation coefficient matrix was generated by means of Pearson’s correlation; subsequently, the CN was obtained by the relationship between the largest and smallest eigenvalue of the X’X correlation matrix. When CN ≤ 100 is verified, the occurrence of weak multicollinearity is considered; while 100 < 1,000 is deemed moderate to severe multicollinearity, and CN ≥ 1,000, severe multicollinearity (Montgomery & Peck 1982). The VIF was obtained for each variable, on the inverse diagonal of the X’X correlation matrix; when a VIF value > 10 is obtained, the occurrence of severe multicollinearity is considered (Hair Jr. et al. 2009). Therefore, variables with VIF > 5 were excluded, not only for VIF > 10, when severe multicollinearity was already considered.

Path analyses were performed, considering GY as the main character, as a function of the explanatory characters (NN, DMN, MTG, PL, PCG, OCG, FCG, ACG, PACG, SACG, OACG, LACG and LAG (Cruz et al. 2012). The direct and indirect effects of the explanatory traits on grain yield were estimated through path analysis under multicollinearity condition (ridge path analysis), including a correction factor (k). To this end, it was considered that each explanatory traits has a direct effect on grain yield and acts indirectly through its effects on the other explanatory traits. However, a constant “k” was added to the diagonal of the X’X correlation matrix in order to reduce the variance associated with the least squares estimator of the path analysis. Thus, the X’X  = X’Y system of normal equations became (X’X + k)  = X’Y. The addition of values of the constant “k” was tested, and its lowest value from which the path coefficients stabilized was chosen (Cruz et al. 2012). This resulted in 2 forms of path analysis: excluding variables and including a correction factor (ridge path analysis).

Then, the canonical correlation analysis was performed (Hotelling 1936), in which the variables assessed were divided into 3 groups: morphological variables: NN and DMN; productive variables: MTG and GY; and bromatological variables: PL, PCG, OCG, FCG, ACG, PACG, SACG, OACG, LACG and LAG. The canonical correlations of the groups were compared by cross-canonical loadings, confronting the morphological variables with productive variables, productive variables with bromatological variables, and morphological variables with bromatological variables.

Due to the severe multicollinearity observed in the canonical correlation analysis, the palmitic acid (PACG), stearic acid (SACG), oleic acid (OACG), linoleic acid (LACG) and linolenic acid (LAG) variables were excluded in all scenarios assessed in order for the same set of variables to be obtained and for the objective of this study to be met.

Additionally, the effects of the parameters were isolated and removed from the mathematical model referring to the experimental design. For a bifactorial experiment, in the randomized block design, the mathematical model is characterized by: Yijk = m + ai + dj + (ad)ij + bk + eijk;, where Yijk is the value observed in the plot; m is the effect of the overall mean; ai is the effect of the first factor (Factor A), co-inoculation of growth-promoting bacteria; dj is the effect of the second factor (Factor D), phosphorus doses (P2O5); (ad)ij is the effect of the interaction (A×D) of the factors; bk is the random effect of the block, and eijk is the random effect of the experimental error of the experimental unit. The effects were stratified, removing the parameter effects from the mathematical model, with the following scenarios being worked with:

General)   Y i j k = m + a i + d j + ( a d ) i j + b k + e i j k ;   (Traditional)
Predicted)   Y i j k = m + e i j k   (No treatment effect) ;
A 1 )   Y i j k = m + a 1 + e i j k ; a 1 = co-inoculation of Bradyrhizobium spp. + Azospirillum spp . ;
A 2 )   Y i j k = m + a 2 + e i j k ; a 2 = co-inoculation of Bradyrhizobium spp. + Pseudomonas fluorescens ;
A 3 )   Y i j k = m + a 3 + e i j k ; a 3 = co-inoculation of Bradyrhizobium spp.+ Bacillus subtilis ;
A 4 )   Y i j k = m + a 4 + e i j k ; a 4 = co-inoculation of Bradyrhizobium spp. + Bacillus subtilis + Bacillus megaterium ;
A 5 )   Y i j k = m + a 5 + e i j k ; a 5 = co-inoculation of Bradyrhizobium spp. +Azospirillum spp. + Pseudomonas fluorescens ;
A 6 )   Y i j k = m + a 6 + e i j k ; a 6 = co-inoculation of Bradyrhizobium spp. + Azospirillum spp. + Bacillus subtilis ;
A 7 )   Y i j k = m + a 7 + e i j k ; a 7 = co-inoculation of Bradyrhizobium spp. + Azospirillum spp. + Bacillus subtilis + Bacillus megaterium ;
A 8 )   Y i j k = m + a 8 + e i j k ; a 8 = co-inoculation of Bradyrhizobium spp. +Azospirillum spp. + Pseudomonas fluorescens + Bacillus subtilis + Bacillus megaterium ;
A 9 )   Y i j k = m + a 9 + e i j k ;   a 9 = Control (no bacteria) ;
D 1 )   Y i j k = m + d 1 + e i j k ;   d 1 = dose of 0 kg h a 1 o f P 2 O 5 ;
D 2 )   Y i j k = m + d 2 + e i j k ;   d 2 = dose of 50 kg h a 1 o f P 2 O 5 ;
D 3 )   Y i j k = m + d 3 + e i j k ;   d 3 = dose of 100 kg h a 1 o f P 2 O 5 ;
D 4 )   Y i j k = m + d 4 + e i j k ;   d 4 = dose of 150 kg h a 1 o f P 2 O 5 .

Thus, there were a total of 14 stratification scenarios concerning combinations of removal of parameter effects from the mathematical model, which were compared to the general model (Traditional) in the path analysis and canonical correlation analysis.

For the aforementioned analyses, the criterion for comparing the path coefficients (direct and indirect) and cross-canonical loadings was the change in the sign of the coefficients (positive or negative) and a change greater than 50% (>50%) in the magnitude of the values, to be considered a significant change, for the scenarios of removing parameters from the model compared to the general model (Traditional), as established by (Sgarbossa et al. 2024a).

The ‘MVN’ statistical package was used to test the statistical assumptions in order to assess multivariate normality (Korkmaz et al. 2014), and to assess multicollinearity in the VIF methods with the ‘faraway’ package (Faraway 2022) and CN with ‘pracma’ (Borchers 2023). To generate the results of the path analysis and canonical correlation analysis, the ‘metan’ (Olivoto & Lúcio 2020) and ‘yacca’ (Butts 2022) statistical packages were used, and so were graphical representations with ‘ggplot2’ (Wickham 2016). In all statistical analyses, a 5% error probability level was adopted, using the R software (R Core Team 2024).

Weather conditions

Data on air temperature – minimum, average and maximum – and accumulated rainfall were obtained from the automatic meteorological station of the National Institute of Meteorology (INMET), linked to the Federal University of Santa Maria, located approximately 5,987 m away from the experimental area. During soybean cultivation, an average air temperature of 22.9 °C was recorded, ranging from 3.5 °C to 37.2 °C, with an accumulated rainfall of 490.1 mm. Therefore, during the cultivation period, weather conditions were favorable for the development of soybeans, since the air temperature, which results in better growth and development of the crop, is concentrated in the range of 20 °C to 30 °C, with water requirement varying from 450 mm to 800 mm in the cycle (Zanon et al. 2022).

RESULTS AND DISCUSSION

The multicollinearity diagnosis showed a violation of the statistical assumption in all scenarios assessed (Figure 1). In the path analysis excluding variables and considering VIF > 5, there was an alteration in the change referring to the occurrence of variables causing multicollinearity. Therefore, in the scenarios with removal of parameters from the mathematical model compared to the general model, a greater number of variables causing multicollinearity was observed, the majority of which were bromatological ones.

Figure 1
Path analysis of direct (diagonal) and indirect (off-diagonal) effects on grain yield (GY), excluding variables with VIF >5, which cause multicollinearity of the general model, and scenarios with the removal of parameters from the mathematical model. Legend: General mathematical model: a) Yijk = m + ai + dj + (ad)ij + bk + eijk; b) Predicted: Yijk = m + eijk; c) model Yijk = m + ai + eijk, of Bradyrhizobium spp., co-inoculated with: Azospirillum spp.; d) Pseudomonas fluorescens; e) Bacillus subtilis; f) Bacillus subtilis + Bacillus megaterium; g) Azospirillum spp. + Pseudomonas fluorescens; h) Azospirillum spp. + Bacillus subtilis; i) Azospirillum spp. + Bacillus subtilis + Bacillus megaterium; j) Azospirillum spp. + Pseudomonas fluorescens + Bacillus subtilis + Bacillus megaterium; k) Control; l) model Yijk = m + dj + eijk, of the phosphorus (P2O5) dose levels, 0 kg ha-1; m) 50 kg ha-1; n) 100 kg ha-1; o) 150 kg ha-1. GY: grain yield (kg ha-1); MTG: mass of one thousand grains (g); NN: number of nodules; DMN: dry mass of nodules; PL: phosphorus in the leaf (g kg-1); PCG: protein content in the grain (%); OCG: oil content in the grain (%); FCG: fiber content in the grain (%); ACG: ash content in the grain (%); PACG: palmitic acid content in the grain (%); SACG: stearic acid content in the grain (%); OACG: oleic acid content in the grain (%); LACG: linoleic acid content in the grain (%); LAG: linolenic acid in the grain (%).

For the general model, it was necessary to remove the LACG variable. For the predicted model, the removal of the ACG variable circumvented the violation of the statistical assumption. In scenario A1, the PCG, ACG, LACG and LAG variables were excluded. In A2, DMN, OCG, ACG, PACG and LACG were excluded. For scenario A3, the PCG, OCG, PACG, LACG and LAG variables were excluded. In A4, the OCG, FCG, PACG and LAG variables were excluded. For A5, the PCG, OCG, OACG and LAG variables were excluded. In A6, the DMN, PL, PCG, FCG, PACG, SACG and LACG variables were excluded to resolve multicollinearity. In scenario A7, the PL, ACG, OACG, LACG and LAG variables were excluded. For A8, the NN, FCG, PACG, SACG, OACG and LACG variables were excluded. Likewise, in A9, PL, PCG, ACG, OACG and LAG were excluded. For D1, PCG and LACG were removed. Similarly, in D2, PCG, PACG and LACG were removed. In D3, the PCG and LACG variables were removed. For scenario D4, PCG and OACG were removed, circumventing the multicollinearity problem.

Based on the occurrence of severe multicollinearity in all scenarios and the need to exclude different variables and different numbers of variables in each scenario, the choice was to proceed with the path analysis including a correction factor (ridge) (Figure 2).

Figure 2
Path analysis of direct (diagonal) and indirect (off-diagonal) effects with GY (grain yield), including a correction factor (k), ridge path analysis, for the general model and scenarios with the removal of parameters from the mathematical model. Legend: General mathematical model: a) Yijk = m + ai + dj + (ad)ij + bk + eijk; b) Predicted: Yijk = m + eijk; c) model Yijk = m + ai + eijk, of Bradyrhizobium spp., co-inoculated with: Azospirillum spp.; d) Pseudomonas fluorescens; e) Bacillus subtilis; f) Bacillus subtilis + Bacillus megaterium; g) Azospirillum spp. + Pseudomonas fluorescens; h) Azospirillum spp. + Bacillus subtilis; i) Azospirillum spp. + Bacillus subtilis + Bacillus megaterium; j) Azospirillum spp. + Pseudomonas fluorescens + Bacillus subtilis + Bacillus megaterium; k) Control; l) model Yijk = m + dj + eijk, of the phosphorus (P2O5) dose levels, 0 kg ha-1; m) 50 kg ha-1; n) 100 kg ha-1; o) 150 kg ha-1. GY: grain yield (kg ha-1); MTG: mass of one thousand grains (g); NN: number of nodules; DMN: dry mass of nodules; PL: phosphorus in the leaf (g kg-1); PCG: protein content in the grain (%); OCG: oil content in the grain (%); FCG: fiber content in the grain (%); ACG: ash content in the grain (%); PACG: palmitic acid content in the grain (%); SACG: stearic acid content in the grain (%); OACG: oleic acid content in the grain (%); LACG: linoleic acid content in the grain (%); LAG: linolenic acid in the grain (%).

Comparing the coefficients of determination (R²) (Table I) for the general model and the scenarios in which parameters were removed from the mathematical model, the form of analysis changes the quality of the model’s fit to the data. With the method, excluding variables that cause multicollinearity, the general model presented a coefficient of 0.27, while the stratification of effects relating to the levels of the first factor, co-inoculation of bacteria, represented 0.35 (A3) to 0.91 (A9). The levels of the phosphorus dose factor, in their turn, varied from 0.35 (D4) to 0.52 (D2).

Table I
Coefficient of determination (R²) for the scenarios of removing parameters from the mathematical model, percentage of change, compared to the general model.

With the method of including a correction factor (k) while maintaining the complete matrix of variables, the general model presented a coefficient of 0.26, while the removal of parameters from the mathematical model, working with the levels of the first factor, co-inoculation of bacteria, represented 0.49 (A4) to 0.90 (A9). The levels of the phosphorus dose factor varied from 0.32 (D4) to 0.50 (D2 and D3). For the predicted model, it represented only 0.12, with the exclusion of variables and the inclusion of a correction factor, showing that the best representations came from the individual effect of the treatments, with the removal of parameters from the mathematical model.

Compared to the general model, including a correction factor (k) in the path analysis matrix, the representation quality of the coefficient of determination (R²) increased from 86.93% (A4) to 245.18% (A9). For the second factor, the increase was from 23.62% (D4) to 91.90% (D2).

Traditional path analysis with variable elimination is more suitable than ridge path analysis, including a correction factor (k) in the diagonal matrix (Toebe & Filho 2013). As observed in pioneering studies, the exclusion of variables responsible for multicollinearity was more efficient than the inclusion of a correction factor, providing less multicollinearity and greater model precision (higher R² and lower residual) (Hoerl & Kennard 1970, Olivoto et al. 2017).

In the ridge path analysis (including a correction factor), as shown in Table II, the stratification of effects with the removal of parameters from the mathematical model compared to the general model, for the direct effects with GY of the levels of the first factor, co-inoculation of bacteria, caused a change in magnitude (>50%) from 7.69% (A3) to 61.64% (A2). As for the phosphorus dose factor levels, the changes were from 15.38% (D4) to 31.46 (D1 and D2). Considering the change in direction of the coefficients, for the levels of bacterial co-inoculation, the change ranged from 23.08% (A7) to 69.23% (A3 and A6). For phosphorus dose levels, there was a change from 30.77 (D2) to 69.23% (D4).

Table II
Change in magnitude and direction of the path analysis coefficients in the direct and indirect effects, with the main variable being grain yield (GY), of the scenarios of removing parameters from the mathematical model compared to the general model.

With respect to the indirect effects with GY, with the removal of parameters from the mathematical model, considering the first factor, there was a change in magnitude (>50%) from 30.13% (A5) to 47.44% (A2). For the second factor, there was a change from 37.82% (D4) to 51.28% (D2). However, it varied from 30.13% (A5) to 51.28 (D2). When analyzing the change in the direction of the associations, for the bacterial co-inoculation factor, there was a change from 42.95% (A2) to 55.77% (A5). For the dose factor, there was a change from 33.97% (D2) to 51.92% (D4). Thus, the removal of parameters from the mathematical model led to changes in the direction of the coefficients ranging from 33.97 (D4) to 55.77% (A5).

Similarly, Sgarbossa et al. (2024b) found that removing parameters from the model resulted in changes of 10.5% and 13.3% in direction, and 24.7% and 23.0% in magnitude, for the direct and indirect effects in path analyses. Likewise, by stratifying scenarios in the path analysis, they obtained a maintenance of only 3.30% and 20% in the linear correlation values, 3.30% and 30% in the direct effects and 7.33% and 24.67% in the indirect effects, with and without fungicide application (Sgarbossa et al. 2024a).

In the canonical correlation analysis by cross-canonical loads, for the group of morphological variables with productive variables (Table III), there was significance only for the first canonical pair of scenario D3, with 0.15 of determination coefficient (R²), considered low, as well as a canonical correlation (r) of 0.46. This resulted in a 98.99% increase in the model’s explanation (R²), as well as a 69.51% increase in the canonical correlation (r), in relation to the general model.

Table III
Canonical correlation analysis by cross-canonical loadings of the groups of morphological and productive variables for the general model and of the scenarios of removing parameters from the mathematical model.

In the canonical correlation analysis by cross-canonical loading, for the group of productive variables with bromatological variables (Table IV), there was high significance by the chi-square test at 5% probability of error for the first and second canonical pair (1st and 2nd CP) of the general model, with a coefficient of determination (R²) of 0.35 and 0.21 and canonical correlation (r) of 0.59 and 0.45 for the 1st and 2nd canonical pair, respectively. There was also statistical significance for scenario A6 in the 1st and 2nd CP with R² of 0.73 and 0.59, as well as canonical correlation (r) of 0.85 and 0.77, respectively. Likewise, for scenario D3, in the 1st CP, there was a significant effect with R² of 0.44 and canonical correlation (r) of 0.66.

Table IV
Canonical correlation analysis by cross-canonical loads of the groups of productive and bromatological variables for the general model and of the scenarios of removing parameters from the mathematical model.

Therefore, the effect of treatment A6 (Bradyrhizobium spp. + Azospirillum spp. + Bacillus subtilis) presented, in addition to significance, a higher R² of 0.73 and 0.59, respectively, in the 1st and 2nd CP, of the relationship between the group of productive and bromatological variables, as well as a high correlation, of 0.87 and 0.77, as previously mentioned. The other two scenarios, general model (traditional) and D3 (dose of 100 kg ha-1 P2O5), although statistically significant, explain little of the relationship between the groups.

Comparing the change (%) in the magnitude of the coefficient of determination (R²), of scenario A6 (Bradyrhizobium spp. + Azospirillum spp. + Bacillus subtilis) in relation to the general model, there was an increase in explanation of 108.57% and 180.95, respectively, in the 1st and 2nd CP, as well as 44.07% and 71.11% in the canonical correlation (r). In scenario D3, in relation to the general model, the increase in R² was 25.71%, in the 1st CP, and 11.86% in the canonical correlation.

In the canonical correlation analysis by cross-canonical loading, for the groups of morphological variables with bromatological variables (Table V), there was significance for the 1st CP of scenario A1 with a coefficient of determination (R²) of 0.79 and canonical correlation (r) of 0.89. There was also statistical significance for scenario A9 in the 1st CP with R² of 0.75, as well as a canonical correlation (r) of 0.86. Likewise, scenario D4 in the 1st CP had a significant effect with R² of 0.38 and canonical correlation (r) of 0.62.

Table V
Canonical correlation analysis by cross-canonical loads of the groups of morphological and bromatological variables for the general model and of the scenarios of removing parameters from the mathematical model.

Therefore, the effect of treatment A1 (Bradyrhizobium spp. + Azospirillum spp) presented, in addition to significance, a higher R², with 0.79, and high canonical correlation (0.89), of the morphological and bromatological variables. Comparing the change in explanation (R²) of scenario A1 (Bradyrhizobium spp. + Azospirillum spp.) in relation to the general model, in the 1st CP, there was an increase of 338.89% in R² and 111.90% in the canonical correlation (r).

However, scenarios A9 (Control - no use of bacteria) and D3 (dose of 150 kg ha-1 P2O5), despite being statistically significant, explain less of the relationship between the groups. In scenarios A9 and D4 in relation to the general model, the increase in R² was 316.67% and 111.11%, and, in the canonical correlation (r), it stood at 104.76% and 47.62%, respectively, in the 1st CP.

Studies in soybean crops using canonical correlations have found that the agronomic and physiological characteristics of seeds are not independent, since more productive plants, with a greater number of pods and high oil content, are associated with seeds with a high germination percentage and emergence rate (Pereira et al. 2017). Also, analyzing segregating F5 soybean families, proving to be promising in breeding programs in terms of physiological and nutritional aspects, with characteristics such as crude protein content, crude fiber, height of insertion of the first pod, plant height and number of pods with 4 grains being dominant and determinant for the establishment of segregating generations (Carvalho et al. 2021).

In beans, canonical correlations were estimated between the group of variables consisting of primary and secondary agronomic traits; the number of pods per plant, pod length and plant type traits showed high and positive magnitude, making it possible to conclude that plants with a greater number of pods, precocity and erect architecture should be selected to increase grain yield (Abreu et al. 2021).

Therefore, the importance of studies that stratify effects to better understand the linear relationships between variables is reinforced, since the conclusions may be redundant based on the general model. Based on this study, the effects of parameters originating from the mathematical model, originating from the experimental design, should be considered in multivariate analyses, as they influence the linear relationships between traits. Analyzing the effects of the factors under study individually allows for discovering valuable information about their effect on linear relationships, in addition to helping in the planning of experiments when it comes to choosing traits to be assessed that are relevant in this area and that are influenced by the treatments. However, further studies are needed, in different environments and scenarios.

CONCLUSIONS

In the ridge path analysis, the removal of parameters from the mathematical model increased the explanatory capacity of the models from 23.62% to 245.18%.

In the path analysis, the removal of parameters from the mathematical model in the direct effects with grain yield caused a change in magnitude from 7.69% to 61.64%, and in direction from 23.08% to 69.23%. In the indirect effects, there was a change in the magnitude of the coefficients from 30.13% to 51.28% and from 33.97% to 55.77% in the direction, in relation to the general model.

In the canonical correlations, in relation to the general model, in the productive variables with bromatological variables, treatment Bradyrhizobium spp. + Azospirillum spp. + Bacillus subtilis significantly increased the explanatory capacity, 108.57% and 180.95%, in the 1st and 2nd canonical pair. Similarly, the latter increased by 44.07% and 71.11% in canonical correlation. For the group of morphological variables with bromatological variables, Bradyrhizobium spp. + Azospirillum spp. increased by 338.89% in explanatory capacity and 111.90% in canonical correlation, for the 1st canonical pair.

Acknowledgements

We thank the members of the Coxilha Large Crop Management Research Group and the Agricultural Experimentation Research Group of the Federal University of Santa Maria for their assistance with this project.

  • Data availability
    The datasets generated and analyzed during the current study are available from the corresponding author upon reasonable request.

References

  • ABREU HKA, FACHINELLI R & CECCON G. 2021. Canonical correlation among morphological traits and yield components of cowpea. Agrarian 14: 314-322.
  • BLALOCK HM. 1963. Correlated Independent Variables: The Problem of Multicollinearity. Social Forces 42: 233-237.
  • BORCHERS HW. 2023. pracma: Practical Numerical Math Functions. R package version 2.4.4. Available at: https://cran.r-project.org/package=pracma
    » https://cran.r-project.org/package=pracma
  • BOX GEP & COX DR. 1964. An Analysis of Transformations. JRSSB (Methodological) 26: 211-252.
  • BUTTS CT. 2022. yacca: Yet Another Canonical Correlation Analysis Package. Vienna: Comprehensive R Archive Network (CRAN). R package version 1: 4-2. Available at: https://CRAN.R-project.org/package=yacca.
  • CARAFFA M, PIRES JLF, RUGERI AP, RIFFEL CT, HARTER LSH, DANIELOWSKI R & PIZZANI R. 2019. Indicações técnicas para a cultura da soja no Rio Grande do Sul e em Santa Catarina, safras 2018/2019 e 2019/2020. Três de Maio/RS: Setrem, 105 p.
  • CARMO CAFS, ARAÚJO A, CARLOS A, BERNARDI C, FRANCISCO M & SALDANHA C. 2000. Métodos de análise de tecidos vegetais, Rio de Janeiro: Embrapa Solos, 41 p.
  • CARVALHO IR, DE SOUZA VQ, NARDINO M, FOLLMANN DN, SCHMIDT D & BARETTA D. 2015. Correlações canônicas entre caracteres morfológicos e componentes de produção em trigo de duplo propósito. Pesqui Agropecu Bras 50: 690-697.
  • CARVALHO IR, SILVA JAG, LORO MV, SARTURI MVDR, HUTRA DJ & LAUTENCHLEGER F. 2021. Inter-relações nutracêuticas canônicas da soja e suas reflexões sobre o melhoramento genético. Agropecuária Catarinense 34: 67-75.
  • CONAB - COMPANHIA NACIONAL DE ABASTECIMENTO. 2024. Acompanhamento da Safra Brasileira de Grãos, 12 ed., Brasília, DF, 117 p.
  • CQFS - COMISSÃO DE QUÍMICA E FERTILIDADE DO SOLO. 2016. Comissão de Química e Fertilidade do Solo. Sociedade Brasileira de Ciência do Solo - Núcleo Regional Sul, 11th ed.
  • CRUZ CD, REGAZZI A & CARNEIRO P. 2012. Modelos biométricos aplicados ao melhoramento genético. 4th ed. Viçosa, MG: Editora UFV, p. 514.
  • DAHER RF, PEREIRA AV, MENEZES BR S, CASSARO S, NOVO AAC, FURLANI EP, JÚNIOR ATA, PEREIRA MG, STIDA WF & VIDAL AKF. 2018. Canonical correlations among morpho-agronomic and chemical traits in hybrids between elephant grass and millet. Aust J Crop Sci 12: 210-216.
  • DIEL MI ET AL. 2020. Canonical correlations in agricultural research: Method of interpretation used leads to greater reliability of results. Int J Innov Educ Res 8: 171-181.
  • FAO - FOOD AND AGRICULTURE ORGANIZATION. 2024. FAOSTAT: Crops and livestock products. Rome: Food and Agriculture Organization of the United Nations. Available at: https://www.fao.org/faostat/en/#data/QCL
    » https://www.fao.org/faostat/en/#data/QCL
  • FARAWAY J. 2022. Faraway: Functions and Datasets for Books by Julian Faraway. Vienna: Comprehensive R Archive Network (CRAN).
  • FEHR WR & CAVINESS CE. 1977. Stages of soybean development, Ames, Yowa. Yowa State University of Science and Technology, 11 p.
  • FOLLMANN DN, CARGNELUTTI FILHO A, QUEIRÓZ DE SOUZA V, NARDINO M, CARVALHO IR, DEMARI GH, FERRARI M, DE PELEGRIN AJ & SZARESKI VJ. 2017. Relações lineares entre caracteres de soja safrinha. Revista de Ciências Agrárias 40: 213-221.
  • GRAHAM MH. 2003. Confronting multicollinearity in ecological multiple regression. Ecology 84: 2809-2815.
  • HAIR JR JF, BLACK WC, BABIN BJ, ANDERSON RE & TATHAM RL. 2009. Análise de Correlação Canônica. Análise Multivariada de Dados. 6th ed. Porto Alegre: Bookman, p. 455-498.
  • HOERL AE & KENNARD RW. 1970. Ridge Regression: Biased Estimation for Nonorthogonal Problems. Technometrics 12: 55-67.
  • HOTELLING H. 1936. Relations between two sets of variates. Biometrika 28: 321-377.
  • KORKMAZ S, GOKSULUK D & ZARARSIZ G. 2014. MVN: An R Package for Assessing Multivariate Normality. R package version 5.8. Available at: https://cran.r-project.org/package=MVN
    » https://cran.r-project.org/package=MVN
  • MONTGOMERY DC & PECK EA. 1982. Introduction to linear regression analysis. New York: John Wiley & Sons, p. 504.
  • OLIVOTO T & LÚCIO AD. 2020. Metan: An R package for multi-environment trial analysis. Methods Ecol Evol 11: 783-789.
  • OLIVOTO T, SOUZA VQ, NARDINO M, CARVALHO IR, FERRARI M, PELEGRIN AJ, SZARESKI VJ & SCHMIDT D. 2017. Multicollinearity in Path Analysis: A Simple Method to Reduce Its Effects. Agron J 109: 131-142.
  • PASQUINI C. 2018. Near infrared spectroscopy: A mature analytical technique with new perspectives – A review. Anal Chim Acta 1026: 8-36. https://doi.org/10.1016/j.aca.2018.04.004.
    » https://doi.org/10.1016/j.aca.2018.04.004
  • PAVINATO PS, CHERUBIN MR, SOLTANGHEISI A, ROCHA GC, CHADWICK DR & JONES DL. 2020. Revealing soil legacy phosphorus to promote sustainable agriculture in Brazil. Scientific Reports 10: 1-11.
  • PEREIRA EM, SILVA FM, VAL BHP, PIZOLATO NETO A, MAURO AO, MARTINS CC & UNÊDA-TREVISOLI SH. 2017. Canonical correlations between agronomic traits and seed physiological quality in segregating soybean populations. Genet Mol Res 16(2). doi: 10.4238/gmr16029547.
  • R CORE TEAM. 2024. R: A Language and Environment for Statistical Computing. Available at: https://www.r-project.org/
    » https://www.r-project.org/
  • ROYSTON JP. 1982. An Extension of Shapiro and Wilk’s W Test for Normality to Large Samples. JRSS Series C 31: 115-124.
  • SANTOS HG, JACOMINE PKT, ANJOS LHC & OLIVEIRA VA. 2018. Sistema brasileiro de classificação de solos, 5th ed., Brasília, DF: Embrapa, p. 356.
  • SGARBOSSA J ET AL. 2024a. Implications of removing model parameters on the linear relationships of trials with oat. Cienc Rural 54(12): e20230379. http://doi.org/10.1590/0103-8478cr20230379.
    » https://doi.org/10.1590/0103-8478cr20230379
  • SGARBOSSA J, LÚCIO ADC, DA SILVA JAG, CARON BO, DIEL MI, OLIVOTO T, NARDINI C, ALESSI O & LAMBRECHT DM. 2024b. Multivariate assumptions and effect of model parameters in path analysis in oat crop. Crop Pasture Sci 75. DOI: 10.1071/CP23135.
  • SILVA CM, LIMA GW, MEZZOMO HC, SIGNORINI VS, OLIVEIRA AB & NARDINO M. 2022. Canonical correlations between high and low heritability wheat traits via mixed models. Cienc Rural 53(2): e20210798. https://doi.org/10.1590/0103-8478cr20210798.
    » https://doi.org/10.1590/0103-8478cr20210798
  • TAPIADOR FJ, MORENO R & NAVARRO A. 2019. Consensus in climate classifications for present climate and global warming scenarios. Atmos Res 216: 26-36.
  • TOEBE M & FILHO AC. 2013. Não normalidade multivariada e multicolinearidade na análise de trilha em milho. Pesqui Agropecu Bras 48: 466-477.
  • WENDT L, MALAVASI MM, DRANSKI JAL, MALAVASI UC & GOMES JUNIOR FG. 2017. Relação entre testes de vigor com a emergência a campo em sementes de soja. Agrária 12: 166-171.
  • WICKHAM H. 2016. ggplot2: Elegant Graphics for Data Analysis. Springer-Verlag New York. https://doi.org/10.1007/978-3-319-24277-4
    » https://doi.org/10.1007/978-3-319-24277-4
  • WRIGHT S. 1923. The Theory of Path Coefficients a Reply to Niles’s Criticism. Genetics 8(3): 239-255. doi: 10.1093/genetics/8.3.239.
    » https://doi.org/10.1093/genetics/8.3.239
  • ZANON A, SILVA M, TAGLIAPIETRA E, CERA J & BEXAIRA K. 2022. Ecofisiologia da soja visando altas produtividades. 2nd ed., Santa Maria, 432 p.
  • ZUFFO AM, AGUILERA JG, RATKE RF, STEINER F, OLIVEIRA AM & FONSECA WL. 2020. Análise de trilha em soja submetida a fontes e doses de nitrogênio inoculadas com Bradyrhizobium japonicum em solos com diferentes níveis de fertilidade. Res Soc Dev 9(7): e203973813. DOI: https://doi.org/10.33448/rsd-v9i7.3813.
    » https://doi.org/10.33448/rsd-v9i7.3813

Data availability

The datasets generated and analyzed during the current study are available from the corresponding author upon reasonable request.

Publication Dates

  • Publication in this collection
    16 Feb 2026
  • Date of issue
    2026

History

  • Received
    14 Jan 2025
  • Accepted
    18 Aug 2025
location_on
Academia Brasileira de Ciências Rua Anfilófio de Carvalho, 29, 3º andar, 20030-060 Rio de Janeiro RJ Brasil, Tel: +55 (21) 2391-7901 - Rio de Janeiro - RJ - Brazil
E-mail: aabc@abc.org.br
rss_feed Stay informed of issues for this journal through your RSS reader
Go to top Report error