Multiple linear regression (Model fitness and prediction viability)
The full regression model (M
1) exhibited a high R
2 value of 0.648, indicating that approximately 65% of the variability in seed yield could be explained by the combination of selected traits in Table 1. This shows a strong predictive capacity of the model, especially in the context of biological and agricultural research where multiple factors interact. The Durbin-Watson statistic of 1.895 suggests that residuals are not autocorrelated, confirming the independence of errors: an essential assumption for reliable regression modelling. In contrast, the null model (M
0), which included no predictors, had an R
2 of 0.000, clearly demonstrating the improvement brought by the inclusion of agronomic variables. This model provides a strong foundation for predicting seed yield based on key morphological traits, essential for yield forecasting and varietal improvement.
Model significance
Table 2 of ANOVA (F = 32.441, p<.001) confirms that the model is statistically significant. This means that the set of predictor variables collectively has a strong effect on seed yield and the observed result is unlikely to be due to chance. This is a crucial validation step showing that the model is not only fitting well but is also meaningful in terms of its statistical inference.
Contribution of individual predictors
From the regression coefficients in Table 3, it was found that the Number of pods per plant emerged as the most influential predictor (β = 0.698, p<0.001). This suggests that increasing the number of pods has a substantial positive impact on seed yield. This is intuitive and aligns with earlier agronomic studies, which identify pod number as a direct determinant of reproductive success in legumes. Number of seeds per pod also had a strong, statistically significant influence (β = 0.314, p<.001), reinforcing the role of reproductive efficiency.
Similarly, a significant but moderate contribution was shown by pod length, which indicated that longer pods with more seeds probably (β = 0.176, p = 0.025). A negligible and non-significant effects of plant height and number of primary branches was found (p = 0.459 and 0.978, respectively). The findings states that these traits do not influence seed yield directly
(Sree et al., 2025), in this context of genotype environment and can serve more in a structural role
(Marzhan et al., 2022). To improve field pea productivity in breeding programs with targeted trait selection, the distinction between yield contributing and non-contributing traits is important (
Tiwari, 2020). Further studies should emphasize towards pod-related traits rather than vegetative growth traits.
Descriptive statistics
Basic statistics is required essentially in regression modelling to get foundational insights in the central tendency and variability of agronomical traits in field pea sample population (n=94).
Seed yield per plant
Moderate variability among genotypes was found with average seed yield per plant of 4.93 gms, with standard deviation of 0.86 gms. A precise population mean and adequate sample size was suggested by a small standard error of 0.089 gm. Therefore, it seems a potential in selection of high yielding plants through breeding.
Plant height
A variation in vertical growth habit was suggested with average plant height of 87.33 cm and 17.6 cm of standard deviation. Although, plant height only can relate to biomass as shown in results (Table 3), indicating that taller plants necessarily do not have effect on seed yield.
Number of primary branches per plant
A mean of 2.08 and low variability; 0.475 of standard deviation shows uniformity in this trait under existing environmental condition. It reveals that this trait is less crucial for yield improvement due to its low variability.
Pods per plant
A good diversified range of reproductive output was found with mean of 7.69 per plant and standard deviation of 1.76 in number of pods per plant. This provides a wide range to breeders with scope for selection and improvement of high pod bearing varieties, suggesting strongest correlation with seed yield in the regression model.
Number of seeds per pod
The mean seeds per pod was 3.54 and SD = 0.518 suggests moderate variability. Since this trait reflect a positive, strong and valuable impact on seed yield (Table 3), increasing number of seed per pod could be a direct strategy to increase yield.
Pod length
The averaged length of pod is 3.70 cm, with 1.01cm standard deviation, showing a significant degree of variability. This characteristic may be an indirect selection criterion, potentially affecting the quantity or size of seeds housed within, given its notable yet moderate impact on production. Hence, the descriptive statistics indicate that the characteristics like pods per plant, seeds per pod and pod length do not only exhibit enough variation to enable selection, but also supports the positive predictive power of this characteristic in the regression analysis. Traits like branching and plant height are more durable but influence of yield is very less. As a result, Table 4 attests to the dataset’s significant diversity in yield-contributing characteristics, which is necessary for significant regression and PCA analysis. Additionally, plant breeders aiming to optimize reproductive traits might use these statistics as baseline benchmarks.
Diagnostic plots
Residual diagnostics
Residuals versus Predicted diagnostics shows no pattern which supports homoscedasticity [Fig 1A (a-d)]. Standard which we found the Standardized Residuals to be approximately normal and also our Q-Q Plot which reported that the data points fell along a diagonal line which in turn confirmed normality. Also, we saw Random Scatter in Residuals versus Covariate plots which in turn confirmed the linear relationship between predictors and seed yield as well as the assumption of homoscedasticity. In terms of the overall Assessment of Model Assumptions; that is residual distribution and normality we had support in our diagnostic plots which in turn validated the model assumptions
(Ribeiro et al., 2025). We noted that the standardized residual histogram was normal. Also, in the Q-Q plot we had residual points which fell along the diagonal which in turn confirmed normality. That we saw no out of the ordinary patterns in Residuals versus Predictors plot which in turn supported the assumption of homoscedasticity. These results in turn gave us confidence in the regression estimates and their large-scale use. Similarly, in Residuals versus Individual Covariate diagnostics [Fig 1B (a-e)], the Scatter plots of residuals against each covariate showed random dispersion with no systematic trends, confirming the linearity and independence of predictors. This is particularly important in agricultural datasets where multicollinearity and interaction effects can distort results.
Individual predictor analysis (Partial and marginal effect plots)
The strong, positive linear relationships with seed yield demonstrated by pod per plant and seeds per pod, even after adjusting for other variables. These are the strong indicators of potential yield and should be prioritized in breeding. Pod length showed a moderate relationship, indicating an auxiliary role. Also, plant height and branches per plant had almost flat trends, reinforcing their lack of predictive utility for yield.
Principal component analysis (PCA)
The results from PCA are in Table 5, which shows that; overall, MSA is above 0.5, acceptable for PCA. However, harvest index and biological yield have low values (<0.5). Pods per plant (0.718) and Plant height (0.804) show strong sampling adequacy. Biological yield (0.377) and Harvest index (0.330) have low MSA values (<0.5), meaning these variables may not fit well into the component structure.
Table 6 results show that, Bartlett’s test is significant, supporting factorability of the correlation matrix. The observations in Bartlett’s test (χ
2 = 859.547, p<0.001) confirms the dataset is suitable for PCA and the model χ
2 = 347.024, p<0.001, supports factor model validity.
Table 7 show that the Seed yield loads moderately on two components (RC1 and RC2), showing it is influenced by both yield structure and physiological maturity traits. PCA was used to explore the multivariate structure of yield-related traits and reduce dimensionality.
The kaiser-meyer-olkin (KMO) measure of 0.513 and significant Bartlett’s test (p<.001) confirmed that the data were suitable for PCA. However, the biological yield (0.377) and harvest index (0.330) exhibited inadequate sample adequacy and should be taken cautiously.
Component loadings and trait clustering
Promax-rotated loadings in Table 8, revealed the following: Vegetative growth dimension showed by RC1 by loading strongly on biological yield and plant height. The traits like pods per plant, seed yield and harvest index, aligning with yield structure is showed by RC2. In RC3 and RC4 includes reproductive precision traits like seeds per pod and pod length. Interestingly in both RC1 and RC2 we see seed yield which tells us of a complex relationship between growth and reproductive elements. This puts forth that it is a function of both size of the biomass and also separate yield related traits.
Many different factors play into the determination of seed yield which in turn indicates that it isn’t a product of a single group of variables. What we see in the diverse grouping of traits like biological yield, plant height and pods per plant is a reflection of the very many aspects that go into the formation of yield
(Patel et al., 2023).
The first three components explain over 56% of total variance, supporting dimensional reduction and helping focus on dominant traits. In the figure,
i.e., the Scree Plot shows a sharp drop after 3 components. It justifies retaining 3-4 principal components (above Kaiser’s rule of eigenvalue >1), which confirms retention of 3 components as ideal. In the figure,
i.e., the path diagram shows the causal paths between the latent variable (PC) and the observed variable, confirming the multi-trait effect on seed yield, showing how yield is influenced by multiple underlying factors rather than a single dominant trait.
Parallel analysis
Components to retain (Parallel analysis and scree plot)
The scree plot showed a clear bend after the third component and parallel analysis confirmed the persistence of three components based on actual versus simulated eigenvalues (Fig 2). These components together explained 56.9% of the total variance, which suggests that a significant portion of trait variability can be explained through the three major axes (Table 9). This dimensionality reduction simplifies future modelling and trait-targeting efforts by highlighting the core structures underlying traits
(Angel et al., 2021).
Path diagram
The path diagram depicted (Fig 3) how the latent components (PCs) influence the observed traits. This diagram presents that many traits which are integrated determine seed yield which in turn presents the multi-dimensional aspect of plant productivity (
Ranjani and Jayamani, 2024;
Jain et al., 2023). In the development of selection indices in crop improvement programs which require at the same time the improvement of related traits this is a useful tool
(Kumar et al., 2024).
In our study we identified number of pods per plant, number of seeds per pod and pod length as the main determinants of seed yield in pea crop. The model we present is very statistical in nature and we did test its assumptions in great detail. PCA, we used to bring out some meaningful trait groups which in turn simplified the complex trait relationships.